How we find the newer version: seven engines, one verdict
CitationLab Team · August 2026 · 8 min read
Asking one database whether a newer version of a paper exists gets you one
database's answer. Asking seven and reconciling them is a different question, and the order in
which you ask turns out to matter more than the number.
This is the panel the update and find passes consult, why it is ordered the way it is, and
what the ledger afterwards tells you that the result alone does not.
The panel
Seven, and they overlap heavily. That is deliberate — the point is not coverage for its own
sake but corroboration: a candidate returned by three independent sources is a different kind of
answer from one returned by a single scrape.
Free-first is a design rule, not an economy
The order is fixed, and it runs from the most structured and openly licensed sources to the
least. Crossref, then the curated and open catalogues, then the scraped tiers, then the general
web.
The obvious reason is cost, and it is the less interesting one. The real reason is that
the quality of the metadata degrades in exactly that order.
A Crossref record is deposited by the publisher: the authors are the authors, the DOI is
authoritative, and where a preprint has a published version the relationship is explicitly
recorded. A scraped result is a best reading of a rendered page — usually right, occasionally a
conference listing mistaken for an article. If you take the structured answer when there is one,
you never have to adjudicate between a good record and a plausible one.
Ask the publisher's registry first and the scraped tiers rarely get a turn. That is
the point of the ordering — not thrift, but that you should not have to weigh a deposited record
against a parsed page.
Formatting follows the same rule. Where several sources describe the same work, the entry is
built from the most structured record available — Crossref, then Semantic Scholar, then
OpenAlex — rather than from whichever answered fastest.
Seven answers, one verdict
The engines are queried in parallel and return candidates, not conclusions. Reconciling them
is a separate step, and it is deterministic: candidates describing the same work are collapsed,
identifiers are preferred over titles as evidence, and what emerges is a small ranked set rather
than seven lists.
Then, and only then, does anything reach you — as a proposal, with the original beside it, for
you to accept or reject. Nothing is swapped because several databases agreed. Agreement is
evidence, not authority. The same principle governs the missing-reference search described in
how to find citations missing from your
bibliography.
The ledger, and why it prints the failures
After a run, every engine reports: how many times it was asked, how much came back, how much
was actually used, and how many calls errored.
Two things are visible there that a bare result would hide.
The first is that gap. Crossref returned thirty-eight candidates and twenty-two were used —
the rest were rejected as not the same work. A tool reporting only what it accepted would look
more decisive and tell you nothing about how much it discarded.
The second is the row of zeros. An engine that was asked nothing was switched off, or declined
the subject; an engine asked and returning nothing found nothing. Both are facts worth having,
and both are invisible unless the ledger prints the engines that contributed no answers. A panel
that only listed its successes would be a panel you could not audit.
What "asked" actually means for each source
The engines are not queried identically, and the differences are the reason the panel works
rather than merely being long.
That last category is why the ledger matters. An engine with zero in the "asked" column did
not fail — it was never the right question. Reading a row of zeros as a broken integration is
the most common misreading of the ledger, and the column headings are what prevent it.
Why not simply use the biggest index?
Because the failure modes are not shared. Each source is missing different things — a small
publisher absent from one catalogue, a discipline underrepresented in another, a thesis that
exists only in an institutional repository.
Querying several is not about being thorough for its own sake. It is that the gaps do not
line up, so what one misses another usually holds. And when none of them holds it, the answer
comes back empty — which is a true answer, and the one a single-source search is least able to
give you honestly.
What this does not claim
More engines do not make a proposal correct. Seven sources agreeing that a newer paper exists
says nothing about whether it answers the same claim the original was cited for — and that
judgement stays with you, which is why nothing is applied without your decision. The reasoning
around what makes a replacement legitimate is in
updating references before submission.
Nor is a coverage gap a defect in your bibliography or any suggestion of misconduct. A source
no catalogue holds is usually a legitimate work published somewhere small, and the honest result
is the tool saying it found nothing rather than offering the nearest thing it could see.