Updating References · The panel it consults

How we find the newer version: seven engines, one verdict

CitationLab Team · August 2026 · 8 min read
ASKED IN ORDER OF HOW STRUCTURED THE ANSWER IS CROSSREF — the publishers' registry PUBMED — curated, biomedical SEMANTIC SCHOLAR — links versions OPENALEX — open, wide GOOGLE SCHOLAR — the long tail GOOGLE BOOKS — editions WEB PASS — grey literature metadata quality degrades down this list — so the top is asked first RECONCILED same work collapsed, identifiers preferred one ranked set A PROPOSAL agreement is evidence, not authority AND THE LEDGER PRINTS WHAT WAS DISCARDED Crossref asked 41 returned 38 taken 22 the gap is the screening doing its work AN ENGINE THAT ANSWERED NOTHING STILL GETS A ROW switched off, declined the subject, or found nothing — all three are facts worth having

Asking one database whether a newer version of a paper exists gets you one database's answer. Asking seven and reconciling them is a different question, and the order in which you ask turns out to matter more than the number.

This is the panel the update and find passes consult, why it is ordered the way it is, and what the ledger afterwards tells you that the result alone does not.

The panel

Who is asked, and what each is good at
  CROSSREF          the publishers' own registry
                    DOIs, version relationships,
                    the record of record

  PUBMED            biomedical literature, curated
                    declines politely on anything
                    outside its subject

  SEMANTIC SCHOLAR  broad coverage, good at linking
                    versions of the same work

  OPENALEX          open catalogue, wide, strong
                    where publishers are small

  GOOGLE SCHOLAR    finds what indexes miss —
                    theses, working papers, the
                    long tail

  GOOGLE BOOKS      books and edition statements

  WEB PASS          grey literature: reports,
                    standards, government pages

Seven, and they overlap heavily. That is deliberate — the point is not coverage for its own sake but corroboration: a candidate returned by three independent sources is a different kind of answer from one returned by a single scrape.

Free-first is a design rule, not an economy

The order is fixed, and it runs from the most structured and openly licensed sources to the least. Crossref, then the curated and open catalogues, then the scraped tiers, then the general web.

The obvious reason is cost, and it is the less interesting one. The real reason is that the quality of the metadata degrades in exactly that order.

A Crossref record is deposited by the publisher: the authors are the authors, the DOI is authoritative, and where a preprint has a published version the relationship is explicitly recorded. A scraped result is a best reading of a rendered page — usually right, occasionally a conference listing mistaken for an article. If you take the structured answer when there is one, you never have to adjudicate between a good record and a plausible one.

Ask the publisher's registry first and the scraped tiers rarely get a turn. That is the point of the ordering — not thrift, but that you should not have to weigh a deposited record against a parsed page.

Formatting follows the same rule. Where several sources describe the same work, the entry is built from the most structured record available — Crossref, then Semantic Scholar, then OpenAlex — rather than from whichever answered fastest.

Seven answers, one verdict

The engines are queried in parallel and return candidates, not conclusions. Reconciling them is a separate step, and it is deterministic: candidates describing the same work are collapsed, identifiers are preferred over titles as evidence, and what emerges is a small ranked set rather than seven lists.

Then, and only then, does anything reach you — as a proposal, with the original beside it, for you to accept or reject. Nothing is swapped because several databases agreed. Agreement is evidence, not authority. The same principle governs the missing-reference search described in how to find citations missing from your bibliography.

Run the panel against your own bibliography — extraction and cross-check are free

The ledger, and why it prints the failures

After a run, every engine reports: how many times it was asked, how much came back, how much was actually used, and how many calls errored.

An engine ledger from one update run
  engine             asked  returned  taken  errors

  Crossref              41       38      22       0
  Semantic Scholar      41       31       9       0
  OpenAlex              41       27       4       0
  PubMed                12        7       2       0
  Google Scholar        19       11       1       0
  Google Books           0        0       0       0
  Web pass               6        3       0       0

  the gap between RETURNED and TAKEN is
  the screening doing its work

Two things are visible there that a bare result would hide.

The first is that gap. Crossref returned thirty-eight candidates and twenty-two were used — the rest were rejected as not the same work. A tool reporting only what it accepted would look more decisive and tell you nothing about how much it discarded.

The second is the row of zeros. An engine that was asked nothing was switched off, or declined the subject; an engine asked and returning nothing found nothing. Both are facts worth having, and both are invisible unless the ledger prints the engines that contributed no answers. A panel that only listed its successes would be a panel you could not audit.

What "asked" actually means for each source

The engines are not queried identically, and the differences are the reason the panel works rather than merely being long.

Not every engine is asked the same question
  BY IDENTIFIER
  where your entry already has a DOI, the
  registry is asked about that DOI — an
  exact question with an exact answer

  BY WORK
  where there is no identifier, the query is
  authors plus title plus year, and several
  catalogues are asked in parallel

  BY SUBJECT FIT
  PubMed is only asked where the citation
  looks biomedical — asking it about a
  literature-theory paper wastes a call and
  returns noise

  NOT ASKED AT ALL
  an engine switched off in configuration,
  or one whose subject the citation is not

That last category is why the ledger matters. An engine with zero in the "asked" column did not fail — it was never the right question. Reading a row of zeros as a broken integration is the most common misreading of the ledger, and the column headings are what prevent it.

Why not simply use the biggest index?

Because the failure modes are not shared. Each source is missing different things — a small publisher absent from one catalogue, a discipline underrepresented in another, a thesis that exists only in an institutional repository.

Querying several is not about being thorough for its own sake. It is that the gaps do not line up, so what one misses another usually holds. And when none of them holds it, the answer comes back empty — which is a true answer, and the one a single-source search is least able to give you honestly.

What this does not claim

More engines do not make a proposal correct. Seven sources agreeing that a newer paper exists says nothing about whether it answers the same claim the original was cited for — and that judgement stays with you, which is why nothing is applied without your decision. The reasoning around what makes a replacement legitimate is in updating references before submission.

Nor is a coverage gap a defect in your bibliography or any suggestion of misconduct. A source no catalogue holds is usually a legitimate work published somewhere small, and the honest result is the tool saying it found nothing rather than offering the nearest thing it could see.

See the ledger from your own run — free to start
Filed under: Updating References academic-engines crossref human-in-the-loop
Share: Post on X Share Email

Keep reading