It missed 27 citations, then told me it had found the last one
CitationLab Team · August 2026 · 7 min read
ASKED FOR EVERY CITATION IN ONE CHAPTER
A LANGUAGE MODEL CANNOT REPORT WHAT IT NEVER LOOKED AT
WHAT IT RETURNED
“THAT IS THE LAST ONE”
SAID 27 TIMES, EACH TIME WITH CERTAINTY
WHAT THE CHAPTER HELD
HOLLOW = NEVER REACHED. NOTHING JOINS THEM,
BECAUSE NOTHING EVER LOOKED.
What it returned, beside what the chapter actually held. The hollow entries are joined to nothing — nothing ever looked at them, and nothing could say so.
In November 2025 I asked a general-purpose AI assistant to find the missing
citations in a single chapter of a thesis. A short chapter. Mechanical work — the kind of
task that looks beneath a language model rather than beyond it. It missed twenty-seven.
The misses were not the interesting part. Any tool has a recall rate, and a recall rate
can be measured, published and improved. What made this different was the answer I got
every time I pointed at something it had skipped. I would ask what about this one? ,
and it would apologise, add the citation, and tell me — with no hedging whatsoever — that
this was now the last one.
Twenty-seven times. Each correction was delivered with the same complete confidence as
the report that had preceded it.
I wrote it up the same day, months before there was any product to sell — as a reply to an
open request from that assistant's own makers for examples of it getting things wrong:
The specific assistant matters less than it looks. That model has had nine months of
releases since, and it would miss fewer today — every model would. What none of them has
acquired in those nine months is the ability to tell you whether they missed any, because
that capability is not a matter of scale. It is a matter of architecture, and it is the one
the task actually requires.
That afternoon is the reason CitationLab exists, and it is the reason the product is
built in an order that looks backwards to anyone who has been told that AI is the fast way
to do this. Today, CitationLab is live.
The failure has a name, and it is not hallucination
Invented references are the failure everybody knows about. A model produces a citation
to a paper that was never written, complete with a plausible DOI and a real-sounding
journal. It is embarrassing, it is well documented, and — crucially — it is
self-revealing . You click the identifier, it resolves to nothing, and you
delete the reference. We have written about that failure
in its own right , because it deserves the
attention it gets.
The failure in front of me was the mirror image, and it gets almost none. The model did
not add anything false. It omitted, silently, and then reported completion — because
reporting completion is what the end of a list looks like in everything it has ever been
trained on.
A language model has no inventory of your document. It has no counter that decrements as
it works through a chapter, no set of things-not-yet-examined that must be empty before it
can stop. It produces the next plausible token, and eventually the next plausible token is
the one that ends a list. Nothing in that process distinguishes having finished from having
stopped.
Ask a model what it missed and you are asking it to describe a place it never
visited. It will answer, because answering is what it does. The answer carries no
information.
This is not a complaint about one vendor, and it is not something a larger model fixes.
A better model misses fewer. It still cannot tell you whether it missed any — and for a
thesis, that second property is the one that matters. A tool that finds eighty per cent of
your problems and tells you it found all of them is worse than no tool, because it converts
an open question into a closed one.
Read the amber lines on their own. Every one of them is false, every one of them is
stated without qualification, and there is no version of that conversation in which the
model volunteers the twenty-eighth. The only thing driving the process forward was that I
already knew what was in my own chapter — which is precisely the knowledge someone reaching
for this tool does not have.
So the rules run first
CitationLab inverts the ordering that most tools use, and the inversion is the whole
design rather than an implementation detail.
Your document is parsed structurally. It is not read by a model, summarised by a model,
or described to a model. Every in-text citation and every bibliography entry is extracted by
rule, and each one carries a character offset, so that every finding points at a real
location in your file rather than at a paraphrase of it. When the system says a citation is
unmatched, you can go and look at it.
Matching happens next, and matching is arithmetic rather than opinion. Either a citation
has a corresponding bibliography entry or it does not. Either a reference entry is cited
somewhere in the body or it is an orphan . These are
countable facts about a document, and a count that comes from enumeration is a count you can
trust — it has a denominator.
The denominator is the part worth dwelling on, because it is the thing the November
conversation never had. When a check reports that eleven citations are unmatched, the useful
question is not whether eleven is a big number. It is eleven out of what — and
whether the tool can show you the other side of that fraction. An enumeration can. It knows
how many citations were in the chapter because it counted them, one at a time, at known
positions in your file. A generated list cannot, because there was never a moment at which
anything was counted.
Verification runs against seven independent engines, because no single index is complete:
Crossref, PubMed, Semantic Scholar, OpenAlex, Google Scholar, Google Books, and general web
discovery for grey literature. Crossref does not hold most books. PubMed declines anything
outside biomedicine by design. A government report or an NGO working paper lives in neither.
Reconciling seven partial answers into one verdict is ordinary engineering — we have written
up how that reconciliation resolves disagreements in
seven engines, one verdict . It is slower than
asking a model, and it is the reason the answer means something when it arrives.
None of this is clever. That is rather the point. Every step in it is the kind of thing
you could check by hand if you had a month, which is exactly what makes it checkable at all
— and it is why we can say how the system behaves across
thousands of test documents rather
than reporting how it felt on the last one.
Cross-check runs on your own document, not a sample file — upload the chapter you are
actually uncertain about and see what pairs up.
See what a check costs
What this looks like from the other side of the desk
It is worth being concrete about why any of this matters, because "your references have
errors" is easy to file under things to worry about later.
An examiner reading your thesis does not audit your bibliography. They read, and
occasionally something stops them — a claim they want to follow up, a paper they know, a
name that looks wrong. Then they turn to the reference list. If the entry is not there, the
question is no longer about the reference. It is about whether the rest of the list is like
this, and that question cannot be answered by looking at one entry. It can only be answered
by checking, which is not what they came to do, and it changes the register of the whole
viva.
This is the asymmetry that makes REFerence [ IN ] tegrity worth an afternoon of your time. A
clean list earns you nothing; nobody praises a thesis for having working DOIs. A broken one
costs you disproportionately, because a single visible failure is evidence about a
population the reader cannot see. The work is not about the eleven citations. It is about
not handing someone a reason to start counting.
And the failures that produce this are almost never dramatic. They are the
same author spelled two ways across chapters
written eighteen months apart, a reference manager that dropped four entries during a
library migration, a chapter reordering that left three numbered citations pointing at the
wrong rows. Each is individually trivial. Collectively they are the reason the check exists.
Four checks, one document
The launch version does four things, and they are deliberately separable — you run what
your document needs rather than a single opaque pass that reports a grade.
Cross-check reconciles every in-text citation against every
bibliography entry, in both directions. It reports citations with no entry, and entries
nothing cites. The first ten pairs are
free , which is enough to tell you whether there is a problem at all.
Find missing takes what your bibliography lacks and searches the seven
engines for the real work, returning candidates with identifiers for you to accept or
reject. Nothing is auto-selected.
Update old finds preprints that have since been published, superseded
editions and dead links, and proposes replacements
without breaking your numbering .
Verify resolves identifiers and confirms each one points at the paper
you actually cited — the check that catches a DOI that is real, resolvable, and attached
to the wrong article.
Each produces a report you can hand to a supervisor, and each finding carries the offset
that lets you go straight to the line in your own file. There is no step at which the
system knows something about your document that it will not show you.
AI comes last, and it never decides
There is a real role for a model in this work, and it arrives at the end. Once the rules
have settled everything countable, what remains is a residue of genuine ambiguity — two
papers by the same team in the same year with near-identical titles, an author name that
survived a transliteration badly, a conference paper that also exists as a journal article.
Judging those is exactly the kind of reading a model is good at.
So that is what it is asked to do, and only that. It reads a candidate and offers a
reading. It does not choose. It does not write into your document.
Every correction requires your
click — not as a safety-theatre step, but because the alternative is a document that
changed in ways nobody can enumerate afterwards.
The ordering is the argument. If a model decides what to look for, its blind spots become
your blind spots and stay invisible, because the same faculty that produced the gap also
writes the report about the gap. If it is consulted last, on a residue that was defined
without it, its blind spots are just a slower review. Same model, same limitations,
completely different consequences — determined entirely by where in the pipeline it sits.
What the Score is, and what it is not
The Ref[ In ]
Score summarises how sound your referencing is across those checks. It is worth being
precise about what that sentence means, because the surrounding vocabulary in this field has
been poisoned by tools that do something else entirely.
The Score is a referencing-quality measure. It is not a similarity score, and it
is not a finding of misconduct. That line is printed on the cover of every report,
and it is there because the distinction is real. A thesis with broken references is, almost
without exception, a thesis assembled over several years with ordinary tools by someone who
was also doing the research. Reference managers drop entries during an import. Numbering
shifts when a supervisor asks for chapters three and four to swap. A preprint acquires a
journal DOI two years after you cited it, and the version you cited stops being the version
of record. None of that is dishonesty. All of it is visible to an examiner who checks.
Start with the chapter you are worried about
Not a demo file. The one you have been avoiding. Cross-check will tell you within a
couple of minutes whether there is a problem in it, and a real proportion of the time the
answer is that there is not — which is worth knowing too.
That is the answer the assistant could never give me on that afternoon in November. Not
because it was a bad model, but because nothing in it could distinguish a chapter it had
read completely from a chapter it had stopped reading. Twenty-seven times it told me it was
done. It was right on the twenty-eighth, and it had no way of knowing that either.
CitationLab is live as of today. Run a cross-check on your own thesis and find out what
is actually in it.
Start a check