AI & Integrity · Launch note

It missed 27 citations, then told me it had found the last one

CitationLab Team · August 2026 · 7 min read
ASKED FOR EVERY CITATION IN ONE CHAPTER A LANGUAGE MODEL CANNOT REPORT WHAT IT NEVER LOOKED AT WHAT IT RETURNED “THAT IS THE LAST ONE” SAID 27 TIMES, EACH TIME WITH CERTAINTY WHAT THE CHAPTER HELD HOLLOW = NEVER REACHED. NOTHING JOINS THEM, BECAUSE NOTHING EVER LOOKED.
What it returned, beside what the chapter actually held. The hollow entries are joined to nothing — nothing ever looked at them, and nothing could say so.

In November 2025 I asked a general-purpose AI assistant to find the missing citations in a single chapter of a thesis. A short chapter. Mechanical work — the kind of task that looks beneath a language model rather than beyond it. It missed twenty-seven.

The misses were not the interesting part. Any tool has a recall rate, and a recall rate can be measured, published and improved. What made this different was the answer I got every time I pointed at something it had skipped. I would ask what about this one?, and it would apologise, add the citation, and tell me — with no hedging whatsoever — that this was now the last one.

Twenty-seven times. Each correction was delivered with the same complete confidence as the report that had preceded it.

I wrote it up the same day, months before there was any product to sell — as a reply to an open request from that assistant's own makers for examples of it getting things wrong:

Posted 22 November 2025 · see the original

“I asked @Grok to scan for citations to update missing references in a thesis chapter. Purely mechanic work in a rather brief chapter. It missed TWENTY SEVEN (27) citations. I ask ‘what about this one’, and it kept CONFIDENTLY saying this is the last ONE…27 TIMES!”

The specific assistant matters less than it looks. That model has had nine months of releases since, and it would miss fewer today — every model would. What none of them has acquired in those nine months is the ability to tell you whether they missed any, because that capability is not a matter of scale. It is a matter of architecture, and it is the one the task actually requires.

That afternoon is the reason CitationLab exists, and it is the reason the product is built in an order that looks backwards to anyone who has been told that AI is the fast way to do this. Today, CitationLab is live.

The failure has a name, and it is not hallucination

Invented references are the failure everybody knows about. A model produces a citation to a paper that was never written, complete with a plausible DOI and a real-sounding journal. It is embarrassing, it is well documented, and — crucially — it is self-revealing. You click the identifier, it resolves to nothing, and you delete the reference. We have written about that failure in its own right, because it deserves the attention it gets.

The failure in front of me was the mirror image, and it gets almost none. The model did not add anything false. It omitted, silently, and then reported completion — because reporting completion is what the end of a list looks like in everything it has ever been trained on.

A language model has no inventory of your document. It has no counter that decrements as it works through a chapter, no set of things-not-yet-examined that must be empty before it can stop. It produces the next plausible token, and eventually the next plausible token is the one that ends a list. Nothing in that process distinguishes having finished from having stopped.

Ask a model what it missed and you are asking it to describe a place it never visited. It will answer, because answering is what it does. The answer carries no information.

This is not a complaint about one vendor, and it is not something a larger model fixes. A better model misses fewer. It still cannot tell you whether it missed any — and for a thesis, that second property is the one that matters. A tool that finds eighty per cent of your problems and tells you it found all of them is worse than no tool, because it converts an open question into a closed one.

Figure 1 · the shape of the conversation, reconstructed with invented names
you  list every in-text citation in this chapter

ai   Okoro (2019), Baptiste & Lindqvist (2021), Ferreira (2018),
     Nwachukwu et al. (2020) … fourteen in total.
     That is the complete set for this chapter.

you  what about Achterberg et al. (2017)?

ai   You are right, my apologies — Achterberg et al. (2017) as well.
     That is now the last one.

you  and Sandoval (2022)?

ai   Good catch. Sandoval (2022) too.
     That completes the list.

you  and Ferreira (2018) is cited twice — the second one?

ai   Correct, the second occurrence as well.
     The list is now complete.

     …twenty-three more times.

Read the amber lines on their own. Every one of them is false, every one of them is stated without qualification, and there is no version of that conversation in which the model volunteers the twenty-eighth. The only thing driving the process forward was that I already knew what was in my own chapter — which is precisely the knowledge someone reaching for this tool does not have.

So the rules run first

CitationLab inverts the ordering that most tools use, and the inversion is the whole design rather than an implementation detail.

Your document is parsed structurally. It is not read by a model, summarised by a model, or described to a model. Every in-text citation and every bibliography entry is extracted by rule, and each one carries a character offset, so that every finding points at a real location in your file rather than at a paraphrase of it. When the system says a citation is unmatched, you can go and look at it.

Matching happens next, and matching is arithmetic rather than opinion. Either a citation has a corresponding bibliography entry or it does not. Either a reference entry is cited somewhere in the body or it is an orphan. These are countable facts about a document, and a count that comes from enumeration is a count you can trust — it has a denominator.

The denominator is the part worth dwelling on, because it is the thing the November conversation never had. When a check reports that eleven citations are unmatched, the useful question is not whether eleven is a big number. It is eleven out of what — and whether the tool can show you the other side of that fraction. An enumeration can. It knows how many citations were in the chapter because it counted them, one at a time, at known positions in your file. A generated list cannot, because there was never a moment at which anything was counted.

Verification runs against seven independent engines, because no single index is complete: Crossref, PubMed, Semantic Scholar, OpenAlex, Google Scholar, Google Books, and general web discovery for grey literature. Crossref does not hold most books. PubMed declines anything outside biomedicine by design. A government report or an NGO working paper lives in neither. Reconciling seven partial answers into one verdict is ordinary engineering — we have written up how that reconciliation resolves disagreements in seven engines, one verdict. It is slower than asking a model, and it is the reason the answer means something when it arrives.

None of this is clever. That is rather the point. Every step in it is the kind of thing you could check by hand if you had a month, which is exactly what makes it checkable at all — and it is why we can say how the system behaves across thousands of test documents rather than reporting how it felt on the last one.

Cross-check runs on your own document, not a sample file — upload the chapter you are actually uncertain about and see what pairs up.

See what a check costs

What this looks like from the other side of the desk

It is worth being concrete about why any of this matters, because "your references have errors" is easy to file under things to worry about later.

An examiner reading your thesis does not audit your bibliography. They read, and occasionally something stops them — a claim they want to follow up, a paper they know, a name that looks wrong. Then they turn to the reference list. If the entry is not there, the question is no longer about the reference. It is about whether the rest of the list is like this, and that question cannot be answered by looking at one entry. It can only be answered by checking, which is not what they came to do, and it changes the register of the whole viva.

This is the asymmetry that makes REFerence [IN]tegrity worth an afternoon of your time. A clean list earns you nothing; nobody praises a thesis for having working DOIs. A broken one costs you disproportionately, because a single visible failure is evidence about a population the reader cannot see. The work is not about the eleven citations. It is about not handing someone a reason to start counting.

And the failures that produce this are almost never dramatic. They are the same author spelled two ways across chapters written eighteen months apart, a reference manager that dropped four entries during a library migration, a chapter reordering that left three numbered citations pointing at the wrong rows. Each is individually trivial. Collectively they are the reason the check exists.

Four checks, one document

The launch version does four things, and they are deliberately separable — you run what your document needs rather than a single opaque pass that reports a grade.

Each produces a report you can hand to a supervisor, and each finding carries the offset that lets you go straight to the line in your own file. There is no step at which the system knows something about your document that it will not show you.

AI comes last, and it never decides

There is a real role for a model in this work, and it arrives at the end. Once the rules have settled everything countable, what remains is a residue of genuine ambiguity — two papers by the same team in the same year with near-identical titles, an author name that survived a transliteration badly, a conference paper that also exists as a journal article. Judging those is exactly the kind of reading a model is good at.

So that is what it is asked to do, and only that. It reads a candidate and offers a reading. It does not choose. It does not write into your document. Every correction requires your click — not as a safety-theatre step, but because the alternative is a document that changed in ways nobody can enumerate afterwards.

The ordering is the argument. If a model decides what to look for, its blind spots become your blind spots and stay invisible, because the same faculty that produced the gap also writes the report about the gap. If it is consulted last, on a residue that was defined without it, its blind spots are just a slower review. Same model, same limitations, completely different consequences — determined entirely by where in the pipeline it sits.

What the Score is, and what it is not

The Ref[In] Score summarises how sound your referencing is across those checks. It is worth being precise about what that sentence means, because the surrounding vocabulary in this field has been poisoned by tools that do something else entirely.

The Score is a referencing-quality measure. It is not a similarity score, and it is not a finding of misconduct. That line is printed on the cover of every report, and it is there because the distinction is real. A thesis with broken references is, almost without exception, a thesis assembled over several years with ordinary tools by someone who was also doing the research. Reference managers drop entries during an import. Numbering shifts when a supervisor asks for chapters three and four to swap. A preprint acquires a journal DOI two years after you cited it, and the version you cited stops being the version of record. None of that is dishonesty. All of it is visible to an examiner who checks.

Start with the chapter you are worried about

Not a demo file. The one you have been avoiding. Cross-check will tell you within a couple of minutes whether there is a problem in it, and a real proportion of the time the answer is that there is not — which is worth knowing too.

That is the answer the assistant could never give me on that afternoon in November. Not because it was a bad model, but because nothing in it could distinguish a chapter it had read completely from a chapter it had stopped reading. Twenty-seven times it told me it was done. It was right on the twenty-eighth, and it had no way of knowing that either.

CitationLab is live as of today. Run a cross-check on your own thesis and find out what is actually in it.

Start a check
Filed under: AI & Integrity missing-citations deterministic-checking academic-engines
Share: Post on X Share Email

Keep reading

AI & Integrity

Can AI invent fake references? Ours can't — by construction

The operation that produces fabricated references — generate a plausible source — is not implemented. What the model is actually asked, and where references come from instead.

Read the post →
Updating References

How we find the newer version: seven engines, one verdict

Crossref, PubMed, Semantic Scholar, OpenAlex and the scraped tiers, asked in a fixed order. Why free-first is a metadata-quality rule.

Read the post →
AI & Integrity

Deterministic first, AI second — and why the order is the whole design

Why CitationLab settles what it can with rules before any model is consulted, and what that ordering buys you in a reference check you can actually verify.

Read the post →
The blog

Browse every post

Case files, guides and engineering notes — everything we learn checking real theses.

All posts →