REFerence [IN]tegrity for universities — the check that stopped running
CitationLab Team · August 2026 · 11 min read
Two gates at submission. One is automated and universal; the other was manual, and quietly closed.
Every thesis your institution receives is screened for plagiarism. It happens
automatically, it happens to everyone, and nobody has to remember to ask for it.
Who checks the references?
Not as a formatting matter — whether the entries are in the right style — but the
substantive question: does every citation in the text correspond to a real entry in the
bibliography, and does every entry correspond to something the candidate actually cited?
In almost every institution we have looked at, the honest answer is that nobody does. It
is in no workflow, on no checklist, and assigned to no role.
It used to happen
This is not a check that was never invented. Supervisors did it, in the margins of a
final read. Departmental administrators did it for the formatting. Subject librarians
did it when a student brought them a list. It was slow and it was partial, but it
happened often enough that a broken bibliography usually got caught by someone.
What ended it was arithmetic.
Two thousand hours is not a resourcing problem to be solved with more diligence. It
is a task nobody can do at that volume, and so — without any decision being taken, at
any meeting, by anyone — it stopped being done. The supervisors who used to catch these
things now supervise more candidates. The librarians who used to sit with a list now
support a research office. The check did not get abolished. It got squeezed out.
No institution decided to stop checking references. The volume decided, and
everybody adjusted around it.
Four things happened, and none of them was a decision
Volume is the headline cause but not the whole one. Three other changes arrived in the
same period, and each of them made the check easier to stop doing.
Supervision load rose faster than supervision time. A supervisor with
three candidates can afford a final read that includes the bibliography. The same
supervisor with eleven cannot, and the part that goes first is always the part that is
mechanical, unrewarded and invisible if skipped. Nobody announces that they have stopped
checking references; they simply run out of the evening in which they used to.
The subject librarian moved. The role that used to sit closest to
this — someone who knew the literature of a discipline and would work through a list with
a student — has largely been restructured into research support, data management and
open-access compliance. Those are real jobs and the restructure was rational. But the
person who once caught a broken reference by recognising that the journal did not publish
under that name in that year is now three teams away from the thesis.
Reference managers made lists look correct. This one is subtle and it
matters most. Before automated formatting, a sloppy bibliography looked sloppy — uneven
punctuation, inconsistent capitalisation, entries in the wrong order. Visible untidiness
was a proxy for unchecked work, and examiners used it as one. A reference manager removes
the untidiness while leaving every substantive defect untouched. The list is now uniform,
beautifully punctuated, and just as likely to point at nothing.
And the plagiarism gate absorbed the attention. When an institution
installs one automated integrity check at submission, it is reasonable — and wrong — to
feel that integrity is now covered. Plagiarism screening answers a different question
from REFerence [IN]tegrity: it asks whether text was taken from elsewhere, not whether the
sources cited actually exist and match their citations. A thesis can pass a similarity
check cleanly while a tenth of its citations point at nothing, because those are
orthogonal properties. The gate is doing its job. It is simply not doing this one.
Nobody knows how big the gap is
Because the check is not run, there is no figure. Institutions cannot tell you what
proportion of their submitted theses carry a citation that resolves to nothing, and
neither could we until we started measuring it.
Preliminary scanning of public submissions — hundreds of theses per institution, from
the last five years — puts the proportion carrying at least one unresolved citation
closer to two in five than to one in twenty. We publish the full
methodology, the denominators and the per-institution distribution in the first
Ref[In] Index
in August 2027, and we would rather you treated the shape as provisional until then.
Whatever the final number, most of what sits inside it is bookkeeping. A gap between
a citation and a reference list is usually the residue of a long project done by a busy
person: a source added in a late revision, a surname transliterated two ways across four
years, an edition superseded while a chapter sat waiting for comments. Treating that as a
discipline matter would be both wrong and destructive, and it is why the framing an
institution adopts matters as much as the measurement.
Inside the same number, though, is a smaller category an institution should want
counted separately: citations whose source cannot be found to exist. That is not
bookkeeping. Presenting a fabricated source in support of a claim is academic misconduct
under every code we have read, and it is the one thing here a check can demonstrate
rather than infer. How often it appears, and what is done about it, are decisions for the
institution — but not knowing the figure is a choice too.
AI moved the problem in both directions at once
The timing is the part institutions tend to underestimate. Two things happened in the
same eighteen months, and they push opposite ways.
The first is that generative tools now produce references. Not maliciously — a
candidate asks for sources on a topic and receives a plausible, correctly formatted list
in which some entries do not exist. We have written about
what a candidate should do when that has already
happened. The volume of references entering theses without ever having been seen in a
real record has gone up sharply.
The second is that the obvious remedy does not work. Asking a general-purpose model to
verify a reference list produces confident, fluent, unreliable answers — it will as
happily invent a confirmation as invent a citation, because both are the same operation
to it. An institution that adopts AI checking without asking what the model is resolving
against has bought a second layer of the first problem.
The distinction is the whole design of what we build: deterministic resolution
against real records first, with judgement applied only to what genuinely remains
ambiguous, and every proposed change shown to a human before anything is written. It is
slower to build and it is the only version that an institution can defend.
See the measurement on your own submissions. Universities can
request a report across their public theses — the same arithmetic, with the per-thesis
evidence behind every line, and nothing published about anyone.
Talk to us about an institutional report
Where this lands inside a university
The question we are asked next is always operational: whose job is this? The answer is
that it belongs to units you already fund, doing work they already do — by hand, at a
fraction of the coverage they would choose.
Academic writing units and writing centres. What happens now: a
candidate books a consultation about referencing, and the tutor spends most of the hour
doing by eye what the student could not — reading down the bibliography, spot-checking
against the text, finding three or four problems before the time runs out. The tutor is
a skilled reader being used as a comparison engine, and the coverage is whatever fits in
sixty minutes. What changes: both people arrive with the full list already in front of
them, and the hour goes on the judgement calls — which orphan to cut, whether an old
source still earns its place, how to phrase a secondary citation. That is what a writing
tutor is actually for, and it is the part no tool does.
Referencing labs and citation clinics. Where these exist they are
usually the most under-equipped good idea in the building: a drop-in session, staffed by
people who know referencing well, with no instrument behind them. The clinic can advise;
it cannot diagnose at speed. What changes is the shape of the session — a candidate
arrives with a document and leaves with a specific worklist rather than general advice.
It also changes who can staff it, because reading a report and discussing it is a
different skill from holding six hundred citations in your head, and the second one is
rare.
Referencing training programmes and bootcamps. Referencing is nearly
always taught from rules and invented examples, which is why it does not stick: the
examples are clean and the participant's own list is not. Running a live check on a
participant's real bibliography, in the room, changes the lesson from a style guide into
a diagnosis. In our experience the moment that lands is not the missing entries — it is
the orphan nobody expected, or the two spellings of one surname, because those are the
errors a careful person is certain they do not make.
Reference-management conferences and workshops. These already command
budget, attendance and a slot in the academic calendar; what they generally lack is
something a participant leaves holding. A session that ends with each attendee having a
report on their own work is more
persuasive than any slide about why referencing matters, and it converts a talk into a
service the institution can point at afterwards.
Graduate schools and doctoral training centres. This is the owner if
you want coverage rather than opportunity. Everything above is opt-in, which means it
reaches the candidates who already worry about referencing — reliably the ones who need
it least. A graduate school can make the check a standard step in the doctoral timeline,
applied to everyone, at a point where a correction is cheap: after the literature review
settles, again before submission. That is the version that changes the institutional
number rather than helping individuals inside it.
Library research-support teams. Already the referral endpoint when a
researcher arrives with a list that will not behave, already trusted to be
non-judgemental about it, and already holding the bibliographic expertise to interpret an
ambiguous result. That trust is not incidental — it is exactly what a quality check needs
and what a compliance office cannot provide. A candidate will bring a library a problem
they would not bring anyone else, which is the whole reason the library is the right
home for the awkward findings.
The point of listing them separately is that each answers a different question, and an
institution that buys this for one of them is usually solving for the wrong one. If the
goal is to help the candidates who ask, the writing centre is the right home. If the goal
is to change what the institution submits, only the graduate school can do it, because
only it reaches the people who never asked.
The four objections, answered
"Students should learn to do this themselves." They should, and this
teaches it better than the current arrangement does. The status quo is not students
learning to check references — it is students not checking them and nobody finding out. A
report that shows a candidate their own nine missing entries teaches the lesson in a way
no handbook chapter has ever managed, because it is about their work rather than an
example. The manual skill worth preserving is judgement — which source to keep, whether
an old paper still carries the claim — and that is precisely the part a check hands back
to the person rather than taking away.
"This creates work for supervisors." It removes work, and moves what
remains earlier. The current cost is not zero: it is paid in the final read, in
corrections after examination, and in the awkward conversation when an examiner finds
something a supervisor did not. What a report changes is that the finding arrives as a
worklist owned by the candidate, at a point where fixing it is a lookup rather than a
revision.
"What about data protection?" A fair question and one with a specific
answer: the check needs the citations and the reference lines, not the argument, the
results or the data — we set out
exactly what leaves a document, including
the part that is not reassuring. For institutional measurement the constraint is tighter
still, because it runs only on theses the university has already published openly.
"What if our number is bad?" It will be, on the first measurement,
and so will everyone else's — that is what a check nobody has been running produces. The
question worth asking is not whether the first number is flattering but whether you would
rather know it before or after someone else measures it. A figure you discovered
yourself, with a plan attached, is a different conversation from a figure that arrives in
a sector report.
What it costs to start
Less than the conversation usually assumes, because none of the above is a new
activity. Every one of those units is currently funded to support referencing and is
doing it with the instruments available — which is to say, by reading. The decision in
front of an institution is not whether to start spending on this. It is whether the
spending already committed should continue to buy coverage of a few dozen students a
term.
The practical starting point is smaller still: measure first. Ask what the figure is
for your own submissions before deciding what, if anything, to do about it. A number
makes the conversation concrete, and it costs nothing to be curious about.
Start with the measurement. We will run your public submissions and
show you where your institution actually stands — with the evidence behind every line,
and nothing about any student, thesis or supervisor published anywhere.
See plans