For Institutions · The case

REFerence [IN]tegrity for universities — the check that stopped running

CitationLab Team · August 2026 · 11 min read
EVERY THESIS PASSES TWO GATES SUBMITTED GATE 1 — PLAGIARISM automated every submission EXAMINED GATE 2 — REFERENCES nobody's job in no workflow BY HAND ~5 hrs per thesis ONE SEASON 2,000 hrs 400 submissions No institution decided to stop checking references. The volume decided, and everybody adjusted around it. Ref[In]
Two gates at submission. One is automated and universal; the other was manual, and quietly closed.

Every thesis your institution receives is screened for plagiarism. It happens automatically, it happens to everyone, and nobody has to remember to ask for it.

Who checks the references?

Not as a formatting matter — whether the entries are in the right style — but the substantive question: does every citation in the text correspond to a real entry in the bibliography, and does every entry correspond to something the candidate actually cited? In almost every institution we have looked at, the honest answer is that nobody does. It is in no workflow, on no checklist, and assigned to no role.

It used to happen

This is not a check that was never invented. Supervisors did it, in the margins of a final read. Departmental administrators did it for the formatting. Subject librarians did it when a student brought them a list. It was slow and it was partial, but it happened often enough that a broken bibliography usually got caught by someone.

What ended it was arithmetic.

The check, expressed as work — illustrative, but the scale is real
One doctoral thesis          ~600 citations
                             ~350 reference entries

Forward pass                 does each citation reach an entry?
Backward pass                is each entry actually cited?

At 20 seconds per decision   about 5 hours, per thesis,
                             for one person, done carefully

A mid-sized graduate school   400 submissions in a season
                             = 2,000 hours of somebody's attention

Two thousand hours is not a resourcing problem to be solved with more diligence. It is a task nobody can do at that volume, and so — without any decision being taken, at any meeting, by anyone — it stopped being done. The supervisors who used to catch these things now supervise more candidates. The librarians who used to sit with a list now support a research office. The check did not get abolished. It got squeezed out.

No institution decided to stop checking references. The volume decided, and everybody adjusted around it.

Four things happened, and none of them was a decision

Volume is the headline cause but not the whole one. Three other changes arrived in the same period, and each of them made the check easier to stop doing.

Supervision load rose faster than supervision time. A supervisor with three candidates can afford a final read that includes the bibliography. The same supervisor with eleven cannot, and the part that goes first is always the part that is mechanical, unrewarded and invisible if skipped. Nobody announces that they have stopped checking references; they simply run out of the evening in which they used to.

The subject librarian moved. The role that used to sit closest to this — someone who knew the literature of a discipline and would work through a list with a student — has largely been restructured into research support, data management and open-access compliance. Those are real jobs and the restructure was rational. But the person who once caught a broken reference by recognising that the journal did not publish under that name in that year is now three teams away from the thesis.

Reference managers made lists look correct. This one is subtle and it matters most. Before automated formatting, a sloppy bibliography looked sloppy — uneven punctuation, inconsistent capitalisation, entries in the wrong order. Visible untidiness was a proxy for unchecked work, and examiners used it as one. A reference manager removes the untidiness while leaving every substantive defect untouched. The list is now uniform, beautifully punctuated, and just as likely to point at nothing.

What formatting stopped telling you
Before          a list assembled by hand
  uneven punctuation, mixed capitalisation
  → looked unchecked, and usually was
  → an examiner's eye caught the signal

Now             a list assembled by a manager
  uniform, correctly styled, alphabetised
  → looks checked
  → the signal is gone, the defects are not

The proxy broke. Nothing replaced it.

And the plagiarism gate absorbed the attention. When an institution installs one automated integrity check at submission, it is reasonable — and wrong — to feel that integrity is now covered. Plagiarism screening answers a different question from REFerence [IN]tegrity: it asks whether text was taken from elsewhere, not whether the sources cited actually exist and match their citations. A thesis can pass a similarity check cleanly while a tenth of its citations point at nothing, because those are orthogonal properties. The gate is doing its job. It is simply not doing this one.

Nobody knows how big the gap is

Because the check is not run, there is no figure. Institutions cannot tell you what proportion of their submitted theses carry a citation that resolves to nothing, and neither could we until we started measuring it.

Preliminary scanning of public submissions — hundreds of theses per institution, from the last five years — puts the proportion carrying at least one unresolved citation closer to two in five than to one in twenty. We publish the full methodology, the denominators and the per-institution distribution in the first Ref[In] Index in August 2027, and we would rather you treated the shape as provisional until then.

Whatever the final number, most of what sits inside it is bookkeeping. A gap between a citation and a reference list is usually the residue of a long project done by a busy person: a source added in a late revision, a surname transliterated two ways across four years, an edition superseded while a chapter sat waiting for comments. Treating that as a discipline matter would be both wrong and destructive, and it is why the framing an institution adopts matters as much as the measurement.

Inside the same number, though, is a smaller category an institution should want counted separately: citations whose source cannot be found to exist. That is not bookkeeping. Presenting a fabricated source in support of a claim is academic misconduct under every code we have read, and it is the one thing here a check can demonstrate rather than infer. How often it appears, and what is done about it, are decisions for the institution — but not knowing the figure is a choice too.

AI moved the problem in both directions at once

The timing is the part institutions tend to underestimate. Two things happened in the same eighteen months, and they push opposite ways.

The first is that generative tools now produce references. Not maliciously — a candidate asks for sources on a topic and receives a plausible, correctly formatted list in which some entries do not exist. We have written about what a candidate should do when that has already happened. The volume of references entering theses without ever having been seen in a real record has gone up sharply.

The second is that the obvious remedy does not work. Asking a general-purpose model to verify a reference list produces confident, fluent, unreliable answers — it will as happily invent a confirmation as invent a citation, because both are the same operation to it. An institution that adopts AI checking without asking what the model is resolving against has bought a second layer of the first problem.

Two questions that sound alike and are not
"Does this reference look right?"      a language question
                                       answerable from plausibility alone
                                       — and therefore answerable wrongly

"Does this reference exist, and is
 it the one the citation points at?"   a lookup question
                                       answerable only against a real
                                       record, or not at all

The distinction is the whole design of what we build: deterministic resolution against real records first, with judgement applied only to what genuinely remains ambiguous, and every proposed change shown to a human before anything is written. It is slower to build and it is the only version that an institution can defend.

See the measurement on your own submissions. Universities can request a report across their public theses — the same arithmetic, with the per-thesis evidence behind every line, and nothing published about anyone.

Talk to us about an institutional report

Where this lands inside a university

The question we are asked next is always operational: whose job is this? The answer is that it belongs to units you already fund, doing work they already do — by hand, at a fraction of the coverage they would choose.

Academic writing units and writing centres. What happens now: a candidate books a consultation about referencing, and the tutor spends most of the hour doing by eye what the student could not — reading down the bibliography, spot-checking against the text, finding three or four problems before the time runs out. The tutor is a skilled reader being used as a comparison engine, and the coverage is whatever fits in sixty minutes. What changes: both people arrive with the full list already in front of them, and the hour goes on the judgement calls — which orphan to cut, whether an old source still earns its place, how to phrase a secondary citation. That is what a writing tutor is actually for, and it is the part no tool does.

Referencing labs and citation clinics. Where these exist they are usually the most under-equipped good idea in the building: a drop-in session, staffed by people who know referencing well, with no instrument behind them. The clinic can advise; it cannot diagnose at speed. What changes is the shape of the session — a candidate arrives with a document and leaves with a specific worklist rather than general advice. It also changes who can staff it, because reading a report and discussing it is a different skill from holding six hundred citations in your head, and the second one is rare.

Referencing training programmes and bootcamps. Referencing is nearly always taught from rules and invented examples, which is why it does not stick: the examples are clean and the participant's own list is not. Running a live check on a participant's real bibliography, in the room, changes the lesson from a style guide into a diagnosis. In our experience the moment that lands is not the missing entries — it is the orphan nobody expected, or the two spellings of one surname, because those are the errors a careful person is certain they do not make.

Reference-management conferences and workshops. These already command budget, attendance and a slot in the academic calendar; what they generally lack is something a participant leaves holding. A session that ends with each attendee having a report on their own work is more persuasive than any slide about why referencing matters, and it converts a talk into a service the institution can point at afterwards.

Graduate schools and doctoral training centres. This is the owner if you want coverage rather than opportunity. Everything above is opt-in, which means it reaches the candidates who already worry about referencing — reliably the ones who need it least. A graduate school can make the check a standard step in the doctoral timeline, applied to everyone, at a point where a correction is cheap: after the literature review settles, again before submission. That is the version that changes the institutional number rather than helping individuals inside it.

Library research-support teams. Already the referral endpoint when a researcher arrives with a list that will not behave, already trusted to be non-judgemental about it, and already holding the bibliographic expertise to interpret an ambiguous result. That trust is not incidental — it is exactly what a quality check needs and what a compliance office cannot provide. A candidate will bring a library a problem they would not bring anyone else, which is the whole reason the library is the right home for the awkward findings.

The same instrument, six different jobs
Writing centre     one-to-one · depth over coverage
                   → the hour goes on judgement, not comparison

Citation clinic    drop-in · speed over depth
                   → a worklist instead of general advice

Training           cohort · teaching over fixing
                   → their own list as the worked example

Conference         event · persuasion over process
                   → an artefact each attendee takes away

Graduate school    everyone · coverage over depth
                   → the number moves, not just individuals

Library            referral · trust over throughput
                   → the awkward findings land somewhere safe

The point of listing them separately is that each answers a different question, and an institution that buys this for one of them is usually solving for the wrong one. If the goal is to help the candidates who ask, the writing centre is the right home. If the goal is to change what the institution submits, only the graduate school can do it, because only it reaches the people who never asked.

The four objections, answered

"Students should learn to do this themselves." They should, and this teaches it better than the current arrangement does. The status quo is not students learning to check references — it is students not checking them and nobody finding out. A report that shows a candidate their own nine missing entries teaches the lesson in a way no handbook chapter has ever managed, because it is about their work rather than an example. The manual skill worth preserving is judgement — which source to keep, whether an old paper still carries the claim — and that is precisely the part a check hands back to the person rather than taking away.

"This creates work for supervisors." It removes work, and moves what remains earlier. The current cost is not zero: it is paid in the final read, in corrections after examination, and in the awkward conversation when an examiner finds something a supervisor did not. What a report changes is that the finding arrives as a worklist owned by the candidate, at a point where fixing it is a lookup rather than a revision.

"What about data protection?" A fair question and one with a specific answer: the check needs the citations and the reference lines, not the argument, the results or the data — we set out exactly what leaves a document, including the part that is not reassuring. For institutional measurement the constraint is tighter still, because it runs only on theses the university has already published openly.

"What if our number is bad?" It will be, on the first measurement, and so will everyone else's — that is what a check nobody has been running produces. The question worth asking is not whether the first number is flattering but whether you would rather know it before or after someone else measures it. A figure you discovered yourself, with a plan attached, is a different conversation from a figure that arrives in a sector report.

What it costs to start

Less than the conversation usually assumes, because none of the above is a new activity. Every one of those units is currently funded to support referencing and is doing it with the instruments available — which is to say, by reading. The decision in front of an institution is not whether to start spending on this. It is whether the spending already committed should continue to buy coverage of a few dozen students a term.

The practical starting point is smaller still: measure first. Ask what the figure is for your own submissions before deciding what, if anything, to do about it. A number makes the conversation concrete, and it costs nothing to be curious about.

Start with the measurement. We will run your public submissions and show you where your institution actually stands — with the evidence behind every line, and nothing about any student, thesis or supervisor published anywhere.

See plans
Filed under: For Institutions refin-report refin-score institutions universities audience academic-writing-units reference-labs training flagship
Share: Post on X Share Email

Keep reading

Institutions

The Ref[In] Index

The measurement behind the figure in this post — what it counts, what it never publishes, and how to re-run it yourself.

Read more →
Institutions

What a Ref[In] Report tells a supervisor

The artefact those six units would actually hand to a candidate — and how to use it without turning quality into accusation.

Read more →
AI & Integrity

ChatGPT gave you citations that don't exist

The first half of the AI problem, from the candidate's side of it.

Read the case →
Institutions

Citation hygiene at scale: quality, not misconduct

What changes when referencing is treated as a cohort-wide quality problem rather than a fault in one student.

Coming soon