Citation hygiene at scale: quality, not misconduct
CitationLab Team · August 2026 · 10 min read
The same finding, sent through two doors. One produces a corrected bibliography; the other produces a defensive student and no correction at all.
An institution that starts measuring REFerence [IN]tegrity discovers something
uncomfortable within a week: the numbers are worse than anyone expected, and they are
worse across the board rather than concentrated in a few documents.
What happens next determines whether the measurement was worth taking. There are two
available readings of a cohort-wide error rate, they lead to opposite operational
responses, and only one of them results in better reference lists.
What a rate at scale actually tells you
Suppose a graduate school checks a season's submissions and finds that a substantial
minority carry at least one citation that resolves to nothing. The instinctive reading
is that a substantial minority of candidates were careless.
That reading does not survive contact with the arithmetic.
When a defect appears at a rate, in a population selected and trained for care, the
defect is telling you about the process rather than the people in it. Human attention
does not fail uniformly at that scale unless the task itself is beyond it — and
comparing six hundred citations against three hundred entries, by reading, is beyond it.
A finding that is evenly distributed across a cohort is a description of the
system the cohort is working inside.
This is not a charitable interpretation offered to be kind. It is the one that
predicts what you will find when you look, and it is the reason the
measurement we publish annually is
reported per institution rather than per document. The institution is the unit at which
the cause actually lives.
What it costs to call it misconduct
The other reading — that these are integrity failures on the part of candidates —
costs an institution three things, and it costs them quickly.
The candidates stop telling you the truth. A student who believes a
referencing report can be used against them will not bring an uncertain reference to a
supervisor to discuss. They will quietly delete the citation, which removes the evidence
and leaves the claim unsupported. You have converted a fixable bookkeeping problem into
an unfixable argumentative one.
The supervisors stop running it. Nobody wants to be the person who
initiated a conduct process against their own candidate over a missing bibliography
entry, and a supervisor who suspects that is where a report leads will simply not
generate one. The check dies at exactly the point it would have been useful.
And the finding is usually wrong on its own terms. The archetypal
case is not a fabricated source. It is
a surname transliterated two ways across four
years — one real source, listed twice, cited once, presenting as an orphan and a
missing entry simultaneously. There is no version of that which is a conduct matter, and
an institution that treats it as one will have spent a hearing on a hyphen.
The one finding that is not a quality matter
Everything above is about a rate, and a rate is a property of a process. There is one
result that does not belong in that argument, and an institution should hold it apart.
A citation whose source cannot be found to exist is not bookkeeping. Presenting work
that was never written in support of a claim is academic misconduct under every code we
have read, and unlike the rest of this it is a finding rather than an inference. It is
also rare relative to the noise around it, which is precisely why it gets lost when a
report is read as a single number: the nine hundred bookkeeping items bury the one that
matters.
Counting the two separately is the practical consequence of everything in this post.
A cohort's missing-entry rate tells you about your process. A cohort's unresolvable-source
count tells you something else entirely, and conflating them wastes both.
What a supervisor should do with a report
Treat it as a worklist, in the order it is cheapest to clear. The missing entries
first, because they are usually a single lookup. The orphans second, because most are
cut literature and the decision is simply whether to cite or remove. The weak matches
last, and with judgement, because that is the only part where a person adds something a
machine cannot.
The framing that works in the room is that the report is about the document, and both
people are on the same side of it. We wrote a fuller guide to
reading one without turning quality into
accusation, but the short version is that the first sentence out of a supervisor's
mouth decides which conversation follows.
And there is a specific caution worth carrying into that meeting. A report showing an
implausibly large number of failures is far more likely to be a bad reading than a bad
candidate — a thesis was once
told 90% of its citations had failed, and the real cause was one wrong assumption
about referencing style. Before anyone discusses anything with a student, the report
itself has to be sound.
Run it across a cohort, not a case. An institutional report shows
where a department's reference lists actually stand, with per-thesis evidence — and
nothing published about any student, supervisor or document.
Talk to us about an institutional report
Three things an institution should never do with this
Never rank people. Not students, not supervisors, not departments
against each other. The moment a number is attached to a person it stops measuring the
document and starts measuring how well that person manages the number. Everything after
that is theatre.
Never route it into a conduct process by default. There are real
integrity cases in academia and they look nothing like this — they involve sources that
do not exist, results that were not obtained, or text that was not written. A missing
bibliography entry is not evidence of any of those, and a process that cannot tell the
difference will be both unjust and, quickly, ignored.
Never let the score become the goal. A cohort can be taught to
produce clean reports without producing better scholarship — cite less, cite safer,
avoid the source you have not fully read. That is a worse outcome than the problem, and
it is what happens when a measure is used to judge rather than to improve.
What good looks like
The institutions doing this well have converged on roughly the same posture, and it
is unglamorous. The check runs early, when a correction is cheap and nobody is close to
a deadline. It runs for everyone, so nobody is selected for it and no one has to ask. It
produces a worklist owned by the candidate. And the aggregate is watched at cohort
level, where a rate is genuinely informative, while the individual result stays between
the candidate and whoever is helping them.
That last division is the whole design. The institution learns whether its process is
working; the individual gets help with their document; and neither of those uses
contaminates the other. The moment the aggregate is assembled from named individuals, or
the individual result is reported upward, you lose both.
None of this requires believing the best about everybody. It requires noticing that a
defect appearing at a rate, across a population, in a task nobody is equipped to do by
hand, is a description of the task — and that fixing tasks is considerably more
productive than investigating people.
Start with the measurement, not the intervention. See where your
cohort's reference lists actually stand before deciding what to change — with the
evidence behind every line, and nothing attributed to anyone.
See plans