The Ref[In] Index — measuring REFerence [IN]tegrity across universities
CitationLab Team · August 2026 · 10 min read
One score per institution, published annually. Nothing below the institution is ever published — and anyone can re-run the measurement on the same public theses.
Every thesis submitted to a university is screened for plagiarism. Almost none is
checked for whether its references resolve. That is not an oversight anyone decided on;
it is what happened when the volume of submissions outgrew the number of people who
could reasonably compare four hundred citations against three hundred entries by hand.
The consequence is that nobody knows the size of the problem. Not the universities,
not the examiners, not us. There is no figure, because there is no measurement.
The Ref[In] Index
is our attempt to make one. The first edition publishes in August 2027 —
one year after launch — and annually after that. This post is the programme rather than
the paper: what it counts, what it will never publish, and how anyone who doubts a
number can check it themselves.
What it counts
The unit of measurement is a submitted thesis, and the question asked of it is narrow
and mechanical: does every in-text citation resolve to an entry in the bibliography, and
is every entry in the bibliography actually cited?
That is deliberately not a judgement about scholarship. It says nothing about whether
the sources were well chosen or the argument is sound. It is closer to an audit of the
document's internal bookkeeping — the same
balance we look for in any single report,
aggregated.
The figures above are made up to show the SHAPE of a line; the real ones publish in August 2027 and not before. The corpus is public theses only — submissions a university has
already made openly available. Institutions have been or will be contacted about their
own results. We do not scan anything that is not already published, and we do not need
to: the open corpus is large enough to measure.
How a thesis is actually scanned
The method has to be stated if the number is going to be argued with, so here it is at
the level of detail an institutional reader will want.
The corpus is a repository, not a sample we chose. For each
institution we take the theses its own repository publishes openly, across a defined
submission window. We do not select within that — no filtering by department, by
supervisor, by year within the window, or by anything that would let us shape the result.
A corpus you curate is a measurement of your curation.
Some documents are excluded, and the exclusions are declared. A thesis
whose text is not machine-readable — a scanned image with no text layer — cannot be
checked and is counted as excluded rather than as passing. So is a document with no
identifiable reference list. Every excluded document is reported as a count alongside the
score, because an exclusion rate is itself information: a repository that yields many
unreadable files is telling you something about its own deposit process.
The first two questions are self-contained. Whether a citation has a
matching entry is answerable from the document alone — no external lookup, no judgement
about whether the source is any good. That is deliberate: it is the part of the
measurement that cannot be disputed on grounds of what our databases happen to contain.
The third question is the one that needs the world. Checking whether
an identifier retrieves a real record depends on bibliographic databases being available
and correct, and they are neither perfectly. A DOI that fails to resolve on the day we
check is not proof of anything, which is why an unresolved identifier is reported
separately from a missing entry rather than pooled into one number.
What we do not do is judge the scholarship. Nothing in the method asks whether a source
was well chosen, whether it supports the claim it is attached to, or whether the
literature is current. Those are the interesting questions and they are not measurable at
this scale; pretending otherwise would make the whole thing worthless.
What is never published, at any point
This matters more than the score, so it is worth stating flatly rather than in a
policy page nobody reads.
The published unit is the institution. One number per university,
per year. Nothing below that line is published, cited, quoted or shared — not a student
name, not a thesis title, not a supervisor, not a department, not an example. There is
no "worst offenders" section, and there will not be one.
The reason is not squeamishness. A measurement whose value depends on exposing the
people inside it is a bad measurement, because it changes the behaviour it is trying to
observe. Nobody submits honestly to a survey that can name them.
The unit of analysis is an institution's reference lists. Never its students,
never its supervisors, never a single document held up as an example.
And to be explicit about the thing an institution will reasonably ask: the score is
not a misconduct rate, and reading it as one would be a mistake. A gap between a citation
and a reference list is overwhelmingly the ordinary residue of a long project — a source
added in a late revision, a surname transliterated twice, an edition superseded while a
chapter sat with a supervisor. A score of 84 does not mean 16% of a university's students
did something wrong. It means one document in six has bookkeeping to finish.
The Index does, separately, count how often a cited source cannot be resolved to any
record at all — which is a different category, and a smaller one. Citing work that does
not exist is academic misconduct wherever it is found, and an institution is entitled to
know how often it appears in what it publishes. We report that figure at the same level
as everything else: per institution, never per document, and never with a name attached
to it. What an institution does with its own number is its business, not ours.
Why the number is reproducible, and why that is the point
An index published by the company selling the remedy deserves scepticism. We would
apply it ourselves. So the design answer is that you do not have to take our word
for any of it.
That is the difference between an assertion and a measurement. A number nobody can
reproduce is a marketing claim wearing a percentage sign. A number anyone can reproduce
on public documents, with the method stated, is something an institution can argue with
— and being arguable is what makes it worth publishing.
The same discipline is why we
test every change against a corpus of
real theses rather than against examples we constructed. A measurement built on
documents you chose is a measurement of your choices.
Want to see your own institution's figures before the Index publishes?
Universities can request their report — the same arithmetic, on their own public
submissions, with the per-thesis evidence behind every line.
Talk to us about an institutional report
What the preliminary scans show
Scanning is under way and the paper follows the first edition. We are not going to
pre-announce a headline figure with the precision of a finished result, because the
denominator is still growing and a number quoted now would be quoted for years.
What we can say is the shape. Across the institutions scanned so far — hundreds of
public theses each, submitted within the last five years — the proportion of documents
carrying at least one citation that does not resolve to their reference list is
closer to two in five than to one in twenty. That is not a rounding
error at the edge of a well-run process. It is the signature of a check that is not
being run at all.
Treat that as preliminary and provisional, because it is. The full methodology, the
denominators, the per-institution distribution and the confidence we place in each
figure publish with the first edition in August 2027.
Three objections worth taking seriously
"You are marking your own homework." Correct, and it is the reason
the verification chain above exists rather than a paragraph promising rigour. We sell the
remedy, so our number should carry no authority on its own — it carries authority only
to the extent that someone else can reproduce it. If an institution re-runs the check on
the same public theses and gets a materially different figure, that is publishable and we
will publish it. An index that cannot be embarrassed is not a measurement.
"Comparing disciplines is unfair." Partly true and worth stating.
A humanities thesis with nine hundred citations, many to editions and archival material
with no identifiers, is a different object from an experimental thesis with two hundred
citations that mostly carry DOIs. The first has more surface for defects and less
external verifiability. We report the identifier-resolution figure separately for exactly
this reason, and an institution's own report breaks down by discipline where the
repository allows it — but a single institutional score does average over a mix, and no
amount of care makes that comparison as clean as it looks.
"A published number will distort behaviour." Yes — that is what
published numbers do, and the honest position is that we want it to distort behaviour in
one specific direction and have thought about the others. If institutions respond by
running checks earlier and helping candidates fix lists, the measure has done its job. If
they respond by teaching people to cite fewer, safer, more easily verifiable sources, it
has done harm, and it would be a real cost of publishing at all.
The failure we are most worried about
There is a sharper version of that last objection, and since nobody else has raised it
yet we will raise it ourselves.
The Index measures theses an institution publishes openly. The cheapest possible way
to improve a score is therefore not to fix any reference lists at all — it is to publish
fewer theses. An institution that closed its repository, or embargoed by default, would
disappear from the measurement entirely.
A measurement that punishes openness would be worse than no measurement,
because open repositories are worth more to scholarship than our index is.
We do not have a complete answer to this and will not pretend to. What we can do is
refuse to build the incentive: no institution is ever listed as absent or non-compliant,
there is no coverage table implying that a smaller repository is a worse one, and the
published figure carries the count of theses it was computed from so that a small corpus
reads as a small corpus rather than as a result. And if we see repositories closing in a
pattern that follows publication, we will say so in the edition that observes it, because
an index that hid its own side effects would have failed at the one thing it exists to
do.
If you run a repository and think this risk is real for your institution, we would
rather hear it before the first edition than after.
What to do in the year before the first edition
The Index publishes in August 2027. An institution reading this in 2026 has a year in
which the number is not yet public and is still entirely movable, and that is the most
useful position anyone in this will ever be in.
Ask for your own figure now. We will run an institution's public
submissions and return the same arithmetic the Index will use, with the per-thesis
evidence attached. It is the same method, on the same corpus, a year early — so it is not
a preview of your score, it is your score, taken before anybody else can see it.
Then decide what it means before it is public. The worst version of
the next eighteen months is an institution learning its figure at the same moment
everyone else does, with no analysis, no plan and a press office asking for a line. The
best version is a figure you have already broken down by department, already explained to
yourself, and already begun to move — at which point publication is an opportunity rather
than an event.
And treat the first edition as a baseline, not a verdict. One
measurement describes a starting point. What an institution is actually judged on, by
anyone paying attention, is the second and third editions — whether the number moved once
it was known. That is the only comparison in this that carries information about an
institution rather than about the sector.
Why annually, and why publish at all
Annually, because a single measurement is a fact and a series is an argument. One
edition says the check is not being run. Five editions say whether anything changed
after people knew.
Published, because the alternative is worse. We could measure this privately and use
it only in sales conversations, and the number would count for nothing precisely because
nobody could check it. Putting it in public, with the corpus identified and the method
stated, is the only version of this that an institution has any reason to believe — and
the only version that lets a university that disagrees demonstrate we are wrong.
If a university's own re-run produces a different number, we want to know. That is
not a risk of publishing an index; it is the reason for publishing one.
See what the measurement looks like on one document. A
Ref[In] Report
is the per-thesis version of an Index line — every citation matched, every gap shown,
with the evidence attached and nothing changed without a click.
See plans