For Institutions · The measurement

The Ref[In] Index — measuring REFerence [IN]tegrity across universities

CitationLab Team · August 2026 · 10 min read
THE Ref[In] INDEX ONE SCORE PER INSTITUTION, PER YEAR ANONYMISED INSTITUTIONS 0 50 100 WHERE MOST SIT PUBLISHED — THE INSTITUTION SCORE, AND NOTHING ELSE student names thesis titles departments, supervisors NEVER PUBLISHED AUG 2027 2028 2029 FIRST EDITION Ref[In]
One score per institution, published annually. Nothing below the institution is ever published — and anyone can re-run the measurement on the same public theses.

Every thesis submitted to a university is screened for plagiarism. Almost none is checked for whether its references resolve. That is not an oversight anyone decided on; it is what happened when the volume of submissions outgrew the number of people who could reasonably compare four hundred citations against three hundred entries by hand.

The consequence is that nobody knows the size of the problem. Not the universities, not the examiners, not us. There is no figure, because there is no measurement.

The Ref[In] Index is our attempt to make one. The first edition publishes in August 2027 — one year after launch — and annually after that. This post is the programme rather than the paper: what it counts, what it will never publish, and how anyone who doubts a number can check it themselves.

What it counts

The unit of measurement is a submitted thesis, and the question asked of it is narrow and mechanical: does every in-text citation resolve to an entry in the bibliography, and is every entry in the bibliography actually cited?

That is deliberately not a judgement about scholarship. It says nothing about whether the sources were well chosen or the argument is sound. It is closer to an audit of the document's internal bookkeeping — the same balance we look for in any single report, aggregated.

The shape of an institutional line — illustrative figures, not a result
Institution              REF-4471   (anonymised in the published Index)
Theses scanned                 412   public submissions, last 5 years
Citations examined         196,180

  resolved to a reference     87.1%
  cited, not in the list       9.4%
  listed, never cited          3.5%

Ref[In] Index score             84   published
Everything above this line      —    never published

The figures above are made up to show the SHAPE of a line; the real ones publish in August 2027 and not before. The corpus is public theses only — submissions a university has already made openly available. Institutions have been or will be contacted about their own results. We do not scan anything that is not already published, and we do not need to: the open corpus is large enough to measure.

How a thesis is actually scanned

The method has to be stated if the number is going to be argued with, so here it is at the level of detail an institutional reader will want.

The corpus is a repository, not a sample we chose. For each institution we take the theses its own repository publishes openly, across a defined submission window. We do not select within that — no filtering by department, by supervisor, by year within the window, or by anything that would let us shape the result. A corpus you curate is a measurement of your curation.

Some documents are excluded, and the exclusions are declared. A thesis whose text is not machine-readable — a scanned image with no text layer — cannot be checked and is counted as excluded rather than as passing. So is a document with no identifiable reference list. Every excluded document is reported as a count alongside the score, because an exclusion rate is itself information: a repository that yields many unreadable files is telling you something about its own deposit process.

What "resolves" means, precisely
For each in-text citation, one of:

  MATCHED       an entry in this document's own reference
                list corresponds to it

  NOT IN LIST   no entry corresponds — the citation points
                at nothing in the bibliography

For each reference-list entry, one of:

  CITED         the text cites it somewhere
  NEVER CITED   it is listed and never used

Separately, for entries carrying an identifier:

  RESOLVES      the identifier retrieves a real record
  DOES NOT      it retrieves nothing, or something else

The first two are internal to the document. The third
reaches outside it, and is the slower half of the scan.

The first two questions are self-contained. Whether a citation has a matching entry is answerable from the document alone — no external lookup, no judgement about whether the source is any good. That is deliberate: it is the part of the measurement that cannot be disputed on grounds of what our databases happen to contain.

The third question is the one that needs the world. Checking whether an identifier retrieves a real record depends on bibliographic databases being available and correct, and they are neither perfectly. A DOI that fails to resolve on the day we check is not proof of anything, which is why an unresolved identifier is reported separately from a missing entry rather than pooled into one number.

What we do not do is judge the scholarship. Nothing in the method asks whether a source was well chosen, whether it supports the claim it is attached to, or whether the literature is current. Those are the interesting questions and they are not measurable at this scale; pretending otherwise would make the whole thing worthless.

What is never published, at any point

This matters more than the score, so it is worth stating flatly rather than in a policy page nobody reads.

The published unit is the institution. One number per university, per year. Nothing below that line is published, cited, quoted or shared — not a student name, not a thesis title, not a supervisor, not a department, not an example. There is no "worst offenders" section, and there will not be one.

The reason is not squeamishness. A measurement whose value depends on exposing the people inside it is a bad measurement, because it changes the behaviour it is trying to observe. Nobody submits honestly to a survey that can name them.

The unit of analysis is an institution's reference lists. Never its students, never its supervisors, never a single document held up as an example.

And to be explicit about the thing an institution will reasonably ask: the score is not a misconduct rate, and reading it as one would be a mistake. A gap between a citation and a reference list is overwhelmingly the ordinary residue of a long project — a source added in a late revision, a surname transliterated twice, an edition superseded while a chapter sat with a supervisor. A score of 84 does not mean 16% of a university's students did something wrong. It means one document in six has bookkeeping to finish.

The Index does, separately, count how often a cited source cannot be resolved to any record at all — which is a different category, and a smaller one. Citing work that does not exist is academic misconduct wherever it is found, and an institution is entitled to know how often it appears in what it publishes. We report that figure at the same level as everything else: per institution, never per document, and never with a name attached to it. What an institution does with its own number is its business, not ours.

Why the number is reproducible, and why that is the point

An index published by the company selling the remedy deserves scepticism. We would apply it ourselves. So the design answer is that you do not have to take our word for any of it.

The verification chain — every link is open
the theses          public, already published by the institution
the check           the same product anyone can run
the corpus          identified by the institution's own repository
the arithmetic      stated: resolved / missing / orphaned, per thesis

so                  any institution, journalist or rival can re-run the
                    measurement on the same documents and compare

and                 we hold the per-thesis evidence behind every score,
                    available to the institution the score belongs to

That is the difference between an assertion and a measurement. A number nobody can reproduce is a marketing claim wearing a percentage sign. A number anyone can reproduce on public documents, with the method stated, is something an institution can argue with — and being arguable is what makes it worth publishing.

The same discipline is why we test every change against a corpus of real theses rather than against examples we constructed. A measurement built on documents you chose is a measurement of your choices.

Want to see your own institution's figures before the Index publishes? Universities can request their report — the same arithmetic, on their own public submissions, with the per-thesis evidence behind every line.

Talk to us about an institutional report

What the preliminary scans show

Scanning is under way and the paper follows the first edition. We are not going to pre-announce a headline figure with the precision of a finished result, because the denominator is still growing and a number quoted now would be quoted for years.

What we can say is the shape. Across the institutions scanned so far — hundreds of public theses each, submitted within the last five years — the proportion of documents carrying at least one citation that does not resolve to their reference list is closer to two in five than to one in twenty. That is not a rounding error at the edge of a well-run process. It is the signature of a check that is not being run at all.

Treat that as preliminary and provisional, because it is. The full methodology, the denominators, the per-institution distribution and the confidence we place in each figure publish with the first edition in August 2027.

Three objections worth taking seriously

"You are marking your own homework." Correct, and it is the reason the verification chain above exists rather than a paragraph promising rigour. We sell the remedy, so our number should carry no authority on its own — it carries authority only to the extent that someone else can reproduce it. If an institution re-runs the check on the same public theses and gets a materially different figure, that is publishable and we will publish it. An index that cannot be embarrassed is not a measurement.

"Comparing disciplines is unfair." Partly true and worth stating. A humanities thesis with nine hundred citations, many to editions and archival material with no identifiers, is a different object from an experimental thesis with two hundred citations that mostly carry DOIs. The first has more surface for defects and less external verifiability. We report the identifier-resolution figure separately for exactly this reason, and an institution's own report breaks down by discipline where the repository allows it — but a single institutional score does average over a mix, and no amount of care makes that comparison as clean as it looks.

"A published number will distort behaviour." Yes — that is what published numbers do, and the honest position is that we want it to distort behaviour in one specific direction and have thought about the others. If institutions respond by running checks earlier and helping candidates fix lists, the measure has done its job. If they respond by teaching people to cite fewer, safer, more easily verifiable sources, it has done harm, and it would be a real cost of publishing at all.

The failure we are most worried about

There is a sharper version of that last objection, and since nobody else has raised it yet we will raise it ourselves.

The Index measures theses an institution publishes openly. The cheapest possible way to improve a score is therefore not to fix any reference lists at all — it is to publish fewer theses. An institution that closed its repository, or embargoed by default, would disappear from the measurement entirely.

A measurement that punishes openness would be worse than no measurement, because open repositories are worth more to scholarship than our index is.

We do not have a complete answer to this and will not pretend to. What we can do is refuse to build the incentive: no institution is ever listed as absent or non-compliant, there is no coverage table implying that a smaller repository is a worse one, and the published figure carries the count of theses it was computed from so that a small corpus reads as a small corpus rather than as a result. And if we see repositories closing in a pattern that follows publication, we will say so in the edition that observes it, because an index that hid its own side effects would have failed at the one thing it exists to do.

If you run a repository and think this risk is real for your institution, we would rather hear it before the first edition than after.

What to do in the year before the first edition

The Index publishes in August 2027. An institution reading this in 2026 has a year in which the number is not yet public and is still entirely movable, and that is the most useful position anyone in this will ever be in.

Ask for your own figure now. We will run an institution's public submissions and return the same arithmetic the Index will use, with the per-thesis evidence attached. It is the same method, on the same corpus, a year early — so it is not a preview of your score, it is your score, taken before anybody else can see it.

Then decide what it means before it is public. The worst version of the next eighteen months is an institution learning its figure at the same moment everyone else does, with no analysis, no plan and a press office asking for a line. The best version is a figure you have already broken down by department, already explained to yourself, and already begun to move — at which point publication is an opportunity rather than an event.

And treat the first edition as a baseline, not a verdict. One measurement describes a starting point. What an institution is actually judged on, by anyone paying attention, is the second and third editions — whether the number moved once it was known. That is the only comparison in this that carries information about an institution rather than about the sector.

Why annually, and why publish at all

Annually, because a single measurement is a fact and a series is an argument. One edition says the check is not being run. Five editions say whether anything changed after people knew.

Published, because the alternative is worse. We could measure this privately and use it only in sales conversations, and the number would count for nothing precisely because nobody could check it. Putting it in public, with the corpus identified and the method stated, is the only version of this that an institution has any reason to believe — and the only version that lets a university that disagrees demonstrate we are wrong.

If a university's own re-run produces a different number, we want to know. That is not a risk of publishing an index; it is the reason for publishing one.

See what the measurement looks like on one document. A Ref[In] Report is the per-thesis version of an Index line — every citation matched, every gap shown, with the evidence attached and nothing changed without a click.

See plans
Filed under: For Institutions refin-index refin-score refin-methodology institutions universities research-integrity verification
Share: Post on X Share Email

Keep reading

Institutions

What a Ref[In] Report tells a supervisor

The per-document version of an Index line — how to read one, and how to use it without turning quality into accusation.

Read more →
Case Files

From 1 thesis to 5,000: how we test

Why the corpus is real documents rather than examples we chose — the same discipline the Index is built on.

Read the case →
Institutions

REFerence [IN]tegrity for universities

The check that stopped running, why it stopped, and what an institution can do about it.

Coming soon
Institutions

Citation hygiene at scale: quality, not misconduct

What changes when referencing is treated as a quality problem across a cohort rather than a fault in one student.

Coming soon