What a Ref[In] Report tells a supervisor
CitationLab Team · July 2026 · 7 min read
The cover, read in order: score to calibrate, missing first, orphans for scale, currency at a glance.
A supervisor with six theses on their desk does not have a spare week
per bibliography. A REFerence [IN]tegrity report exists for exactly that
desk: one page that says how a document's citations and its reference list actually
relate — what matches, what's missing, what's never used — before anyone opens
chapter one.
This post walks through the report we produce — formally the Academic Reference
Integrity Report, the
Ref[In] Report
for short — the way a supervisor reads it: the score first, the two tables second, and
the boundaries of what it claims third. The third part matters most, so we'll say it
here too: it is a quality instrument, not a misconduct detector, and
it is built to keep those two ideas apart.
What the report actually measures
Everything in the report rests on one mechanical exercise: read every in-text
citation, read every reference-list entry, and pair them. Done completely —
with every line accounted for — that pairing
exposes the two defects that matter most to a reader:
- Missing references — a source is cited in the text, but no
entry for it exists in the list. A reader who wants the source cannot get there.
This is the serious one: it breaks the contract a citation makes.
- Orphaned entries — an entry sits in the list, but nothing in
the text ever cites it. Usually the residue of drafting: a source that was cut, a
list that was padded, an import that never got used. Less serious, still worth
clearing — an examiner notices a list that doesn't match its text.
The report presents each as a table you can act on: the citation as written, where
it appears, and what the checker looked for. Beyond the pairing, a currency panel
summarises how old the cited sources run, and an activity summary records what was
checked and how — so the report is legible as evidence, not just as a verdict.
How to read the score
The headline number — the
Ref[In] Score
— compresses that pairing into a single figure of bibliography health. You don't need
the formula to read it well; you need to know three things about how it behaves.
- Only real defects move it. The score is computed from the
missing and orphaned tables — the two things above that a writer should actually
fix.
- Missing weighs heavier than orphaned. A citation pointing at
nothing harms a reader more than an entry nobody cites, and the score reflects that
asymmetry rather than counting all defects alike.
- Open questions never lower it. Where the checker could not
decide — an ambiguous pairing it flags for a human — the ambiguity is listed for
review but costs the writer nothing. A tool that docked points for its own
uncertainty would be grading its confidence, not the bibliography.
The score is designed to be argued with: every point it loses is a row in
a table you can check by hand.
Reading one as a supervisor
The practical routine is short. Glance at the score to calibrate attention. Go
straight to the missing table — those rows are the feedback that saves a student real
pain later, because missing citations are
what examiners chase to the library and fail to find. Skim the orphan table for
scale rather than detail: three orphans are drafting residue, thirty are a
literature-review habit worth a conversation. Then look at the currency panel once —
if a thesis on a fast-moving field leans on decade-old sources, that is a supervision
point no citation table will surface on its own.
Timing matters as much as reading. The report earns most of its value when it
arrives before the supervision meeting rather than during it: the student
runs the draft, fixes the mechanical rows themselves — a missing entry is not a
conversation, it's a chore — and brings the residue that actually needs judgement.
Meetings stop being proofreading sessions; the ten minutes once spent discovering
that chapter 5's references were never entered get spent on whether the argument's
evidence base is the right one. The instrument's job is to remove the findable from
the meeting so the discussable can have it.
What does "good" look like? Not zero. Real, honest bibliographies almost always
carry a stray orphan or one unresolved citation —
a healthy report usually has something in
it. The signal a supervisor is reading is not perfection; it is whether the
defects are few, explained, and shrinking between drafts.
Supervising this term? Have your student run their draft and
bring the report to the meeting — the tables make a better agenda than "check your
references again".
See how students run it
What the report is not
Boundaries, stated plainly, because a document like this gets misused when they
are implied:
- It is not a plagiarism report. It never compares a student's
prose to anyone else's. It compares a text to its own reference list.
- It is not a misconduct finding. A missing reference is a
defect, not an act. The report deliberately has no vocabulary for intent — a
formatting habit and an invented source produce the same row, and distinguishing
them is a human's job, made easier because the row exists.
- It is not a grade. A brilliant argument can sit on a messy
bibliography and a hollow one on a spotless list. The report measures the
scaffolding, so the reader can spend their judgement on the building.
This framing is also why the report works at cohort scale. A department that runs
drafts through the same instrument gets a consistent, anonymisable picture of
citation hygiene across a programme — where the common defects cluster, whether they
shrink after supervision — without any individual ever being accused of anything by a
machine. Quality, not misconduct, is the entire posture.
Why it carries a name
The report and score are named — Ref[In],
for REFerence [IN]tegrity — because an unnamed number is unaccountable. Naming
the standard pins what the figure means across documents, drafts and years: a score
on one thesis is comparable to a score on another because both were produced by the
same published discipline — complete pairing,
conservation of every line, defects counted
only when a human could verify them from the table. When a bibliography has been
refreshed — say after updating
outdated references before submission — re-running the report shows the
improvement in the same currency.
Questions supervisors actually ask us
Can a student game the score?
The crude route — deleting citations to empty the missing table — doesn't work
quietly, because every deletion is a recorded movement in the run's own accounting,
and a bibliography that shrank between drafts is more conspicuous than one with two
honest defects. The productive way to raise the score is indistinguishable from
doing the work: add the missing entries, resolve the orphans. An instrument you can
only game by genuinely improving the document is gamed in the right direction.
Is the score comparable across disciplines?
The pairing mechanics don't care whether the thesis is nursing or numerical
analysis — a citation pointing at nothing is the same defect everywhere, which is
what makes the number comparable. What differs by field is the currency
panel: reference age reads differently in machine learning than in classics, so
age is presented as context alongside the score rather than folded into it.
What does re-running do to the record?
Each run reports the draft it was given; a re-run after fixes is a new snapshot,
not an overwrite of history. The useful supervision artefact is the pair: the
before, the after, and the delta — evidence of process, which for most committees
matters more than either snapshot alone.
Can we run it across a cohort without singling anyone out?
That is the designed use at programme level: aggregated, anonymised
distributions — how many drafts sit in each band, which defect classes dominate —
with individual reports staying between student and supervisor. The programme
learns where to aim teaching; nobody's name is attached to a number. Opt-in,
aggregate-first, and framed as writing support: those aren't courtesies bolted on,
they're the conditions under which measurement helps instead of chills.
The two-minute takeaway
- Score first — calibration, not verdict.
- Missing table — the rows worth a student's immediate attention.
- Orphan table — scale tells you if it's residue or habit.
- Currency panel — one glance, one possible conversation.
- And nothing in it accuses anyone of anything. That part stays human.
For programmes and libraries: we work with institutions on
cohort-level citation quality — anonymised, opt-in, and framed as support.
Ref[In] Reports
for every draft, one standard across the department.
Talk to us