AI & Integrity · Human in the loop

Why every correction needs your click

CitationLab Team · August 2026 · 8 min read
TWO KINDS OF CHANGE, TWO OBLIGATIONS AI FORMED A JUDGEMENT AS IT STANDS Adeyemi K. (2019) Coastal erosion monitoring. AS PROPOSED …Marine Studies, 12(3), 44-61. BECAUSE the raw text names them ACCEPT REJECT nothing applied until you decide NO JUDGEMENT INVOLVED IN THE TEXT (WHO, 2021) IN THE REFERENCE LIST World Health Organization (2021) INITIALS AND YEAR one answer, not an opinion applied — and named on the report you can see it happened, and why An accepted proposal is a decision you made. An applied correction is something that happened to you.

A tool that fixes your references while you are not looking has not saved you any work. It has moved the work to the viva, where you will be asked why an entry changed and will not know.

Most referencing tools are judged on how much they do for you. We think that is the wrong measure, and the reason is not caution about the machine being wrong. It is that a correction nobody approved is a correction nobody can account for — and a bibliography is a claim about where your evidence came from. You have to be able to stand behind every line of it.

So the rule inside CitationLab is narrow and absolute: where the AI has formed a judgement, a person accepts or rejects it before anything is written. Not a progress bar you can cancel. Not an undo. An explicit decision, before the change exists.

The click is not friction. It is the record.

Consider what a supervisor or an examiner can actually ask you. Not "is this reference correct" — they can check that themselves. The harder question is "why is this reference different from the one in your chapter draft six weeks ago?"

If a tool silently normalised it, the honest answer is "I don't know, the software did it". That answer is worse than the original error, because the original error was a mistake and this one sounds like an absence of care.

When the change went through a click, the answer is available: a proposal was made, it named what it wanted to change and why, and you accepted it. The click is what converts a machine suggestion into an authored decision. That is the whole reason it exists.

An accepted proposal is a decision you made. An applied correction is something that happened to you. Only one of those can be defended.

What "advisory" means when the machine is right

The gate is not there because the suggestions are poor. Most of them are good, and the ones that are not are usually obviously not. The gate is there because "usually good" is not a standard anyone can submit under.

Every AI-formed proposal in the product arrives as the same shape: the entry as it stands, the entry as proposed, and the reason. Nothing is pre-selected on your behalf, and nothing is applied while you decide.

What a proposal looks like before you touch it
PROPOSED CHANGE                       reference 41

  as it stands   Adeyemi K. (2019) Coastal erosion
                 monitoring. Marine Studies.

  as proposed    Adeyemi, K. (2019). Coastal erosion
                 monitoring. Journal of Marine
                 Studies, 12(3), 44-61.

  because        the entry names a volume and pages
                 in its raw text that the parser did
                 not read into fields

  ACCEPT   REJECT                    nothing applied yet

Read the because line, because it is doing more than explaining. It is the constraint. Every correction has to be recoverable from the raw text of the entry itself, so the reason can always be checked against something you can see. A proposal that cannot say where it got a field from is a proposal that should not have been made — which is the same rule that stops the model inventing a source it never read. We wrote about that separately in what to do when the citations don't exist.

Rejecting costs you nothing and loses you nothing. The proposal stays visible with your decision beside it, which matters more than it sounds: a rejected suggestion that vanishes looks, three weeks later, exactly like a suggestion that was never made.

The one class of change that does not ask

There is an exception, and it is worth being precise about it because a rule with an undisclosed exception is not a rule.

Some transformations involve no judgement at all. When your text cites (WHO, 2021) and your bibliography lists World Health Organization for the same year, those are the same source. Recognising that is not an opinion. It is the initials of the words, in order, and a matching year — a deterministic test with one answer, which any two people would reach independently.

Asking you to approve that would not be caution. It would be theatre: a dialogue box offering you a choice you have no information to make differently, twenty times in a row, training you to click accept without reading. Consent that is always given stops being consent.

So a transformation with no judgement behind it is applied. But — and this is the half that took a ruling to settle — it is always disclosed.

Two kinds of change, two different obligations
  AI FORMED A JUDGEMENT          NO JUDGEMENT INVOLVED

  proposed to you                applied
  you accept or reject           stated on the report
  your decision is recorded      the rule is named

  "the raw text names a          "(WHO, 2021) merged with
   volume the parser              World Health Organization
   did not read"                  - initials and year"

  nothing applied until          you can see it happened
  you decide                     and why

The disclosure is the load-bearing part. A merge like that is otherwise invisible — the citation and the reference simply line up, and a reader assumes they always did. If it is ever wrong, silence would mean you never find out. Naming it costs one line and turns an invisible transformation into something you can disagree with.

Upload a chapter and see the proposals — extraction and cross-check are free

Why not simply apply everything and let people undo?

Because undo is not symmetric with approval, and the difference is not cosmetic.

To undo a change you must first notice it. In a bibliography of two hundred entries, changed quietly and correctly ninety-five times, the five that were wrong are invisible — they look exactly like the ninety-five, and you have no reason to go looking. Approval inverts that: nothing is hidden, because nothing has happened yet.

There is a second reason, and it is about who the tool is for. A doctoral candidate is not a passive recipient of a corrected document. They are the author, and the references are part of the argument — which source, which edition, which year. A tool that quietly decides those has taken a position on the work itself. Ours declines to.

What survives when the document is rebuilt

A run is not a single pass. You cross-check, you fix things, you upload a revised chapter, you run again. If your decisions did not survive that, the gate would be a tax you pay repeatedly — and paying it repeatedly is what teaches people to click through it.

So decisions are kept in their own record, not glued to the document they were made about. A citation you promoted back from the screened rows stays promoted, and the report records that you promoted it rather than the AI. A reference you edited stays edited when the analysis is rebuilt around it. We wrote about the conservation rule behind that in nothing gets lost.

The practical effect is that the click is asked for once per decision, not once per run. Which is the only way a gate this strict stays tolerable across the months a thesis actually takes.

The report says who decided, not just what changed

There is a category of row that starts life outside your bibliography: things the extractor picked up that look like citations and are not — a table caption, a figure credit, a line of an address. Those are screened out, and screening them out is itself a judgement that can be wrong.

So a screened row can be put back. Either you look at it and say that is a citation, or the AI review looks at it and says the same. Both routes reinstate it, and both are legitimate — but they are not the same event, and the certificate does not pretend they are. Each reinstated row is printed with which of the two put it back.

That distinction is small until somebody asks about it. "The tool flagged this and I overruled it" and "the AI review recovered this" are different statements about how much attention a line received, and the person reading your report is entitled to know which one they are looking at. A record that flattened them into "reinstated" would be tidier and less true.

The same principle runs through the rest of the report. It states what was measured, by what route, and where a human intervened — and it declines to state anything it did not measure. An audit trail that is generous with itself is not an audit trail.

What this does not claim

None of this is a check for misconduct, and the report never reads as one. It reports whether your citations and your references line up with each other and with their sources. A missing entry is nearly always a chapter reordered at midnight, not a person hiding something — and the Ref[In] Report is written to be shown to a supervisor, which is not a document you would want phrased as an accusation.

Nor does the gate make the AI accurate. It makes it accountable. A wrong suggestion you rejected leaves no mark on your thesis; a wrong suggestion applied silently leaves a reference you will one day have to explain. The click is the difference between those two outcomes, and it takes about a second.

Check your references — nothing changes without your say-so
Filed under: AI & Integrity human-in-the-loop research-integrity deterministic-checking
Share: Post on X Share Email

Keep reading