Guides · Reading a sample

Match 10 — what the first ten pairs tell you, and what they don't

CitationLab Team · August 2026 · 8 min read
TEN PAIRS, AND WHAT THEY REACH CITATION ENTRY 8 resolved 2 open WHAT TEN PAIRS REACH is the joining sound? answered well are the sources real? a hint, not an answer is anything missing? cannot reach it at all A missing citation forms no pair — so it is never in your ten. Finding what is absent takes the whole document. FIRST TEN PAIRS — FREE Ref[In]
Ten pairs is a sample, not a survey. It answers one question well and two others not at all — which is worth knowing before you read one.

The first ten matched pairs on any document are free. That is the offer, and this is not really a post about the offer — it is a post about what a sample of ten can honestly tell you, because a diagnostic you misread is worse than no diagnostic at all.

Ten pairs answers one question properly. It gestures at a second. And it says nothing whatsoever about a third, which is the one people assume it covers.

What a matched pair actually is

The unit is not a citation and it is not a reference. It is the join between them, together with the evidence that the join is correct.

One pair, with its working shown
Citation      …outcomes improved in community settings
              (Okonkwo & Ferreira, 2022)…

Entry         Okonkwo, C. B., & Ferreira, L. (2022). Workflow
              variance in regional centres. Health Systems
              Review, 14(3), 201–218.

Resolved      record found · authors agree · year agrees
              · title agrees

Verdict       matched

That last-but-one line is the whole product. Two strings agreeing on a surname and a year is not a match — it is a coincidence that happens constantly, and it is how a citation ends up attached to a paper by a different author entirely. A pair is only matched when there is a real record behind it and you can see what it was.

The question ten pairs answers well

Is the joining sound? That is, when your citations and your entries are compared properly, do they line up — and where they don't, is the reason visible?

Ten is enough for this because joining defects are not rare events sprinkled through a list. They cluster around causes: a reference manager that exported one house style for some entries and another for the rest, a co-author's library merged in late, a transliterated surname that varies across chapters. If a cause like that is present in your document, ten pairs will almost always touch it.

Three ways ten pairs come back, and what each suggests
10 matched, evidence shown for each
   → the joining is working. Whatever else needs
     attention, it is probably not this.

8 matched, 2 open with a stated reason
   → ordinary. Look at the two, fix or dismiss,
     move on. This is the common result.

10 open, or matched on nothing but surname and year
   → stop. Something structural is wrong — most often
     a referencing style read incorrectly, which makes
     every judgement downstream meaningless.

That third row is worth dwelling on, because it is the failure that looks most like catastrophe and is usually the least serious. A thesis was once told that 90% of its citations had failed; the actual defect was a single wrong assumption about which referencing style the document used. If ten out of ten come back open, suspect the reading before you suspect the document.

The five reasons a pair comes back open

"Open" is not one finding, and this is where a sample is most often misread. Five different situations produce it, they carry completely different consequences, and the reason is always stated beside the pair. Read the reasons, not the count.

Same word, five meanings
1  IDENTIFIER WRONG
   the DOI or PMID retrieves nothing, or retrieves a
   different paper
   → usually a transcription slip. Fix the identifier.

2  ENTRY READ BADLY
   the parser ran two fields together, or split an
   author name in the wrong place
   → the source is fine; the entry needs correcting.

3  AMBIGUOUS BETWEEN TWO
   two real works share a surname and a year, and the
   citation does not distinguish them
   → only you know which one you meant.

4  REAL BUT NOT INDEXED
   a thesis, a government report, a standard, an
   older monograph — genuinely published, absent from
   the databases we can reach
   → nothing is wrong. The check cannot confirm it.

5  NO SUCH WORK
   the title, authors and year describe nothing that
   appears to exist anywhere
   → stop here. This is the one that matters.

The distance between rows two and five is the whole reason the reason is shown. A badly parsed entry and a source that does not exist are both "open", and treating them the same way would be like treating a typo and a missing chapter as the same defect.

Row four deserves particular attention, because it is the one people wrongly panic about. Bibliographic databases are excellent at journal articles and patchy at everything else. A 1974 monograph, a national statistics bulletin, an institutional thesis, an ISO standard — these are real, citable and frequently unfindable through an automated lookup. An honest check says "could not confirm", not "does not exist", and the difference is not a hedge. It is the difference between reporting what was observed and asserting what was not.

Row five is the one to act on immediately. If the title, authors and year together describe nothing that appears to exist, do not add it to a worklist for later. Search for it yourself before you do anything else, because the two explanations are that the entry is badly mangled — or that the source was never real, and citing work that was never written is academic misconduct whether or not anyone intended it.

Why ten, specifically

Ten is not a marketing round number, though it is convenient that it reads as one. It is chosen because of how referencing defects are distributed.

If defects were sprinkled randomly across a list, a sample of ten would tell you very little — you would be estimating a rate from a handful of draws, badly. But they are not random. They cluster around causes: one reference manager exporting a different house style for the entries imported before a certain date, a co-author's library merged in late, a transliterated surname that varies between chapters, a section drafted in a hurry after a supervisor meeting.

Why clustering makes a small sample informative
If defects were random, at a low rate
  ten draws would usually return ten clean pairs
  → you would learn almost nothing

Because defects cluster around causes
  a cause that is present usually touches a
  meaningful share of the list
  → ten draws will usually touch it at least once

So ten answers "is a cause present here?" well,
and "how many defects are there?" not at all.

That is the honest scope of the number. Ten pairs is a good instrument for detecting whether something systematic is wrong with how your list was assembled, and a poor instrument for counting anything. If it comes back clean, the useful conclusion is not "my list has no defects" — it is "no obvious cause is operating", which is genuinely worth knowing and is not the same claim.

The question it gestures at

Are my sources real? Each of the ten is resolved against an actual record, so ten pairs gives you ten small pieces of evidence about the authenticity of your list.

Ten is a genuine signal here but a weak one, because fabricated or hallucinated entries are not evenly distributed either — they arrive in the section where a particular tool was used, and a sample may miss that section entirely. If your concern is specifically that some references may not exist, treat ten as a first look rather than an answer. The fuller version of that check is a different exercise.

The question it cannot answer at all

Is anything missing? And this is the one people expect it to cover, so it is worth being blunt.

A pair, by definition, is made of a citation that found an entry. A citation with no entry at all forms no pair — it is not among your ten, and it cannot be, because there was nothing to join it to. The same applies to a reference nobody cited.

The most common defect in a reference list is a citation with no entry. That defect is invisible in a sample of matched pairs, because it is the absence of a pair.

Finding those requires reading the whole document and tallying in both directions — every citation forward to the list, every entry backward to the text. That is a whole-document operation, not a sampling one, and it is the check we describe separately for exactly this reason. Ten free pairs will never surface it, and a service that implied otherwise would be misleading you about its own product.

Run your first ten. Upload a document and see ten pairs resolved with the evidence shown — enough to tell whether the joining is sound before you decide to do anything more.

Try ten pairs

Reading the evidence, not the verdict

Every pair comes with the record it was resolved against, and that column is the one worth your attention — a verdict without evidence is an opinion with better typography.

The same verdict, two very different pairs
PAIR A                                      matched
  citation   (Okonkwo & Ferreira, 2022)
  entry      Okonkwo, C. B., & Ferreira, L. (2022).
             Workflow variance in regional centres.
  resolved   record retrieved · both authors agree ·
             year agrees · title agrees
  → four independent things line up. Believe it.

PAIR B                                      matched
  citation   (Chen, 2019)
  entry      Chen, L. (2019). Adherence in community
             settings.
  resolved   record retrieved · surname agrees ·
             year agrees · initial not stated in the
             citation · three works fit this shape
  → agrees on everything the citation gave, which
    was not much. Check this one yourself.

Both are matched. Only one is safe. The difference is not in the verdict, it is in how much the citation constrained the answer — and a citation of the form (Chen, 2019) constrains it very little, because a surname and a year are a weak identity in a literature where both recur constantly.

This is why we show what was compared rather than a single word. A pair that agrees on authors, year and title against a retrieved record is a different object from one that agrees on a surname and a year, even though a system reporting only verdicts would print the same thing for both. If you read nothing else in your ten, read how many things had to line up.

What ten pairs cost, and why they are free

A matched pair is not a string comparison. It is a lookup — often several, across different bibliographic sources, because the first one queried frequently does not hold the record. That is the expensive part of this product and the reason a whole document is not free.

Ten is offered without charge because the alternative is worse for both sides. A reference check is difficult to evaluate from a description: everyone claims to match citations to references, the differences only become visible on a real document, and asking someone to pay before seeing their own results is asking them to trust a claim they cannot yet test. Ten pairs is the smallest number that shows the actual behaviour — the evidence column, the reasons, the refusal to assert a match on weak agreement — on work the reader knows better than we do.

It also means the first thing you see is your own document rather than a demonstration we prepared. Any tool looks good on a case chosen by the people selling it.

Should you choose which ten?

You can, and mostly you should not.

The instinct is to submit the ten you are least sure about — the older sources, the ones from a co-author, the one you have not opened since your literature review. That gives you a reading of your worst pairs, which tells you about those pairs and nothing about the list.

An unselected ten is more informative, because it is the only version that tells you about the document. If ten arbitrary pairs come back clean, that is evidence about your process; if the ten you cherry-picked come back clean, all you have learned is that your worries were misplaced about those particular sources.

The exception is when you already have a specific suspicion — a section assembled by someone else, or a chapter written during a period you used a particular tool. Then targeting is the right move, because you are not sampling. You are checking a hypothesis.

What to do after your ten

The result points at one of three next steps, and they are genuinely different — a clean ten and a messy ten should not lead to the same action.

Ten clean. No systematic cause is operating, which means the thing most likely to be wrong with your bibliography is not the joining — it is what is absent from it, and that is the one a sample structurally cannot see. Go and do the two-directional tally, by hand if the document is short enough. That is where the defects will be.

A few open, with ordinary reasons. Identifier slips and badly parsed entries, which are twenty minutes of work. Fix them, then treat the list as you would after a clean ten — because those findings tell you nothing about what is missing either.

Open with a pattern. Several pairs failing the same way is the result worth stopping on: it means a cause is present and the ten you saw are a symptom rather than the problem. Before fixing anything individually, work out what the cause was — a particular import, a co-author's section, a chapter written in a particular month — and then check that whole region rather than the sample.

Fixing the ten pairs a sample showed you, without asking what produced them, is the one response that guarantees the rest of the list stays exactly as it was.

And a note on sequencing that saves people real time: do the style check first if you have not. If your document is Vancouver and something reads it as author–date, every pair will come back open for reasons that have nothing to do with your bibliography, and you will spend an afternoon fixing a report rather than a list.

What a poor result means, and does not

If several of your ten come back open, the useful first reading is that your reference list needs work. Merge duplicates, stale identifiers and citations that drifted from their entries are the ordinary residue of assembling a long document, and the reasoning institutions apply across a cohort applies to one person reading their own sample.

Read the reasons, though, and not just the count. "Open" covers two different things. An identifier that does not resolve is usually a typo. A source for which no record exists at all is a different finding — citing work that was never written is academic misconduct, whoever assembled that part of your list — and it is the one result worth stopping on rather than adding to a worklist.

The honest summary of ten free pairs: it will tell you quickly and cheaply whether the joining in your document is sound, hint at whether your sources are real, and stay completely silent about what is missing. Knowing which of those three you are looking at is the difference between a useful check and false reassurance.

Then run the whole thing. The full pass tallies in both directions — every citation, every entry, and the gaps a sample structurally cannot show you.

See plans
Filed under: Guides matched-review pricing citation-matching flagship
Share: Post on X Share Email

Keep reading

Guides

How to find citations missing from your bibliography

The whole-document check a sample cannot do — the one that finds what is absent.

Read the guide →
Case Files

One author, two spellings, zero matches

Why a pair needs evidence: surname and year agreeing is a coincidence, not a match.

Read the case →
Case Files

One orphan, one missing: a healthy report

What an ordinary result looks like on a whole document, so you can calibrate your ten.

Read the case →
Guides

Did someone help with your thesis?

The authenticity check in full, when a sample is not enough to settle it.

Read the guide →