How often this tool is wrong about a human

September 7, 2026 · View on GitHub

Every AI detector is asked the same question by every teacher who picks one up: how often are you wrong? None of them answer it. This is our answer, and the method is in the repository next to it.

It is not an accuracy figure. Accuracy needs machine-written text to measure against, and any collection of that is a sample of whichever models were convenient in whichever month — a number that ages badly and flatters whoever assembled it. What follows needs only writing known to be human, which does not go stale, and it measures the harm this category actually causes: detectors flag 61% of essays by non-native English speakers, and not one of them publishes that figure about itself.

What was measured

  • Corpus signsofai-human-baseline, fingerprint 78bda061bde3dc99
  • Texts 296 (452,184 words)
  • Lengths measured 649 – 9,328 words (median 832)
  • Engine SignsOfAI.Core 0.7.1
  • Run 2026-09-07
  • Target false-positive rate 5%

Every text here was written before 2022 — articles and encyclopedia revisions with a date, and classroom essays from a learner corpus collected years earlier. That is the whole basis for calling it human, and it is a stronger guarantee than any classifier offers about anything. The manifest names each source, its licence and its year, so the claim can be traced rather than trusted.

Sources

GroupTextsLicenceSource
en-anglophone-affiliation21CC BY 4.021 open-access research articles from PLOS; each DOI is in the manifest
en-other-affiliation19CC BY 4.019 open-access research articles from PLOS; each DOI is in the manifest
en-second-language-learner206CC BY-NC-ND 4.0Juffs, A., Han, N.-R. & Naismith, B. (2020). The University of Pittsburgh English Language Institute Corpus (PELIC), v1.0. doi:10.5281/zenodo.3991977 — classroom essays, first submitted version, one per student
en-wikipedia25CC BY-SA 3.0revisions of en.wikipedia.org; each revision URL is in the manifest
es-wikipedia25CC BY-SA 3.0revisions of es.wikipedia.org; each revision URL is in the manifest

The headline

At a threshold of 30/100, this tool flags at most 5% of writing known to be human — 2 of 296 texts in this corpus, an observed 0.7% with a 95% interval of 0.2% – 2.4%.

It covers documents of 649 words and up, because that is what was measured. Nothing shorter was: the corpus has no text below that length, so the boundary below is not supported there and the tool withholds its verdict rather than extrapolating. That is a statement about coverage, not about where the tool breaks — though the direction of the length effect has been measured, and it goes the wrong way: the same documents flagged 0 of 32 whole and 6 of 32 as 400-word excerpts of themselves (Docs/PARAPHRASE.md, section Length). Lowering this floor means measuring short writing people actually composed at that length, not slicing long documents into pieces.

Read the interval, not the percentage. On a small corpus an observed rate is compatible with a much wider range, and the recommendation below is made from the upper end of that range rather than the flattering one — so it stays cautious while the corpus is thin, and it follows the data in whichever direction they move as the corpus grows.

By language

A rate that holds in English and fails in Spanish is not one number, and reporting it as one would hide exactly the failure that matters here.

GroupTextsMedian90th pctHighestThreshold for 5%Best bound it can support
en2718.718.333.8301.4%
es257.215.118.4—13.3%

A dash means this group has too few texts to bound that rate at all — with nothing flagged it still takes roughly seventy-five before the interval alone gets under 5%. That is a statement about the corpus, not the tool.

By writer

The reason the whole exercise exists. If this project cannot show a rate for second-language writers that is comparable to the rate for everyone else, it has the same defect as every tool it criticises, and should say so.

GroupTextsMedian90th pctHighestThreshold for 5%Best bound it can support
en-anglophone-affiliation215.99.014.2—15.5%
en-other-affiliation196.213.618.0—16.8%
en-second-language-learner2069.519.733.8301.8%
en-wikipedia254.910.423.4—13.3%
es-wikipedia257.215.118.4—13.3%

A dash means this group has too few texts to bound that rate at all — with nothing flagged it still takes roughly seventy-five before the interval alone gets under 5%. That is a statement about the corpus, not the tool.

Across these groups the median score runs from 9.5 (en-second-language-learner) down to 4.9 (en-wikipedia), a spread of 4.6 points on a scale of a hundred. The longest tail belongs to en-second-language-learner at 19.7 for the ninetieth percentile. At the boundary this page recommends, 30/100, en-second-language-learner is flagged 2 of 206 (1%, interval 0.3% – 3.5%); the other 4 groups measured there are flagged nothing at all. It also sits highest in median and ninetieth percentile. One step down, at 25/100, en-second-language-learner would be flagged 9 of 206 (4.4%, interval 2.3% – 8.1%) — the only flags anywhere in the corpus at that boundary. That step is why the boundary does not sit at 25. That is the shape of the defect this project criticises, and it is reported here rather than averaged away — smaller than the figures published for other tools, which is a comparison, not an excuse. The groups run from tens of texts to a couple of hundred, and the numbers move as the corpus grows, in whichever direction they move.

Every threshold

Score at or aboveHuman texts flaggedRate95% interval
5253 / 29685.5%81% – 89%
10112 / 29637.8%32.5% – 43.5%
1549 / 29616.6%12.8% – 21.2%
2021 / 2967.1%4.7% – 10.6%
259 / 2963%1.6% – 5.7%
302 / 2960.7%0.2% – 2.4%
350 / 2960%0% – 1.3%
400 / 2960%0% – 1.3%
450 / 2960%0% – 1.3%
500 / 2960%0% – 1.3%
550 / 2960%0% – 1.3%
600 / 2960%0% – 1.3%
650 / 2960%0% – 1.3%
700 / 2960%0% – 1.3%
750 / 2960%0% – 1.3%
800 / 2960%0% – 1.3%
850 / 2960%0% – 1.3%
900 / 2960%0% – 1.3%
950 / 2960%0% – 1.3%
1000 / 2960%0% – 1.3%

What the product does with this number

The tool speaks at 30/100 and nowhere else, taking the boundary from the table above rather than from anybody's judgement. Below it a document gets its score and the reason it gets nothing more: a low score is not evidence that a person wrote something, since a detector that detects nothing also returns a low score, and this project has deliberately never measured how much machine writing it catches. The boundary moves when this page moves — including upward if a larger corpus turns out to be less flattering.

Above it there is one verdict, not a scale of them. This corpus can place a boundary and can say nothing whatever about how much further past it a score has travelled: no text known to be human came close to the upper reaches, and grading "moderate" against "strong" would need machine-written text, which the opening of this page argues against collecting. Interfaces do shade a high score more urgently than a low one, and those shades are a display convention — they are not on this page because nothing measured them.

A language absent from the corpus entirely gets no verdict at all, whatever it scores. A language present but too thin to bound its own rate borrows this boundary and carries its own figure beside it, so the reader weighs the rate that was measured for the writing in front of them rather than the pooled one.

Which rules misfire

Every rule below fired on text no machine wrote, so each hit is a false positive by construction — there is no judgement call to make. This is the most immediately useful thing the exercise produces: it turns "our rules probably have false positives somewhere" into a ranked list.

RuleTexts it fired onShareTotal hits
lex.just10334.8%189
stat.burstiness9532.1%95
rhet.in-conclusion8930.1%99
rhet.rule-of-three7726%143
lex.moreover6622.3%126
rhet.not-only-but6321.3%90
lex.furthermore5518.6%92
lex.crucial5016.9%61
lex.actually4916.6%69
rhet.in-order-to4113.9%87
rhet.in-terms-of227.4%38
rhet.weasel-attribution196.4%32
rhet.in-this-article165.4%24
lex.simply144.7%23
lex.facilitate124.1%22
lex.utilize103.4%30
lex.truly103.4%15
lex.comprehensive93%30
rhet.when-it-comes93%10
syn.serves-as93%10
lex.notably82.7%15
rhet.with-regard-to82.7%11
rhet.on-one-hand82.7%9
lex.profound82.7%8
lex.robust72.4%19

A rule near the top is not automatically wrong. Some tells genuinely appear in human academic prose and the catalog says so. But a rule firing on most human texts is measuring the genre rather than the machine, and should be reweighted or retired.

What this does not tell you

  • Nothing about how much AI writing it catches. That is the other half of the picture and it is not measured here, deliberately. A tool that flags nothing has a perfect false-positive rate.
  • Nothing about text unlike this corpus. These are published articles and the essays of adult learners in a university English programme. A first-year essay by a native speaker is a different population again, and the rate on one does not transfer to the other. Calibrating on your own students' pre-2022 work is the fix, and the same tool does it.
  • The affiliation groups are a proxy, not a fact. Nobody's first language is recorded in a DOI. The learner group is the exception — its corpus records each writer's first language — which is why it exists. Elsewhere the manifest states the reasoning per text so it can be argued with; where it is wrong, the number moves.
  • Hashes prove what this run measured, not that another person extracting the same articles would get identical text. They would not: PDF and HTML extraction differ. Reproducing this needs the extracted texts, not just the manifest.

Re-run it yourself:

dotnet run --project tools/SignsOfAI.Calibration -- run \
    --manifest Docs/Calibration/corpus.json --texts <your texts dir>