AI detector accuracy: every published number

Search for "are AI detectors accurate" and most of what ranks is written by companies with a position to defend. Humaniser and paraphrasing vendors need detectors to look broken. Detector vendors need them to look near-infallible. Both quote real numbers, usually stripped of the corpus, the threshold and the funding context that give those numbers meaning.

This page collects every accuracy and false positive figure we could trace to a source we actually fetched and read, in one table, with the corpus it was measured on and who ran the test. Where a widely quoted number could not be traced to a primary source, it is flagged as unverified rather than repeated. ScriptGrain sells neither AI detection nor tools for evading it, so we have no stake in which way the numbers point.

Every published number, in one table

How to read this table: the same tool can post a 1% false positive rate in one test and 16% in another, and both can be true, because they were measured on different texts at different thresholds. The corpus column and the final column are therefore the most important ones. Vendor self-reports are not assumed to be dishonest, but they are unaudited and measured on corpora the vendor chose, so they are labelled as what they are.

One note on precision: the Stanford figure is reported as 61.3% in the published paper and by Nature; the arXiv version of the same paper states the average as 61.22%. We quote the published figure.

Study or testYearWhat was measuredNumberCorpusWho ran it (independence)
Turnitin (self-report)2023Document-level false positive rateUnder 1% (for documents scored over 20% AI)800,000 academic papers written before ChatGPT's releaseTurnitin (detector vendor)
Turnitin (self-report)2023Sentence-level false positive rateAbout 4%Same 800,000-document pre-ChatGPT testTurnitin (detector vendor)
Turnitin usage telemetry (self-report)2023Share of student submissions containing AI writing9.6% scored over 20% AI; 3.5% scored 80% to 100% AI38.5 million submissions processed to 14 May 2023Turnitin (detector vendor)
OpenAI AI Text Classifier (self-report)2023True positive rate and false positive rateCaught 26% of AI-written text; wrongly flagged 9% of human-written textOpenAI's own challenge set of English texts; tool withdrawn July 2023 for low accuracyOpenAI (vendor, reporting on its own tool)
Liang et al., Stanford (Patterns)2023False positive rate on non-native English writing, 7 detectors61.3% on average; 97.8% of essays flagged by at least one detector; 18 of 91 flagged by all seven91 TOEFL essays by Chinese speakers, written before ChatGPT existedAcademic (Stanford University)
Liang et al., Stanford (Patterns)2023False positive rate on native-speaker writing, same 7 detectors5.2% on average (described as near-perfect)88 essays by US students aged 13 to 14 (ASAP dataset)Academic (Stanford University)
Weber-Wulff et al. (Int. Journal for Educational Integrity)2023Overall classification accuracy of 14 detection toolsAll 14 below 80% accuracy; only 5 above 70%; Turnitin ranked first54 documents per tool: human-written, translated, AI-generated, edited and paraphrasedAcademic (multi-university European team)
Weber-Wulff et al.2023Accuracy on plain human-written text96%Human-written English essaysAcademic (multi-university European team)
Weber-Wulff et al.2023Accuracy on unmodified AI-generated text74%ChatGPT-generated documents, no editsAcademic (multi-university European team)
Weber-Wulff et al.2023Accuracy on AI text lightly edited by a human42%AI-generated text with manual synonym swaps (patchwriting)Academic (multi-university European team)
Weber-Wulff et al.2023Accuracy on machine-paraphrased AI text26%AI-generated text run through the Quillbot paraphraserAcademic (multi-university European team)
Walters (Open Information Science)2023Detectors performing strongly on both AI and human text3 of 16 (Copyleaks, Turnitin, Originality.ai); all tools weaker on GPT-4 than GPT-3.5AI-generated (GPT-3.5 and GPT-4) and human-written essaysAcademic (Southern Illinois University); reported by Nature, July 2026
Bloomberg Businessweek2024False positive rate of GPTZero and Copyleaks1% to 2% of essays falsely flagged, some with near 100% stated certainty500 Texas A&M application essays from summer 2022, before ChatGPT existedIndependent (journalists; no detection or humanising product)
RAID benchmark, Dugan et al. (ACL 2024)2024Robustness of 12 detectors under adversarial attackDetectors described as easily fooled by adversarial attacks, sampling changes and unseen models; vendors' claims of 99%+ accuracy found to lack rigorous validationOver 6 million generations across 11 models, 8 domains and 11 attack typesAcademic (University of Pennsylvania)
Dik, Erdem and Dik (preprint)2025GPTZero false positive rate on human essaysAbout 16%Human-written student essaysAcademic preprint; reported by Nature, July 2026
GPTZero (self-report)2026Claimed false positive rate and mixed-document accuracyFPR of no more than 1%; 96.5% accuracy on mixed AI-and-human documents; 1.1% FPR on TOEFL textsInternal and external benchmarks; corpus sizes not disclosedGPTZero (detector vendor)
Pangram (self-report)2025Claimed false positive rate0.01%, down from 2%, 1% and 0.1% in earlier versionsNot stated on the claim page; Nature separately notes independent assessments rate Pangram among the most accuratePangram (detector vendor)
Nature (its own spot test)2026ZeroGPT verdict on famous pre-AI text95% to 100% AI-generated, across repeated runsExtracts of the 1776 US Declaration of IndependenceIndependent (Nature journalists)
Russell, Karpinska et al. (preprint)2026Share of US newspaper articles flagged as partly or fully AIAbout 9%186,000 articles from 1,500 US newspapers, June to September 2025, analysed with PangramAcademic (Simon Fraser University and colleagues)

Vendor claims vs independent findings

Turnitin's verifiable current claim, stated on its product page and in its May 2023 update, is a document-level false positive rate under 1%, and that claim carries two qualifications in Turnitin's own words: it applies only to documents scored above the 20% AI threshold, and the sentence-level false positive rate is around 4%. A separate "98% accuracy" figure is widely attributed to Turnitin's FAQ, including by the aggregator detectiondrama.com, but we could not trace it to a live Turnitin page on 3 August 2026, so treat it as unverified.

GPTZero is the clearest example of the spread between self-report and independent measurement. The vendor claims a false positive rate of no more than 1%. The RAID academic benchmark recorded 95.7% recall at a fixed 1% false positive rate, a strong result. Yet a 2025 academic preprint measured its false positive rate on human-written student essays at about 16%, a 12-tool comparison reported by fast.io's 2026 aggregation scored it at 52% overall, and Weber-Wulff's team found that half of GPT Zero's positive classifications in their test would have been false accusations. Same tool, five very different numbers, five different corpora and thresholds.

Pangram claims a 0.01% false positive rate, and its own blog describes that figure falling from 2% across successive versions. The page making the claim cites no reproducible benchmark, which keeps it in the vendor-claim column, though Nature's July 2026 reporting notes that independent assessments have rated Pangram among the most accurate detectors available.

On prevalence rather than accuracy: Turnitin's own telemetry to May 2023 (38.5 million submissions) found 9.6% of submissions scored over 20% AI. The aggregator detectiondrama.com attributes later figures to Turnitin press releases, including roughly 11% of 200 million papers in the first year and a rise in mostly-AI submissions from 3.3% in 2023 to 14.8% by early 2026; we did not locate those press releases directly, so treat the later figures as aggregator-reported.

False positives: the documented harms

The single most cited harm is the Stanford finding. Seven widely used detectors, tested on 91 TOEFL essays written by non-native English speakers before ChatGPT existed, wrongly flagged them as AI-generated 61.3% of the time on average. 97.8% of those essays were flagged by at least one detector, and 18 of the 91 were flagged by all seven. The same detectors were near-perfect on essays by native-speaking US teenagers, with an average false positive rate of 5.2%. The study's authors showed the gap is driven by vocabulary and sentence-pattern richness: detectors penalise exactly the constrained, formulaic style that second-language writers are taught.

Bloomberg Businessweek's 2024 test points at the same population and adds another: testing GPTZero and Copyleaks on 500 genuinely human application essays written before ChatGPT's release, it found 1% to 2% falsely flagged, sometimes with near 100% stated certainty, and reported that the students most susceptible are those who write in a generic or mechanical way, whether because they are neurodivergent, speak English as a second language, or simply learned a plain, rule-following style.

Turnitin's own published numbers show why a small headline rate still matters at scale. Its document-level false positive rate is under 1%, but its sentence-level rate is about 4%, and its data shows false positive sentences cluster at the starts and ends of documents and next to genuine AI text: 54% sit immediately beside actual AI writing and 10% are nowhere near any. Turnitin also concedes reliability drops for documents scored under 20% AI, which is why those scores now carry an asterisk, and it raised the minimum document length from 150 to 300 words to reduce errors.

Vanderbilt University did the arithmetic and switched the feature off in August 2023: at a 1% false positive rate, the roughly 75,000 papers it submitted to Turnitin in 2022 would imply around 750 students wrongly flagged in a single year at one university. It also noted Turnitin gives no detailed information about how the detector works, making claims impossible to validate.

The failure mode is not confined to student writing. Nature ran extracts of the 1776 US Declaration of Independence through ZeroGPT repeatedly in 2026 and was told the text was 95% to 100% AI-generated; Pangram's blog attributes this class of error to famous texts being memorised by the language models that perplexity-based detectors rely on. Nature also documented the human cost: a chemistry student whose entirely human PhD application statements came back as almost 100% AI from several free detectors, and who rewrote her genuine writing to be, in her words, less perfect, before applying.

Why the numbers disagree

Different corpora. A false positive rate is a property of a detector and a text population, never of the detector alone. The same seven tools posted a 5.2% false positive rate on native-speaker essays and 61.3% on non-native essays in the same study. Tests built on polished newsroom prose, TOEFL scripts, application essays or synthetic benchmark text are measuring different things and will not agree.

Different thresholds and operating points. Turnitin's headline rate counts only documents over its 20% threshold, and scores under that threshold are marked as less reliable. RAID reports recall at a fixed false positive rate, so a figure like 95.7% recall at 1% FPR describes one chosen trade-off, not the tool's behaviour at other settings. GPTZero's error rate is under 1% if you restrict to its high-confidence predictions, and much higher if you do not. Comparing numbers taken at different operating points is comparing different machines.

Vendor incentives, in both directions. Detector vendors benchmark on corpora they select and rarely publish enough to reproduce the result; the RAID team found commercial claims of 99%+ accuracy unsupported by sufficiently challenging benchmarks. On the other side, humaniser and paraphrasing vendors publish tests designed to make detectors look as broken as possible, because that is their sales pitch. Independent and academic tests sit between the two, which is why the funder column exists.

Modified text collapses the numbers. Weber-Wulff's team measured accuracy of 96% on plain human text and 74% on unmodified AI text, falling to 42% when a human lightly edited the AI text and 26% after machine paraphrasing. They also found roughly 20% of AI-generated text was misattributed to humans overall, and that tools lean towards calling text human when unsure. A 2025 arXiv study, reported second-hand via fast.io's aggregation, found targeted adversarial paraphrasing cut detection rates by an average of 87.88%. Any test corpus containing edited or paraphrased AI text will therefore report far lower accuracy than one using raw model output.

Models and detectors both move. Walters and Elkhatat both found detectors markedly better at catching GPT-3.5 output than GPT-4 output, and newer models write more human-like text still. Pangram's claimed false positive rate changed by two orders of magnitude across its own versions. Every number on this page is dated to its study year and should be read that way.

What a detector score can and cannot support

What the published record supports using a score for: population-level measurement and triage. The 2026 newspaper study that flagged about 9% of articles as partly or fully AI-generated is a legitimate use, and its own lead researcher drew the line precisely: results at that scale can reveal trends, but not the guilt of any one author. A score can also reasonably prompt a closer human look at a document, or a conversation. Turnitin frames its tool this way itself, and GPTZero advises treating results as a conversation starter and not a final verdict.

What the record does not support: treating a score as sole evidence for an accusation. Turnitin states plainly that it does not make a determination of misconduct and that educators must apply their own judgment, because its false positive rate is not zero. OpenAI withdrew its own classifier in July 2023 over low accuracy, at a measured 26% detection rate and 9% false positive rate. Weber-Wulff and colleagues computed the false-accusation risk directly and found six of fourteen tools produced false positives, with half of GPT Zero's positive classifications amounting to false accusations in their test. Academic-integrity researcher Mike Perkins put the consensus to Nature in 2026: detectors can work in controlled tests, but the false positive concerns mean they should not be used for anything sensitive for a student.

The honest bottom line: a detector score is a probability estimate from an unaudited model, measured on someone else's corpus at someone else's threshold. In aggregate, across thousands of documents, these tools measure something real. On any single document, in either direction, the published numbers say the score is weak evidence, weaker still if the writer is a non-native speaker or writes in a plain, formulaic style, and weakest of all on text that has been edited or paraphrased after generation.

Methodology

Independence: ScriptGrain sells neither AI detection nor tools for evading detection, and has no commercial stake in whether detectors look accurate or broken. This page exists because most pages ranking for this query are published by companies on one side of that market or the other.

Every number on this page was taken from a source we fetched and read on 3 August 2026, listed below. Primary sources (the paper itself or the vendor's own page) were used wherever possible; the few figures we could only obtain second-hand are labelled as such in the text.

Widely quoted figures that could not be traced to a primary source, such as the 98% accuracy commonly attributed to Turnitin's FAQ, are flagged as unverified rather than repeated as fact.

Vendor self-reported numbers are included because the reference would be incomplete without them, and are always labelled with the vendor's name in the table's final column.

This page reports what published tests found, including how accuracy degrades under paraphrasing. It contains no guidance on evading detection, and ScriptGrain does not build or endorse tools for doing so.

Last checked: 3 August 2026. Detector versions change quickly; every figure should be read as dated to its study year.

Sources

ScriptGrain · Why ScriptGrain · Sample voice profiles · The Journal · Support