AI detector accuracy: every published number
Search for "are AI detectors accurate" and most of what ranks is written by companies with a position to defend. Humaniser and paraphrasing vendors need detectors to look broken. Detector vendors need them to look near-infallible. Both quote real numbers, usually stripped of the corpus, the threshold and the funding context that give those numbers meaning.
This page collects every accuracy and false positive figure we could trace to a source we actually fetched and read, in one table, with the corpus it was measured on and who ran the test. Where a widely quoted number could not be traced to a primary source, it is flagged as unverified rather than repeated. ScriptGrain sells neither AI detection nor tools for evading it, so we have no stake in which way the numbers point.
Every published number, in one table
How to read this table: the same tool can post a 1% false positive rate in one test and 16% in another, and both can be true, because they were measured on different texts at different thresholds. The corpus column and the final column are therefore the most important ones. Vendor self-reports are not assumed to be dishonest, but they are unaudited and measured on corpora the vendor chose, so they are labelled as what they are.
One note on precision: the Stanford figure is reported as 61.3% in the published paper and by Nature; the arXiv version of the same paper states the average as 61.22%. We quote the published figure.
| Study or test | Year | What was measured | Number | Corpus | Who ran it (independence) |
|---|---|---|---|---|---|
| Turnitin (self-report) | 2023 | Document-level false positive rate | Under 1% (for documents scored over 20% AI) | 800,000 academic papers written before ChatGPT's release | Turnitin (detector vendor) |
| Turnitin (self-report) | 2023 | Sentence-level false positive rate | About 4% | Same 800,000-document pre-ChatGPT test | Turnitin (detector vendor) |
| Turnitin usage telemetry (self-report) | 2023 | Share of student submissions containing AI writing | 9.6% scored over 20% AI; 3.5% scored 80% to 100% AI | 38.5 million submissions processed to 14 May 2023 | Turnitin (detector vendor) |
| OpenAI AI Text Classifier (self-report) | 2023 | True positive rate and false positive rate | Caught 26% of AI-written text; wrongly flagged 9% of human-written text | OpenAI's own challenge set of English texts; tool withdrawn July 2023 for low accuracy | OpenAI (vendor, reporting on its own tool) |
| Liang et al., Stanford (Patterns) | 2023 | False positive rate on non-native English writing, 7 detectors | 61.3% on average; 97.8% of essays flagged by at least one detector; 18 of 91 flagged by all seven | 91 TOEFL essays by Chinese speakers, written before ChatGPT existed | Academic (Stanford University) |
| Liang et al., Stanford (Patterns) | 2023 | False positive rate on native-speaker writing, same 7 detectors | 5.2% on average (described as near-perfect) | 88 essays by US students aged 13 to 14 (ASAP dataset) | Academic (Stanford University) |
| Weber-Wulff et al. (Int. Journal for Educational Integrity) | 2023 | Overall classification accuracy of 14 detection tools | All 14 below 80% accuracy; only 5 above 70%; Turnitin ranked first | 54 documents per tool: human-written, translated, AI-generated, edited and paraphrased | Academic (multi-university European team) |
| Weber-Wulff et al. | 2023 | Accuracy on plain human-written text | 96% | Human-written English essays | Academic (multi-university European team) |
| Weber-Wulff et al. | 2023 | Accuracy on unmodified AI-generated text | 74% | ChatGPT-generated documents, no edits | Academic (multi-university European team) |
| Weber-Wulff et al. | 2023 | Accuracy on AI text lightly edited by a human | 42% | AI-generated text with manual synonym swaps (patchwriting) | Academic (multi-university European team) |
| Weber-Wulff et al. | 2023 | Accuracy on machine-paraphrased AI text | 26% | AI-generated text run through the Quillbot paraphraser | Academic (multi-university European team) |
| Walters (Open Information Science) | 2023 | Detectors performing strongly on both AI and human text | 3 of 16 (Copyleaks, Turnitin, Originality.ai); all tools weaker on GPT-4 than GPT-3.5 | AI-generated (GPT-3.5 and GPT-4) and human-written essays | Academic (Southern Illinois University); reported by Nature, July 2026 |
| Bloomberg Businessweek | 2024 | False positive rate of GPTZero and Copyleaks | 1% to 2% of essays falsely flagged, some with near 100% stated certainty | 500 Texas A&M application essays from summer 2022, before ChatGPT existed | Independent (journalists; no detection or humanising product) |
| RAID benchmark, Dugan et al. (ACL 2024) | 2024 | Robustness of 12 detectors under adversarial attack | Detectors described as easily fooled by adversarial attacks, sampling changes and unseen models; vendors' claims of 99%+ accuracy found to lack rigorous validation | Over 6 million generations across 11 models, 8 domains and 11 attack types | Academic (University of Pennsylvania) |
| Dik, Erdem and Dik (preprint) | 2025 | GPTZero false positive rate on human essays | About 16% | Human-written student essays | Academic preprint; reported by Nature, July 2026 |
| GPTZero (self-report) | 2026 | Claimed false positive rate and mixed-document accuracy | FPR of no more than 1%; 96.5% accuracy on mixed AI-and-human documents; 1.1% FPR on TOEFL texts | Internal and external benchmarks; corpus sizes not disclosed | GPTZero (detector vendor) |
| Pangram (self-report) | 2025 | Claimed false positive rate | 0.01%, down from 2%, 1% and 0.1% in earlier versions | Not stated on the claim page; Nature separately notes independent assessments rate Pangram among the most accurate | Pangram (detector vendor) |
| Nature (its own spot test) | 2026 | ZeroGPT verdict on famous pre-AI text | 95% to 100% AI-generated, across repeated runs | Extracts of the 1776 US Declaration of Independence | Independent (Nature journalists) |
| Russell, Karpinska et al. (preprint) | 2026 | Share of US newspaper articles flagged as partly or fully AI | About 9% | 186,000 articles from 1,500 US newspapers, June to September 2025, analysed with Pangram | Academic (Simon Fraser University and colleagues) |
Vendor claims vs independent findings
Turnitin's verifiable current claim, stated on its product page and in its May 2023 update, is a document-level false positive rate under 1%, and that claim carries two qualifications in Turnitin's own words: it applies only to documents scored above the 20% AI threshold, and the sentence-level false positive rate is around 4%. A separate "98% accuracy" figure is widely attributed to Turnitin's FAQ, including by the aggregator detectiondrama.com, but we could not trace it to a live Turnitin page on 3 August 2026, so treat it as unverified.
GPTZero is the clearest example of the spread between self-report and independent measurement. The vendor claims a false positive rate of no more than 1%. The RAID academic benchmark recorded 95.7% recall at a fixed 1% false positive rate, a strong result. Yet a 2025 academic preprint measured its false positive rate on human-written student essays at about 16%, a 12-tool comparison reported by fast.io's 2026 aggregation scored it at 52% overall, and Weber-Wulff's team found that half of GPT Zero's positive classifications in their test would have been false accusations. Same tool, five very different numbers, five different corpora and thresholds.
Pangram claims a 0.01% false positive rate, and its own blog describes that figure falling from 2% across successive versions. The page making the claim cites no reproducible benchmark, which keeps it in the vendor-claim column, though Nature's July 2026 reporting notes that independent assessments have rated Pangram among the most accurate detectors available.
On prevalence rather than accuracy: Turnitin's own telemetry to May 2023 (38.5 million submissions) found 9.6% of submissions scored over 20% AI. The aggregator detectiondrama.com attributes later figures to Turnitin press releases, including roughly 11% of 200 million papers in the first year and a rise in mostly-AI submissions from 3.3% in 2023 to 14.8% by early 2026; we did not locate those press releases directly, so treat the later figures as aggregator-reported.
False positives: the documented harms
The single most cited harm is the Stanford finding. Seven widely used detectors, tested on 91 TOEFL essays written by non-native English speakers before ChatGPT existed, wrongly flagged them as AI-generated 61.3% of the time on average. 97.8% of those essays were flagged by at least one detector, and 18 of the 91 were flagged by all seven. The same detectors were near-perfect on essays by native-speaking US teenagers, with an average false positive rate of 5.2%. The study's authors showed the gap is driven by vocabulary and sentence-pattern richness: detectors penalise exactly the constrained, formulaic style that second-language writers are taught.
Bloomberg Businessweek's 2024 test points at the same population and adds another: testing GPTZero and Copyleaks on 500 genuinely human application essays written before ChatGPT's release, it found 1% to 2% falsely flagged, sometimes with near 100% stated certainty, and reported that the students most susceptible are those who write in a generic or mechanical way, whether because they are neurodivergent, speak English as a second language, or simply learned a plain, rule-following style.
Turnitin's own published numbers show why a small headline rate still matters at scale. Its document-level false positive rate is under 1%, but its sentence-level rate is about 4%, and its data shows false positive sentences cluster at the starts and ends of documents and next to genuine AI text: 54% sit immediately beside actual AI writing and 10% are nowhere near any. Turnitin also concedes reliability drops for documents scored under 20% AI, which is why those scores now carry an asterisk, and it raised the minimum document length from 150 to 300 words to reduce errors.
Vanderbilt University did the arithmetic and switched the feature off in August 2023: at a 1% false positive rate, the roughly 75,000 papers it submitted to Turnitin in 2022 would imply around 750 students wrongly flagged in a single year at one university. It also noted Turnitin gives no detailed information about how the detector works, making claims impossible to validate.
The failure mode is not confined to student writing. Nature ran extracts of the 1776 US Declaration of Independence through ZeroGPT repeatedly in 2026 and was told the text was 95% to 100% AI-generated; Pangram's blog attributes this class of error to famous texts being memorised by the language models that perplexity-based detectors rely on. Nature also documented the human cost: a chemistry student whose entirely human PhD application statements came back as almost 100% AI from several free detectors, and who rewrote her genuine writing to be, in her words, less perfect, before applying.
Why the numbers disagree
Different corpora. A false positive rate is a property of a detector and a text population, never of the detector alone. The same seven tools posted a 5.2% false positive rate on native-speaker essays and 61.3% on non-native essays in the same study. Tests built on polished newsroom prose, TOEFL scripts, application essays or synthetic benchmark text are measuring different things and will not agree.
Different thresholds and operating points. Turnitin's headline rate counts only documents over its 20% threshold, and scores under that threshold are marked as less reliable. RAID reports recall at a fixed false positive rate, so a figure like 95.7% recall at 1% FPR describes one chosen trade-off, not the tool's behaviour at other settings. GPTZero's error rate is under 1% if you restrict to its high-confidence predictions, and much higher if you do not. Comparing numbers taken at different operating points is comparing different machines.
Vendor incentives, in both directions. Detector vendors benchmark on corpora they select and rarely publish enough to reproduce the result; the RAID team found commercial claims of 99%+ accuracy unsupported by sufficiently challenging benchmarks. On the other side, humaniser and paraphrasing vendors publish tests designed to make detectors look as broken as possible, because that is their sales pitch. Independent and academic tests sit between the two, which is why the funder column exists.
Modified text collapses the numbers. Weber-Wulff's team measured accuracy of 96% on plain human text and 74% on unmodified AI text, falling to 42% when a human lightly edited the AI text and 26% after machine paraphrasing. They also found roughly 20% of AI-generated text was misattributed to humans overall, and that tools lean towards calling text human when unsure. A 2025 arXiv study, reported second-hand via fast.io's aggregation, found targeted adversarial paraphrasing cut detection rates by an average of 87.88%. Any test corpus containing edited or paraphrased AI text will therefore report far lower accuracy than one using raw model output.
Models and detectors both move. Walters and Elkhatat both found detectors markedly better at catching GPT-3.5 output than GPT-4 output, and newer models write more human-like text still. Pangram's claimed false positive rate changed by two orders of magnitude across its own versions. Every number on this page is dated to its study year and should be read that way.
What a detector score can and cannot support
What the published record supports using a score for: population-level measurement and triage. The 2026 newspaper study that flagged about 9% of articles as partly or fully AI-generated is a legitimate use, and its own lead researcher drew the line precisely: results at that scale can reveal trends, but not the guilt of any one author. A score can also reasonably prompt a closer human look at a document, or a conversation. Turnitin frames its tool this way itself, and GPTZero advises treating results as a conversation starter and not a final verdict.
What the record does not support: treating a score as sole evidence for an accusation. Turnitin states plainly that it does not make a determination of misconduct and that educators must apply their own judgment, because its false positive rate is not zero. OpenAI withdrew its own classifier in July 2023 over low accuracy, at a measured 26% detection rate and 9% false positive rate. Weber-Wulff and colleagues computed the false-accusation risk directly and found six of fourteen tools produced false positives, with half of GPT Zero's positive classifications amounting to false accusations in their test. Academic-integrity researcher Mike Perkins put the consensus to Nature in 2026: detectors can work in controlled tests, but the false positive concerns mean they should not be used for anything sensitive for a student.
The honest bottom line: a detector score is a probability estimate from an unaudited model, measured on someone else's corpus at someone else's threshold. In aggregate, across thousands of documents, these tools measure something real. On any single document, in either direction, the published numbers say the score is weak evidence, weaker still if the writer is a non-native speaker or writes in a plain, formulaic style, and weakest of all on text that has been edited or paraphrased after generation.
Methodology
Independence: ScriptGrain sells neither AI detection nor tools for evading detection, and has no commercial stake in whether detectors look accurate or broken. This page exists because most pages ranking for this query are published by companies on one side of that market or the other.
Every number on this page was taken from a source we fetched and read on 3 August 2026, listed below. Primary sources (the paper itself or the vendor's own page) were used wherever possible; the few figures we could only obtain second-hand are labelled as such in the text.
Widely quoted figures that could not be traced to a primary source, such as the 98% accuracy commonly attributed to Turnitin's FAQ, are flagged as unverified rather than repeated as fact.
Vendor self-reported numbers are included because the reference would be incomplete without them, and are always labelled with the vendor's name in the table's final column.
This page reports what published tests found, including how accuracy degrades under paraphrasing. It contains no guidance on evading detection, and ScriptGrain does not build or endorse tools for doing so.
Last checked: 3 August 2026. Detector versions change quickly; every figure should be read as dated to its study year.
Sources
- Turnitin: Understanding false positives within our AI writing detection capabilities (March 2023)
- Turnitin: AI writing detection update from Turnitin's Chief Product Officer (May 2023)
- Turnitin: AI detector product page (false positive rate statement)
- Liang, Yuksekgonul, Mao, Wu and Zou: GPT detectors are biased against non-native English writers (Patterns, 2023)
- Weber-Wulff et al.: Testing of detection tools for AI-generated text (International Journal for Educational Integrity, 2023)
- Nature careers feature: Universities are relying on AI-detection software to catch cheating. How well do the programs work? (July 2026; source for Walters 2023, Dik et al. 2025, the ZeroGPT spot test, the newspaper study and the Perkins and Karpinska quotes)
- Bloomberg Businessweek: AI Detectors Falsely Accuse Students of Cheating (October 2024; excerpt consulted)
- Ars Technica: OpenAI discontinues its AI writing detector due to low rate of accuracy (July 2023; source for OpenAI's 26% and 9% figures)
- Vanderbilt University: Guidance on AI detection and why we're disabling Turnitin's AI detector (August 2023)
- Dugan et al.: RAID, a shared benchmark for robust evaluation of machine-generated text detectors (ACL 2024)
- Dik, Erdem and Dik: preprint on GPTZero reliability (2025; figures via Nature)
- Russell, Karpinska et al.: preprint on AI-generated text in US newspapers (2026; figures via Nature)
- GPTZero: technology and accuracy page (vendor self-report)
- Pangram: Why perplexity and burstiness fail to detect AI (vendor self-report)
- detectiondrama.com: Turnitin AI detection statistics (aggregator; used as a pointer, flagged where not verified against a primary)
- fast.io: AI detector accuracy comparison 2026 (independent aggregation of Scribbr, RAID, ProofreaderPro and Axis Intelligence results)
ScriptGrain · Why ScriptGrain · Sample voice profiles · The Journal · Support