How do AI detectors work? The four methods, and why each one misfires
By Jack Stovell X · Instagram · LinkedIn · published 2026-10-03 · checked 2026-10-02
AI detectors work in one of four ways: they measure how predictable the words are to a language model, compare a text with rewrites of itself or with a second model, run a classifier trained on labelled writing, or look for a hidden watermark. None of them can prove who wrote a text.
The short answer
"AI detector" is a label for four quite different machines. The first asks how surprising your word choices are to a language model. The second nudges the text, or runs it past two related models, and watches how the scores behave. The third is a classifier that has seen a large pile of labelled human and machine writing and learned to tell them apart. The fourth doesn't analyse style at all; it checks for a statistical signature planted when the text was generated.
Each family fails in its own way, and the table attached below sets out what each one measures and where it breaks. The pattern across all four is the same. Every method estimates a probability about a pattern in the words. None of them observes the writing happening.
That distinction matters more than any accuracy figure. So let's take the methods one at a time.
| Method | What it measures | Needs the original model? | Example | Known weakness |
|---|---|---|---|---|
| Predictability (perplexity, burstiness) | How likely a language model finds each word, and how much that varies | A language model to score with | GPTZero's original method (dropped as its core in autumn 2023) | Plain, formulaic human writing reads as predictable too |
| Probability curvature or two-model comparison | Whether the text sits at a peak of a model's probability, or how two models disagree on it | A language model to score with | DetectGPT (2023), Binoculars (2024) | Paraphrasing: DIPPER cut DetectGPT from 70.3% to 4.6% at a 1% false positive rate |
| Trained classifier | Patterns a neural network learned from labelled human and AI text | No | GPTZero, Turnitin and Pangram today | Only as good as its training data; scores vary by corpus and threshold |
| Watermark | A statistical signal the generating model hid in its word choices | No, but the generator must have added it | Kirchenbauer et al. (2023), SynthID-Text (2024) | Absent from any text whose generator added none |
Method one: how predictable the text is
This is the oldest idea, and the one most people mean when they say "AI checkers look at perplexity". GPTZero's own explainer defines perplexity as a measure of "how likely an AI model would have chosen the exact same set of words". If a model would have picked nearly every word you used, the text scores low. Low perplexity reads as machine-like.
Burstiness is the second half. It's how much perplexity varies across a document, and GPTZero says low burstiness points to AI. The logic runs like this. A language model tends to write at an even, probable pace. People wander. A human paragraph might hold one plain sentence, one odd turn of phrase, and a fragment. So variation was read as human, and flatness as machine.
To be fair, the reasoning was sensible for its moment. Early model output really was smoother than most people's writing.
But predictability isn't authorship. A careful writer producing plain, standard prose scores as predictable too. So does anyone writing a legal summary, a lab report or a recipe, where the expected word is the right word. The signal measures a property of the text, and that property has more than one cause.
GPTZero itself has moved on. Its explainer says that "as of autumn 2023, GPTZero no longer uses perplexity and burstiness" as its core method, "because we migrated to a deep-learning based architecture". That's worth knowing, because plenty of guides still describe the old approach as if it were current. The idea survives as a teaching example and a rough intuition. As a production method, at least at GPTZero, it has been replaced.
Method two: comparing against rewrites or a second model
The second family stops asking "is this probable?" and asks something sharper: how does the text behave when you disturb it?
DetectGPT, from Mitchell et al. (2023), starts from an observation. Machine-written text "tends to occupy negative curvature regions of the model's log probability function". In plain terms, a model's own output sits on a local peak. Nudge the wording slightly and the probability drops. Human text doesn't sit on that peak, so small rewrites move its score in either direction. The method compares a text's log probability with slightly rewritten copies of it. It needs no trained classifier. On fake news written by GPT-NeoX 20B, the authors report 0.95 AUROC, against 0.81 for the strongest zero-shot baseline.
Binoculars, from Hans et al. (2024), takes a different route. It scores text by "contrasting two closely related language models". The idea is that a single model's surprise is hard to interpret on its own, because some passages are hard for every model. Putting a second, related model alongside gives a reference point. In the authors' tests, it detects "over 90% of generated samples from ChatGPT (and other LLMs) at a false positive rate of 0.01%".
Those are strong figures, and they're the authors' own, from their own test sets. That's the catch with this whole family. Both were measured on text as the models produced it, and edited text is a different test (see below), so the numbers may not describe your situation.
Method three: classifiers trained on labelled writing
This is what the big commercial tools describe using now. Collect a large corpus of human writing and machine writing, label it, and train a neural network to separate the two. No clever theory of curvature. Just pattern recognition at scale.
Turnitin describes its approach in its FAQ. A submission is "segmented into overlapping sections", and each segment is "given a value between 0 and 1". Each sentence then takes the scores of the segments it sits in. The model is "based on the transformer architecture" and trained on Turnitin's own academic writing, which suits a tool built for coursework. A file needs "at least 300 words of prose text" before it will be assessed at all.
GPTZero's technology page describes "an end-to-end deep learning approach" with "a sentence-by-sentence classification model". It returns three verdicts: human, AI, or mixed. The same page carries the vendor's own line: "No AI detector is 100% accurate." Credit where it's due, that's an honest sentence to publish about your own product.
Pangram's technical report, by Emi and Spero (2024), describes "a transformer-based neural network trained to distinguish" model text from human text. Its distinctive step is "hard negative mining with synthetic mirrors". The report credits it with "orders of magnitude lower false positive rates on high-data domains such as reviews", which makes it a direct attack on the false-positive problem.
Every accuracy claim in this family is the vendor's own, measured on the vendor's own data, at the vendor's own threshold. Independent results sit on the accuracy page, and they tend to be less flattering than the marketing. A classifier is also only as good as the writing it was trained on. Text unlike that corpus, whether a new model's output or an unusual human style, is where it's weakest.
Method four: watermarks
The fourth method changes the question entirely. Instead of studying a finished text, it hides a signal in the text as it's written.
Kirchenbauer et al. (2023) described the green-token idea. Before each word is generated, the model selects "a randomized set of 'green' tokens" and softly promotes them during sampling. Over a long enough passage, a watermarked text contains more green tokens than chance would produce. A statistical test finds that excess without access to the model.
SynthID-Text, from Dathathri et al. (Nature 634, 818, 2024), is the production version: "a production-ready text watermarking scheme" that "modifies only the sampling procedure" and is detected "without using the underlying LLM".
Watermarking is the only approach here that isn't guessing from style. It's also the narrowest. A watermark only exists if the company running the model added one when the text was generated. Text from a model with no watermark has nothing to find, and no amount of detector cleverness changes that. A clean result tells you the signal is absent, which is a different thing from the text being human.
Why editing and paraphrasing break detection
Every method above assumes the text arrives roughly as the model produced it. Real text rarely does.
Krishna et al. (2023) built a paraphraser, DIPPER, to test exactly this. It "drops detection accuracy of DetectGPT from 70.3% to 4.6% (at a constant false positive rate of 1%)". It also evaded watermarking, GPTZero and OpenAI's classifier. Rewording scrambles the statistical fingerprint that each method reads.
The same authors proposed a retrieval defence: the model provider keeps what it generated and searches that archive. It "can detect 80% to 97% of paraphrased generations" at a 1% false positive rate. It matches "semantically-similar generations" rather than exact wording, and it must be maintained by the provider of the model, which has to keep everything it generates.
Independent testing shows the same slope. Weber-Wulff et al. (2023), covering 14 tools, found 96% accuracy on plain human text and 74% on unmodified AI text. On AI text lightly edited by a human, it fell to 42%. On AI text run through a paraphraser, 26%.
Now the other side of the problem. Plain human writing shows the same habits that detectors treat as suspicious. Liang et al., at Stanford (Patterns, 2023), ran seven detectors over 91 TOEFL essays by non-native writers. On average they flagged those essays as AI 61.3% of the time, against 5.2% for essays by US students. The study's authors traced the gap to vocabulary and sentence-pattern richness: the constrained style second-language writers are taught reads as predictable. There's more on this in the post on non-native English writers.
So the errors run both ways. Edited machine text slips through, and careful human text gets caught.
How ScriptGrain's AI-writing check works
ScriptGrain's check belongs to none of the four families, and it's worth saying exactly what it does.
It's free with an account, and it's the AI-writing check, also reachable at POST /v1/ai-detect. A small model (Claude Haiku) reads the text against a written list of AI and human tells. It returns a human score from 0 to 1, along with the signals behind that score. Counts measured in code (sentence lengths and their spread, contractions, dashes, semicolons) go in alongside the text, so the model has the measured figures in front of it.
Then code checks the model's homework. Any signal that quotes a phrase not in the text gets dropped. So does any signal a count disproves. Same text, same score. The check needs at least 50 words.
What it doesn't do matters as much. It doesn't compute perplexity. It isn't a classifier trained on labelled essays. It doesn't look for a watermark. Its output is indicative, never proof, and the signals are there so you can read the reasoning and disagree with it. ScriptGrain does not sell detector evasion.
What a detector score can and cannot tell you
A detector score is a probability from a model, trained on someone's corpus, read at someone's threshold. That sentence contains three places for the answer to change. A different corpus gives a different score. A different threshold turns the same score into a different verdict.
For triage and trends, that's fine. If an editor is sorting a hundred submissions, or a department wants to see whether flagged rates are moving over a term, a noisy signal still helps. As evidence about one document, it's weak. It can't tell you who typed the words, how many drafts came before, or whether a tool helped at some step.
The clearest case for caution is OpenAI's own classifier. The company withdrew it, and the notice says: "As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy." It caught 26% of AI text and wrongly flagged 9% of human text. It was also "very unreliable on short texts (below 1,000 characters)".
That doesn't make detectors useless. It makes them one input among several. If you're weighing what to do about a flagged text, the position on detector evasion sets out where ScriptGrain stands on the wider question.
Limits
Commercial detectors do not publish their models, so the descriptions of GPTZero, Turnitin and Pangram are what each vendor says about its own system. Nobody outside can confirm them, and each can change its model without notice.
The research methods are described as their papers report them, on the models and data those papers tested. Detection results on 2023 models do not carry over unchanged to newer ones, in either direction.
This page explains how detectors work. It does not tell anyone how to make AI text pass as human, and it is not evidence for or against any particular accusation.
Questions
How do AI detectors work in one line?
They estimate the probability that a text was machine-written, using one of four methods: word predictability, comparison with rewrites or a second model, a classifier trained on labelled examples, or a watermark. Each produces a likelihood, not a finding about who wrote the text or how it was produced.
Do AI detectors use perplexity?
Some do. Perplexity measures how likely a model would have chosen your exact words, and low values were read as an AI sign. But GPTZero says that as of autumn 2023 it no longer uses perplexity and burstiness as its core method, having moved to a deep-learning architecture. GPTZero, Turnitin and Pangram all describe trained classifiers.
Does editing remove a watermark?
Heavy rewording can: Krishna et al. report that their paraphraser, DIPPER, also evaded watermarking. Light edits were not measured in the sources this page uses, so treat any claim either way with caution. And a watermark only exists if the company running the model added one when the text was generated; text from a model without one has nothing to find.
Why does human writing get flagged?
Detectors read statistical regularity, and plain human prose can look regular. In Liang et al.'s Stanford study, seven detectors flagged 91 TOEFL essays by non-native writers as AI 61.3% of the time on average, against 5.2% for essays by US students. Simple, standard phrasing resembles machine output to a probability model.
Is ScriptGrain's check a detector?
Not in the usual sense. It has a small model read the text against a written list of AI and human tells, with sentence counts measured in code, then drops any signal the text doesn't support. It uses no perplexity, no classifier trained on labelled essays and no watermark check. Its score is indicative, never proof.
Methodology
Every mechanism on this page is described from its primary source, read on 2026-10-02: the research papers' own abstracts (DetectGPT, Binoculars, the Kirchenbauer watermark, SynthID-Text in Nature, the DIPPER paraphrasing study, Pangram's technical report) and each vendor's own description of its product (GPTZero's technology page and its perplexity explainer, Turnitin's AI writing detection FAQ, OpenAI's notice withdrawing its classifier). Copies of the pages read are kept with the wave's sources.
Accuracy figures are not re-derived here: they come from the sourced table on /reference/ai-detector-accuracy, each with its study, corpus and funder, and the dates its sources were read.
ScriptGrain's own AI-writing check is described from its code (the api-v1 ai-detect handler and the detector guard), not from marketing copy. ScriptGrain sells neither AI detection on its own nor any tool for evading it.
Sources
- GPTZero: Perplexity and burstiness, what is it? (vendor explainer, read 2026-10-02)
- GPTZero: technology page (vendor self-description, read 2026-10-02)
- Mitchell et al. (2023), DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature, arXiv 2301.11305
- Hans et al. (2024), Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text, arXiv 2401.12070
- Turnitin: AI writing detection capabilities FAQ (vendor guide, read 2026-10-02)
- Emi and Spero (2024), Technical Report on the Pangram AI-Generated Text Classifier, arXiv 2402.14873
- Kirchenbauer et al. (2023), A Watermark for Large Language Models, arXiv 2301.10226
- Dathathri et al. (2024), Scalable watermarking for identifying large language model outputs, Nature 634, 818 (doi)
- Krishna et al. (2023), Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense, arXiv 2303.13408
- OpenAI: New AI classifier for indicating AI-written text (with the July 2023 withdrawal notice, read 2026-10-02)
- ScriptGrain: AI detector accuracy, every published number