AI voice patterns in writing: the tells you can measure

By Jack Stovell · 2026-09-30 · Guides

AI writing has a recognisable voice, and it's made of patterns you can count: overused words, a small set of phrases and sentence shapes, even sentence rhythm, and a formal register. Each one also turns up in human writing. So the numbers and the human baseline matter more than any single example.

The short answer

Four layers make up the voice. Vocabulary is the first: a handful of words that models reach for far more often than people do. Phrases and sentence shapes come next, the stock transitions and balanced constructions that recur. Then rhythm: sentences of much the same length, paragraph after paragraph. Last, punctuation and register, meaning heavy em dash use and prose that stays stiffly formal, with few contractions.

Here's the thing. No layer proves anything alone. A habit is something a writer does; a tell is a habit that shows up at a rate humans rarely reach. The catalogue of 16 measurable markers sets out the evidence and the limits for each. What follows is the short tour, with the human baseline beside every pattern.

Words: the vocabulary that gives it away

Start with the most famous tell. Juzek and Ward, in "Why Does ChatGPT Delve So Much?" (COLING 2025), analysed 26.7 million PubMed abstracts, 5.2 billion tokens covering 1975 to May 2024. They identified 21 focal words whose use rose sharply after ChatGPT's release. The word "delves" rose 6,697% in PubMed abstracts from 2020 to 2024.

That's a big number. Read it carefully, though. The magnitude belongs to that corpus, biomedical abstracts, and it doesn't mean the word rose that much in blog posts or newsletters. Kobak et al. (Science Advances, 2025) measured the same kind of excess vocabulary in biomedical abstracts, so the pattern shows up in more than one study of that kind of text. Its reach beyond scientific writing is less certain than the headlines suggest.

Why does it happen? Juzek and Ward report failing to find evidence that the overuse comes from model architecture or from training or fine-tuning data. The focal words are far rarer in corpora such as arXiv and Wikipedia than in ChatGPT output. The hypothesis left standing, not proven, is reinforcement learning from human feedback: preference tuning may have taught models that this vocabulary is what a good answer sounds like.

One more thing, and it matters. Tells date. "Delve" was heavily overused in 2023 and early 2024, then dropped off sharply in 2025. A word list ages as models change, so any list you find today (including the AI-isms glossary entry) is a snapshot, not a verdict.

Phrases and sentence shapes

Words are the easy layer. Phrases and shapes are where counting gets useful.

ScriptGrain's tell list holds 30 phrases in four groups. Ten are generic transitions, the sort of filler that opens a paragraph and says nothing. Think of stock openers like the ones about noting what's important, or the ones that start with a sweeping claim about the age we live in. Eleven are nominalised abstract metaphors, the kind that dress up a plain point in borrowed scenery and abstract texture. Four are hedged abstract closes, the vague sort that trail off into a maybe instead of landing a point. The last five are systematic casual markers, counted only when they recur: "I mean", "you know", "right?", "sort of", "kind of". A single "you know" is a person talking. Ten in a page is a script.

Then the shapes. Ten are detected:

  1. Balancing constructions, "not X but Y", which set up a denial just to knock it down.
  2. Hedging pairs, "perhaps X, perhaps Y".
  3. Contrasts that lean on "less X, more Y".
  4. Closers that run "it's not that X, it's that Y".
  5. Triadic lists, three items in identical grammar.
  6. Semicolon chains, three or more in one sentence.

The other four are em dashes, uniform sentence length, paragraphs with no short sentence, and 300 words of prose without a contraction. An eleventh, the rhetorical question followed by its own answer, is catalogued but not detected.

Density is phrase hits plus the first six shapes, per 1,000 words. Under 1 counts as clean. From 1 to 3 is "some". Over 3 is heavy. Em dashes and the three rhythm meters are reported separately and stay out of the density figure.

Now the baseline, because it's the whole point. In the reference corpus of 299 human pieces, the median density is 0.4 per 1,000. The 90th percentile is 1.6, and the 95th is 2.4. So a human writer landing at 1.2 is unremarkable. A page at 5 is another matter. The AI-cliché checker counts all of this for free and publishes the full list.

Rhythm: when every sentence is the same length

This is the layer readers feel before they can name it. Paragraph after paragraph, sentences of 25 or 30 words, each one built like the last. Nothing snags.

Humans snag. In the reference corpus (299 pieces, 47 sources, 835,655 words, measured and discarded), the median average sentence length is 16.8 words, with a spread from 11.6 at the 10th percentile to 23.4 at the 90th. The more telling figure is variation: standard deviation over mean, median 0.61, running from 0.50 to 0.78. Short sentences, under 8 words, make up a median 20% of sentences, and the range runs from 8% to 40%.

Read that again. One sentence in five, in ordinary human prose, is a short one. That's the rhythm of people thinking on the page: a long, winding sentence that piles up qualifications, then a stop. Fragments too.

The rhythm meters look for the opposite. A paragraph is flagged if every sentence sits inside the same 20 to 40 word band. Another flag fires when a paragraph of three or more sentences has nothing under eight words. This property has a name, burstiness, and it's the one people mean when they say a text feels "flat". To be fair, plenty of careful human writers produce flat paragraphs now and then. It's the run of them that reads as machine.

Punctuation and register

Two habits sit here. Both are easy to measure and easy to over-read.

The first is the em dash, that long horizontal stroke people use to break into a sentence. It's the most talked-about tell of the lot. Some human writers lean on it heavily, so a few in a piece prove nothing. Volume is what counts.

Which brings us to the SGR-002 study, ChatGPT vs a measured voice, run on one author's posts. Bare GPT-5 averaged 17.5 em dashes per draft. It also wrote 2,309 words against a target of 1,393, about two-thirds more than asked. Its tell density, oddly, was low: 0.2 per 1,000 words. Few listed phrases, many dashes, a lot of padding.

The second habit is register. The prompt asked for about 14 contractions per 1,000 words and a formality of 3.7. GPT-5 produced 0.5 contractions per 1,000, and a judged formality of 7.2. Set that against the human baseline: contractions per 1,000 words run 8.1 at the 10th percentile, 19.5 at the median and 32.2 at the 90th. Half a contraction per 1,000 sits nowhere near the human range.

The limits are worth stating. Two briefs, one author, one run per arm, and not blind. It's a measured case, not a law.

There's a sting in the tail, too. ScriptGrain's own drafts produced 2.3 tells per 1,000, almost all of them that same denial-then-correction pattern, despite the generator banning it. That's why the scanner now has to run on every draft. Banning a pattern in the instructions doesn't mean the pattern stays away.

What the patterns cannot prove

They describe text. They don't say who wrote it.

Human writers show every pattern on this page. Someone with a fondness for the em dash, a formal register and a taste for triads will trip half the meters, and no machine was involved. Detectors built on these markers misfire for exactly that reason. Treat a high count as a prompt to look harder, never as a finding.

The useful comparison is different. The better question is "does this look like you?" A writer's own measured voice (contraction rate, sentence spread, where the questions fall, which words recur) gives you a baseline that a population median can't. Draft text that drifts from it is worth a second look, whoever or whatever produced it. That measured voice is what ScriptGrain builds. The population numbers above tell you what's typical; your own numbers tell you what's yours.

Questions

What are the signs of AI writing?

Four layers: overused words, stock phrases and balanced sentence shapes, even sentence rhythm, and a formal register with few contractions plus heavy em dash use. Each layer also appears in human writing, so look at rates. A human median is 0.4 tells per 1,000 words, and the 90th percentile is 1.6. Well above that, look closer.

Can you tell AI writing by the words it uses?

Partly. Juzek and Ward found 21 focal words that rose sharply after ChatGPT's release, but that was measured in PubMed abstracts, not everyday prose. Word tells also age: "delve" faded in 2025. A single word proves nothing, and a whole list is only a snapshot of one model generation.

Why does AI write sentences the same length?

The evidence stops at the symptom. Human sentence-length variation has a median of 0.61, and about 20% of human sentences run under eight words. The checker flags a paragraph whose sentences all sit in one 20 to 40 word band, with nothing short. Nobody has published a proven cause for the evenness itself.

Is using em dashes a sign of AI?

Not by itself. Plenty of human writers use them heavily. Volume is the clue: in SGR-002, bare GPT-5 averaged 17.5 em dashes per draft. That study covered one author and two briefs, so treat it as a measured case. The count matters more than any one dash, and a writer's own usual rate is the fair comparison.

How do I make AI writing sound less like AI?

Measure first. Count contractions, check sentence spread, and look for the listed phrases and shapes. Then edit toward your own baseline, not a generic ideal. Note that instructions alone don't hold: GPT-5 asked for about 14 contractions per 1,000 words produced 0.5. Scan every draft, and trust the count over the brief.

More from the ScriptGrain Journal