AI-isms in marketing copy: 30 phrases and 10 sentence shapes, counted per 1,000 words

By Jack Stovell · published 2026-09-21 · checked 2026-09-20

An AI-ism is a phrase or a sentence shape that models produce far more often than people do. ScriptGrain counts 30 phrases and 10 shapes per 1,000 words and publishes the whole list so anyone can run the count themselves. Not a vibe. A number.

What counts as an AI-ism

Here's the thing: any single phrase on this list can turn up in perfectly normal human writing. That's not the point. The point is frequency. Models over-produce certain phrases and certain sentence shapes at rates human writers don't come close to, and once you're counting per 1,000 words across a few hundred words of copy, the gap shows up reliably. A generic transition like "moreover" isn't a crime. Ten AI-isms in 400 words is a pattern.

The list groups into four phrase categories and ten shapes. Generic transitions are the connective tissue models reach for when they want to sound structured without doing structural work; "furthermore" and "in conclusion" do a job real writers usually do with a full stop and a new paragraph. Nominalised abstract metaphors ("architecture of", "the very texture of") are the models trying to sound literary by turning verbs into ornaments. Hedged abstract closes exist because models are trained to avoid committing to a claim at the exact moment a human writer would just land one. And the casual markers, "I mean", "sort of", "kind of", only count when they recur, because used once they're just how people talk.

The 30 phrases

The full list lives at /tools/ai-cliche-checker, pulled straight from the same detector this page runs on. Four groups, thirty phrases.

Generic transitions: furthermore, moreover, in conclusion, it's worth noting, delve into, in today's world, in an age where, thus it was that, it was a condition of, it is important to note. Ten phrases, and you've read most of them in a LinkedIn post this week.

Nominalised abstract metaphors: architecture of, geography of, constellation of, apparatus, mechanism, the very texture of, possessed of, possessed what could only be described as, in the glow of, the practiced resignation of, conspiratorial. Eleven phrases, and they're the ones that sound like the model is trying too hard to sound like a novelist.

Hedged abstract closes: perhaps enough after all, its own form of grace, something approaching, what could only be described as. Four phrases, and they show up almost exclusively at the end of a paragraph, right when a human writer would just say what they mean.

Systematic casual markers: I mean, you know, right?, sort of, kind of. Five phrases, counted only on recurrence, because context matters here more than anywhere else on the list.

The table below shows each group against its count and gives one example per group so you can see the shape of the thing without downloading the source file.

GroupPhrases
Generic transitions (10)furthermore; moreover; in conclusion; it's worth noting; delve into; in today's world; in an age where; thus it was that; it was a condition of; it is important to note
Nominalised abstract metaphors (11)architecture of; geography of; constellation of; apparatus; mechanism; the very texture of; possessed of; possessed what could only be described as; in the glow of; the practiced resignation of; conspiratorial
Hedged abstract closes (4)perhaps enough after all; its own form of grace; something approaching; what could only be described as
Systematic casual markers (5, recurring only)I mean; you know; right?; sort of; kind of

The 10 sentence shapes

Phrases are the easy part. Sentence shapes are where most detectors give up, and where most AI-isms actually live.

The headline shape is the balancing construction that swaps two things against each other in a single breath. It's the model's favourite move, and it's why ScriptGrain's own scanner exists at all (more on that below). Alongside it sit the hedging pair, the contrast construction, and the closing pivot, three variations on the same trick: set up an expectation, then subvert it, rather than just stating the thing.

Then there's the structural stuff. Triadic parallel lists (three items, identical grammar, marching in a row) sound impressive and say almost nothing. Semicolon chains stack three or more clauses into one sentence because the model has run out of full stops it trusts. Em dashes get their own line entirely, reported separately from density, because they're less a phrase choice and more a nervous tic.

The last three are rhythm meters, not phrase hits: uniform sentence length (every sentence in a paragraph sitting inside the same 20 to 40 word band), no short sentences (a paragraph of three or more sentences with nothing under eight words), and no contractions across 300 words of otherwise contemporary prose. None of these are single sentences you can point to. They're patterns across a paragraph, and they're often the most reliable tell of all, because they're the hardest for a writer, human or model, to fake once they know they're being watched. There's an eleventh shape, rhetorical question then answer, that's catalogued but not detected; it's common enough in human writing that counting it would just add noise.

ShapeCounts toward density
"Not X but Y" balancing constructionsyes
"Perhaps X, perhaps Y" hedging pairsyes
"Less X, more Y" contrastsyes
"It's not that X, it's that Y" closersyes
Triadic parallel listsyes
Semicolon chains (three or more in a sentence)yes
Em dashesreported separately
Uniform sentence length (one 20 to 40 word band)reported separately
No sentence under eight words in a paragraphreported separately
No contractions in 300 wordsreported separately

Density, and the bands

Density is phrase hits plus the first six sentence shapes, per 1,000 words. Em dashes and the three rhythm meters get reported on their own; they don't feed the density number, because they measure something closer to texture than vocabulary.

Three bands: clean is under 1 per 1,000, some is 1 to 3, heavy is over 3. That's it. No fourth category, no nuance beyond those cut-offs, because more granularity than that just invites people to argue with the number instead of using it.

What human writing scores

In the reference corpus of 299 human pieces, the median density is 0.4 per 1,000 words. The 90th percentile is 1.6. The 95th is 2.4.

So a piece with a handful of AI-isms in it isn't damning. It's normal. Even fairly clean human writing drifts into the "some" band occasionally, and a small number of genuinely human pieces sit above 2 per 1,000 without anyone reading them and thinking "a machine wrote this." The bands describe likelihood, not verdicts.

Measure10th percentile25thMedian75th90th
AI-tell density per 1,000 words (299 human pieces)000.41.01.6

What AI drafts score

Measured in SGR-002 (/research/chatgpt-vs-a-measured-voice, 2026-09-20): bare GPT-5 scored 0.2 tells per 1,000 words on this exact list. Genuinely clean, on phrases and shapes alone. And yet it produced 17.5 em dashes per draft and ran 65% longer than asked. Clean on density, loud on everything density doesn't measure.

Here's the more uncomfortable number: ScriptGrain's own drafts scored 2.3 tells per 1,000, almost all of it "not X but Y" constructions, despite the generator explicitly banning that construction at the prompt level. The ban didn't hold. That's why the scanner now runs on every single draft rather than trusting the ban to do the job upfront; a rule you tell a model to follow and a rule you actually verify are two different things, and only one of them shows up in a density count.

Limits

A list is a list. It catches phrases and shapes, not meaning, and a genuinely good human writer can use "moreover" and mean every word of it. Context beats pattern-matching every time you zoom in close enough on a single sentence.

Density needs volume to mean anything. Run this on a 60-word product blurb and you'll get noise, not signal; a few hundred words minimum, ideally more, before the number settles into something worth reading.

And the 95th percentile of human writing sits at 2.4 per 1,000. So "some AI-isms" in a piece of copy is not a verdict on who or what wrote it. It's a number that needs the bands, the sample size, and a bit of judgement sitting next to it, not instead of it. For a broader read on how this fits together, see /reference/markers-of-ai-writing; for the practical side, /tools/ai-slop-checker and /humanize-ai-writing cover checking and fixing respectively.

Questions

What's the difference between an AI-ism and just bad writing?

Bad writing is subjective; AI-isms are counted. A clumsy sentence might just be a clumsy sentence written by a tired human. An AI-ism is specifically a phrase or shape that models produce at rates far above the human baseline, which is why density matters more than any single example. One instance proves nothing. A cluster is data.

Can a human writer trigger a "heavy" density score?

Yes, and it happens. Academic or corporate writing leans on "moreover" and semicolon chains naturally, and can drift into the "some" or even "heavy" band without a model anywhere near it. That's exactly why the bands describe likelihood rather than certainty; the 95th percentile of human writing already sits at 2.4 per 1,000, close to the "heavy" threshold.

Why do em dashes get counted separately from density?

Because they're a rhythm habit, not a vocabulary choice. Bare GPT-5 produced 17.5 em dashes per draft in SGR-002 while scoring a clean 0.2 tells per 1,000 on phrases and shapes. Folding em dashes into density would have hidden that gap entirely, so they're reported on their own line, alongside the three rhythm meters.

Does banning a phrase in the prompt actually stop it appearing?

Not reliably. ScriptGrain's generator explicitly bans "not X but Y", and its own drafts still scored 2.3 tells per 1,000, almost all of them that exact construction. That's the whole reason the scanner runs on every draft rather than trusting the ban. A rule stated once at generation time is not the same thing as a rule enforced afterwards.

How many words do I need before a density score means anything?

More than a paragraph. The count is phrase hits plus sentence shapes per 1,000 words, and on very short copy a single em dash or one "moreover" swings the number wildly. A few hundred words is a reasonable floor; more gives a steadier read, and this is exactly why the reference corpus behind the human baseline runs to 299 full pieces rather than a handful of tweets.

Methodology

The list is ScriptGrain's cliché checker data (version 2026-09-20), published in full; human percentiles come from the reference corpus of 299 public pieces measured 2026-09-20; AI figures from SGR-002 as published. The checker is ScriptGrain's own instrument and another list would count differently.

Sources