AI slop checker: score a passage on the tells, no rewriting

By Jack Stovell · published 2026-09-21 · updated 2026-09-20

Slop is countable. It's listed phrases and sentence shapes per 1,000 words, nothing more mystical than that, and the checker below counts them against what 299 human pieces score on the same measure. Paste your draft and get a number, not a vibe.

Drop your text into the box below. It'll scan for the phrase list and the sentence shapes, tally the density, and tell you where you land against the human reference corpus. No edits, no suggestions, no rewrite. Just the count.

What slop is, counted

Here's the thing: "AI writing sounds off" is not a useful sentence. It's a feeling dressed up as a diagnosis. So we counted instead.

There are 30 tell phrases, split into four groups. Generic transitions like "furthermore" and "it's worth noting". Nominalised abstract metaphors: "architecture of", "apparatus", "the practiced resignation of", eleven of these in total. Hedged abstract closes such as "something approaching" or "its own form of grace". And five casual markers ("I mean", "sort of", "right?") that only count when they recur, because everyone says "kind of" once by accident.

Then there are ten sentence shapes, the ones your eye slides past but your brain still clocks. Balancing constructions with X and Y on either side of a fulcrum. Hedging pairs that refuse to land anywhere. Triadic lists where three items share suspiciously identical grammar. Semicolon chains. Em dashes. And three rhythm meters we'll get to in a second.

Density is phrase hits plus the first six sentence shapes, per 1,000 words. Em dashes and the rhythm meters get reported separately; they don't count toward density, because they measure something different: not vocabulary, but pacing. Clean is under 1 per 1,000. Some is 1 to 3. Heavy is anything past 3. The full phrase list lives at /tools/ai-cliche-checker if you want to see all thirty laid out.

The shapes matter more than the words

Anyone can ban a phrase. Banning a shape is harder, which is exactly why shapes are the better tell.

Take "not X but Y". Something like: "This isn't about speed, but about care." Grammatically fine. Structurally, it's a tic, because real writers rarely need the correction; they just say the thing.

Or the triadic list, three items marching in identical grammar: "a faster workflow, a cleaner draft, a happier client." Neat. Also a giveaway, because natural speech rarely balances three items that tidily unless someone's performing tidiness.

Then the rhythm meters, which are structural rather than lexical. Uniform sentence length: every sentence in a paragraph sitting inside the same 20 to 40 word band, so the whole paragraph hums at one pitch. No short sentences: three or more sentences with nothing under eight words anywhere. And no contractions across 300 words of otherwise contemporary prose, which reads like someone translating themselves into a second, more formal, language.

None of these need a banned word. That's why phrase lists alone catch less than half the problem. Full background on this split lives at /reference/ai-isms-in-marketing-copy.

And the data backs this up harder than expected. In SGR-002, bare GPT-5 produced just 0.2 tells per 1,000 words on the phrase list; genuinely clean by that measure. But it also produced 17.5 em dashes per draft and ran 65% over the requested length. Clean words, loud rhythm. Meanwhile ScriptGrain's own generator, despite explicitly banning "not X but Y", still produced 2.3 tells per 1,000, almost all that exact construction. Bans on words don't hold. That's why the scanner has to run on every draft rather than trusting the rule that was meant to prevent the problem in the first place.

What a clean result does and does not mean

A clean score is not proof of a human. Say that twice if you need to.

It means the passage sits inside the range human writing typically occupies. That's all. In the reference corpus of 299 human pieces, the median density is 0.4 per 1,000 words, the 90th percentile is 1.6, and the 95th percentile is 2.4. So even genuinely human, unedited prose sometimes reads "heavy" by this measure. A single dense paragraph, one clumsy transition, and you're past the median without having typed a word of AI output.

The checker measures tells, not authorship. Treat a low score as "this doesn't trip the obvious wires", not "this passed an audit". For actual detection odds rather than style tells, that's a different job, covered at /ai-detect.

Fixing it without flattening it

Fixing slop is not the same job as disguising it.

The actual edits are specific: break up sentences sitting in the same length band, so the paragraph has some snap to it. Add contractions where the voice would naturally have them. Cut the balancing constructions and just state the claim. Drop one item from every triadic list, because two items said plainly beats three said symmetrically. None of this requires a tone transplant, just attention to the shapes above.

The alternative, running everything through a generic "humaniser", tends to swap one set of tics for another; it doesn't give the writing a voice, it just moves the seams. A rewrite toward a measured, specific voice does more, and it's the difference SGR-001 found: on the same banking brief, average sentence length dropped from about 17 words to about 11, sentences under eight words doubled from 24% to 48%, contractions rose by roughly 70%, and ten em dashes became none. One run, unedited, but the direction is the point. More on building that voice deliberately at /humanize-ai-writing, and the fuller study behind these numbers at /research/chatgpt-vs-a-measured-voice.

Questions

Does a zero score mean the text is human?

No. It means the passage didn't trip the phrase or shape detectors, which is a narrower claim than authorship. Human writing itself sits at a median of 0.4 tells per 1,000 words, and even the 95th percentile only reaches 2.4. Clean just means unremarkable by this specific measure, not verified.

Why do shapes matter more than banned words?

Because words are easy to swap and shapes aren't. A generator can avoid saying "delve into" while still writing uniform 20 to 40 word sentences with no contractions, and that rhythm is what readers actually notice first. SGR-002 found GPT-5 scored 0.2 tells per 1,000 words yet still produced 17.5 em dashes per draft.

Can banning phrases in a prompt actually work?

Not reliably. ScriptGrain's generator explicitly bans "not X but Y" and still produced 2.3 tells per 1,000 words in testing, almost all that exact construction. That's the reason the scanner runs on every draft rather than trusting the ban itself; instructions get ignored more often than expected.

What's the difference between this and the cliché checker?

Same scanner, same count. The cliché checker at [/tools/ai-cliche-checker](/tools/ai-cliche-checker) walks through the phrase list, the thirty terms across four groups; this page explains the sentence shapes and rhythm meters, things like uniform sentence length or missing contractions, which catch structural tells that a clean vocabulary can still hide behind.

Should I aim for a zero score everywhere?

Not necessarily. Median human writing sits at 0.4 tells per 1,000 words, not zero, so obsessively chasing zero can itself flatten a passage into something stiffer than natural prose. Aim for "clean" (under 1 per 1,000) as a sanity check, then focus effort on rhythm and voice rather than chasing an arbitrary floor.