Voice preservation versus detector evasion: our position, stated

By Jack Stovell · published 2026-09-24 · updated 2026-09-20

We build tools that measure whether a piece of writing sounds like a specific person, and we optimize output toward that measurement, not toward beating a detector. Those are different jobs. We only do the first one.

Paste an AI draft next to something you actually wrote and run the check below. It scores voice preservation between two pieces, no profile needed, nothing stored. If a tool claims to preserve your voice, that claim should be measurable, not just plausible.

The position

Evasion optimizes text against a classifier. It edits until a detector says "human" and stops there. The output has to land somewhere; it just doesn't land on you. Preservation optimizes toward one measured writer: your sentence lengths, your contraction habits, your rhythm, checked against a profile built from your own material. Different target, different result. We build the second thing. We don't build the first.

That sounds like a small distinction until you see what it does to the writing. Evasion tools chase a moving target inside someone else's model. Preservation tools chase you.

Why the two goals pull apart

Here's the thing: passing a detector and sounding like you are not the same task, and SGR-002 proved it isn't even correlated in the direction you'd expect.

In that study, one author's five blog posts were used to build a profile, two posts held out as unseen briefs. Each brief ran three ways: bare GPT-5, GPT-5 with a one-page measured prompt, and ScriptGrain's generation against the profile. Against the study profile, ScriptGrain scored 0.83 on voice match, bare GPT-5 scored 0.74, GPT-5 with the prompt scored 0.62.

Then the detector scores came in backwards. Human-confidence went prompt arm 0.78, ScriptGrain 0.67, bare 0.52. So the run that read closest to the actual author scored lower on sounding human than the run that read least like them. A detector can't tell the difference between "this is a real person's voice" and "this text doesn't trip my filters." Those aren't the same signal, and treating them as interchangeable is exactly how evasion tools end up dragging every draft toward the same flattened middle: no one's voice, just an average built to slip past a gate.

The prompt arm is instructive too. It asked for roughly 14 contractions per 1,000 words and a formality score of 3.7. GPT-5 delivered 0.5 contractions per 1,000 words and a judged formality of 7.2, while dutifully obeying the countable rules: no em dashes, more short sentences. It followed the checklist and missed the voice. That's the whole problem in one paragraph.

What we measure and report

Every draft or pasted piece gets scored on two kinds of feature. Counted ones: sentence length and variance, contractions, commas, exclamations, semicolons, brackets, questions, ellipses, pronoun mix, article balance, word length, signature phrases. Judged ones: formality, humor, expressiveness, confidence, opening style, argument structure, rhythm, complexity, clause ordering, specificity. Weighted, calibrated, scored 0 to 1.

Against a real profile, 0.85 to 0.95 reads as voice match, 0.5 to 0.8 is partial, under 0.45 is off. Comparing two pieces with no profile (the free check), 0.80 and above is a match, 0.55 to 0.79 is drifting, under 0.55 is off. The free check needs 120+ words per piece and stores nothing.

Full method's at /reference/voice-measurement-framework. We also publish the tell list AI drafts keep repeating: 30 phrases across four groups, plus ten sentence shapes detectors and readers both flag once they know to look, at /tools/ai-cliche-checker. If you want the numbers behind detector reliability itself, those are at /reference/ai-detector-accuracy, and they cut both ways.

You can run your own draft through /ai-detect for a human-versus-AI read with a signal breakdown. It's free. It just isn't the product.

What we will not build

No bypass. No paraphraser tuned to slip past a classifier. No claim of undetectability, ever, on any page.

Several tools in this space sell "undetectable" as the feature. That's a different business to the one we're running. Our humaniser at /humanize-ai-writing targets natural variation and a measured voice; when a draft doesn't get there, it reports the miss rather than rounding up. That's a design choice, not a limitation we're apologizing for.

What this costs us, and why we accept it

This costs us traffic. People searching for "undetectable AI text" are not going to find that promise here, and some of them will leave for tools that make it. That's the trade, stated plainly.

We'd rather build something that measures voice preservation and tells you honestly when it fails than build something that scores well on a detector and reads like nobody in particular. SGR-002 is the argument for that, not just a preference: the arm that sounded most human to the classifier was the one that matched the author least. Chase the detector and you get that. Chase the voice and you get /preserve-your-voice, which is a narrower promise but a checkable one.

We'd rather lose the search term than win it on a claim we can't stand behind at /trust.

Questions

Is a high detector score the same as sounding human?

No. SGR-002 showed the opposite pattern: the run with the highest detector human-score (prompt arm, 0.78) had the lowest voice match against the study profile (0.62), while ScriptGrain scored lower on the detector (0.67) but highest on voice (0.83). They measure different things.

Can you guarantee my writing won't get flagged?

No, and we won't claim it. Our tools report a miss when preservation fails rather than rounding up. Detector accuracy varies and is documented at [/reference/ai-detector-accuracy](/reference/ai-detector-accuracy); we don't sell undetectability because it isn't a promise we can measure or stand behind.

What counts as a good voice match score?

Against your own profile: 0.85 to 0.95 is a match, 0.5 to 0.8 is partial, under 0.45 is off. Comparing two pieces with no profile, 0.80 and above counts as matched, 0.55 to 0.79 is drifting. Full scoring method is at [/reference/voice-measurement-framework](/reference/voice-measurement-framework).

Does the free check store what I paste in?

No. The free comparison on this page needs 120 or more words per piece and stores nothing. It compares two pieces directly rather than against a saved profile, which is why the thresholds differ slightly from profile-based scoring; the detector at [/ai-detect](/ai-detect) is a separate, free read.

Why do some AI drafts avoid em dashes but still sound artificial?

Because obeying countable rules isn't the same as matching a voice. In SGR-002, the prompted arm hit the mechanical targets, no em dashes, more short sentences, but still landed at 0.5 contractions per 1,000 words against a 14-per-1,000 target and a formality of 7.2 against a target of 3.7. Checklist compliance and voice match diverge easily.