Punctuation and formatting as a fingerprint: dashes, commas, list habits

By Jack Stovell · published 2026-09-24 · checked 2026-09-20

Punctuation is the layer of your writing you think about least, which is exactly why it gives you away. Nobody plans their comma placement. So when a piece of text has zero semicolons, no ellipses, and one suspiciously tidy em dash every few sentences, that's not a style choice. That's a tell.

The punctuation and format attributes

Eight measures do the work here, and two of them aren't strictly punctuation but belong in this room anyway. Comma density (per sentence), exclamation rate (per 1,000 words), semicolon frequency (per 1,000), and dash frequency (per 1,000) are the core marks. Ellipsis usage gets banded rather than scored as a raw number: never at zero, rare under 0.8 per 1,000, occasional under 2.5, frequent above that. Then there's question_mark_in_body, a simple yes or no, and capitalisation_quirks, which is descriptive text rather than a score. List preference rounds out the set: bullets, numbered, inline, mixed, or avoids entirely.

The scoring weights tell you what matters most. Comma density carries a weight of 1.5, with tolerance set at whichever is larger: 0.9, or 70% of your profile value. Exclamation rate is also weighted 1.5, tolerance 2.5 per 1,000 words. Semicolons match that 1.5 weight. Ellipsis habit gets 1, with a half-point still awarded if you're one band off rather than two. Questions in the body score 1, with only a small 0.2 penalty for a mismatch. Parentheticals get their own point too.

Here's the thing about rare habits: if you never use semicolons and the reference profile never uses them either, that counts as agreement. It just never becomes the headline. Nobody's fingerprint is "absence of semicolons". It's the whole hand.

What the reference corpus shows

The numbers behind all this come from 299 recent public pieces across 47 sources: personal blogs, engineering blogs, marketing blogs, newsletters, essays, magazine features, newspapers. That's 835,655 words in total, a median of 1,728 words per piece, built on 2026-09-20. Each piece was fetched, measured with the product's own code, and discarded. No text gets stored.

So what does typical human punctuation actually look like? Commas per sentence run p10 at 0.51 up to p90 at 1.40, with the median sitting at 0.92. That's a wide spread, which tells you comma habit varies enormously between writers who are all, by any reasonable definition, writing normally.

Exclamation marks are rarer than you'd guess: 0 at both p10 and p25, climbing to just 0.4 at the median, 1.8 at p75, and 3.9 at p90. A full quarter of human writing carries no exclamation marks and no semicolons at all. Semicolons follow a near-identical shape: 0, 0, 0.45, 1.6, 3.1. Restraint is the norm, not the exception.

Ellipses are rarer still. Zero at p10, p25, and the median. Only at p75 do they creep to 0.98 per 1,000 words, reaching 2.3 at p90. Put plainly: the median writer uses one ellipsis in a thousand words, or none at all. If your drafts sprinkle them every paragraph, that's not neutral. That's a signature.

Parentheticals sit at 0, 0.48, 1.1, 2.2, 3.6 per 300 words across the percentiles, so occasional asides are normal but constant bracketing isn't. And questions in the body turn out to be far more common than most people assume: at least three quarters of corpus pieces contain one somewhere in the running text, not just in headings.

Measure10th percentile25thMedian75th90th
Commas per sentence0.510.730.921.141.40
Exclamation marks per 1,000 words000.41.83.9
Semicolons per 1,000 words000.451.63.1
Ellipses per 1,000 words0000.982.3
Parentheticals per 300 words00.481.12.23.6

The em dash

The em dash deserves its own section because right now, it's the cleanest signal in the whole dataset. The cliché checker treats it as an AI tell and reports it separately, outside the normal dash density score.

SGR-002 put a number on it. Bare GPT-5, writing without a profile or any steering, produced 17.5 em dashes per draft on average. The profile arm and the prompt arm both produced none. Zero. SGR-001 found the same pattern from the other direction: drafts that started with ten em dashes dropped to none once a profile was applied.

That's not subtle. A machine writing freely reaches for the em dash constantly; a machine writing under constraint drops it entirely, and so does a profiled human voice. Which raises an obvious question: what if you genuinely, unironically love the em dash? Fair enough, some writers do. The answer isn't to ban it from your own voice. It's to make sure your profile actually records that you use them, at what rate, and in what pattern, so the mark reads as your habit rather than as evidence of an unedited draft. A dash used deliberately, at a measured rate, sitting inside a voice profile that says "yes, this one uses dashes", is a different animal to seventeen of them appearing from nowhere.

Lists, questions and capital letters

Formatting habits sit right next to punctuation, because they're just as unconscious. List preference gets banded into five types: bullets, numbered, inline, mixed, avoids. Some writers never reach for a list at all and thread everything into prose instead. Others number everything, even things that don't need numbering. Neither is wrong. Both are identifying.

Question marks inside body text are worth watching closely, given that three quarters of the corpus uses them. A voice that never asks the reader a direct question reads differently to one that does it every few paragraphs, and that's true regardless of formality or subject matter.

Capitalisation quirks are harder to quantify but easy to spot: ironic emphasis via quotation marks, unusual capitalisation of specific terms, consistent lowercase where you'd expect a capital. These get recorded as descriptive text rather than a number, because they resist banding, but they're still part of the fingerprint. You can read more about how these sit alongside the other 45 measures on the writing voice attributes reference page.

How the marks are counted

Every mark gets counted against the length of the text it comes from, which is why rates rather than raw counts. Exclamations, semicolons, ellipses and dashes are all measured per 1,000 words. Commas are measured per sentence, since sentence length varies so much between writers that a straight per-1,000 count would just track verbosity. Parentheticals get their own denominator too: per 300 words, which tracks better with how asides actually cluster in real paragraphs.

The counting method itself lives in the voice measurement framework, and the same pipeline that measures a submitted draft is the one that built the reference corpus in the first place. Consistency there matters more than most people assume.

Limits

Markdown strips things. Bold, italics, some list markers, occasionally a stray dash gets swallowed depending on how a document was exported. That means raw punctuation counts can undercount slightly for pieces that started life in a formatting-heavy editor and got flattened somewhere along the way.

Short texts cause a different problem. Per-1,000-word rates are stable across a 1,728-word median piece; they jump around wildly across 200 words. One ellipsis in a 200-word post reads as 5 per 1,000, which looks frequent, when really it's one mark that happened to land in a small sample. That's why the tools favour longer submissions where possible, and why a single short piece shouldn't be read as a full fingerprint on its own. You can run a sample through writing style analysis or check a draft against known AI patterns with the AI cliché checker to see how these limits play out in practice, and the broader comparison against unprofiled model output sits in ChatGPT vs a measured voice.

Questions

Why does punctuation matter more than word choice for identifying a writer?

Word choice is conscious; you can swap a word deliberately. Punctuation happens beneath that, in rhythm and habit you rarely notice yourself forming. That makes it harder to fake and easier to measure consistently. A writer can vary vocabulary between pieces far more easily than they can vary how often they reach for a semicolon.

Is using em dashes always a sign of AI writing?

No. SGR-002 showed bare GPT-5 producing 17.5 per draft against zero once a profile was applied, which is a strong signal, not an absolute rule. A human writer who genuinely favours the mark should keep using it, provided their profile records the habit at its real rate rather than leaving it unaccounted for.

What counts as a "rare" habit in this system?

Anything sitting near zero on both sides, such as a writer with no semicolons matched against a reference profile that also carries none. It scores as agreement. It just doesn't become the headline finding, because absence of a mark is common across roughly a quarter of the corpus for exclamations and semicolons alike.

Why measure commas per sentence instead of per 1,000 words?

Sentence length varies hugely between writers, from an average of 11.6 words at p10 to 23.4 at p90. A per-1,000-word comma count would mostly just reflect how long someone's sentences run, not how comma-heavy each one is. Per-sentence measurement isolates the actual habit.

Can a short piece of writing still be measured reliably?

Not as reliably as a longer one. Per-1,000 rates swing sharply on small samples: one ellipsis in 200 words looks frequent purely by arithmetic. The reference corpus itself runs on a 1,728-word median for this reason, and shorter submissions should be read with that limit in mind.

Methodology

Attribute definitions and scoring rules from ScriptGrain's profile schema and engine; corpus percentiles computed 2026-09-20 from 299 public pieces; em-dash counts from SGR-001 and SGR-002 as published.

Sources