The syntactic layer: sentence-structure metrics that separate two writers
By Jack Stovell · published 2026-09-24 · checked 2026-09-20
Two writers can post the same average sentence length, 17 words, and read nothing alike. One rides a wave of short punches and long run-ons; the other keeps everything in a tidy 14 to 20 word band. The difference lives in variance, in the share of short sentences, and in how clauses stack up before the point lands.
The seven syntactic attributes
Sentence structure is measured across seven attributes, part of a larger set of 45 covered in the full attribute reference. Average sentence length in words is the obvious one, and the least useful on its own. Sentence-length variance tells you how much that average is lying to you. Complexity preference (simple, compound, complex, or mixed) captures whether clauses pile up or stay separate. Clause ordering (front-loaded, build-to-point, or mixed) captures where the point sits inside the sentence.
Then three punctuation counts: parenthetical rate per 300 words, dash frequency per 1,000, and semicolon frequency per 1,000. None of these are exotic. Together they do more work than most writers expect.
Here's the front-loaded versus build-to-point distinction, illustrated rather than defined. Front-loaded: "The merger fell through because the board couldn't agree on valuation." Build-to-point: "The board couldn't agree on valuation, and after three months of back-and-forth, the merger fell through." Same information. Different shape. Different feel on the way in.
What the reference corpus shows
The reference corpus behind these numbers is 299 recent public pieces from 47 sources: personal blogs, engineering blogs, marketing blogs, newsletters, essays, magazine features, newspapers. 835,655 words in total, a median of 1,728 words per piece. Each one was fetched, measured with the product's own code, then discarded; nothing is stored. Full method is on the measurement framework page.
The percentile table (attached) runs average sentence length, variance, standard deviation over mean, share of short sentences, and the punctuation counts from p10 to p90. A few things jump out reading it. Average sentence length moves from 11.6 words at p10 to 23.4 at p90, a fairly ordinary spread. Variance moves from 47 to 260, more than fivefold. That gap between how much the average shifts and how much the variance shifts is the whole argument of this page.
Semicolons and parentheticals show the same widening pattern. At p10 both sit at zero. By p90, semicolons reach 3.1 per 1,000 words and parentheticals reach 3.6 per 300. So a chunk of writers use neither, ever, and that absence is itself a signal, not a gap in the data. Rare habits near zero on both sides count as agreement between two writers, though they never get headlined in a match report, because "neither of you uses semicolons" isn't a very interesting finding.
| Measure | 10th percentile | 25th | Median | 75th | 90th |
|---|---|---|---|---|---|
| Average sentence length (words) | 11.6 | 14.0 | 16.8 | 20.0 | 23.4 |
| Sentence-length variance | 47 | 69 | 102 | 146 | 260 |
| Variation (standard deviation over mean) | 0.50 | 0.54 | 0.61 | 0.70 | 0.78 |
| Sentences under 8 words, share | 8% | 13% | 20% | 30% | 40% |
| Parentheticals per 300 words | 0 | 0.48 | 1.1 | 2.2 | 3.6 |
| Semicolons per 1,000 words | 0 | 0 | 0.45 | 1.6 | 3.1 |
Rhythm: why variance matters more than average
Here's the thing about average sentence length: it flattens everything. A paragraph of six 17-word sentences and a paragraph with one 4-word fragment, two 30-word sentences, and three mid-length ones can both average out to 17. Read them side by side and nobody would mistake one for the other.
Variance is what the average hides. At p10, corpus variance sits at 47, which reads as tight and controlled: sentences cluster near the mean, rhythm feels smooth, almost metronomic. At p90, variance hits 260, and that's a writer swinging between clipped fragments and long, clause-heavy sentences in the same paragraph. Neither is wrong. They're different instruments.
Share of sentences under 8 words tells a related story. Corpus median is 20%, moving from 8% at p10 to 40% at p90. A writer at the low end almost never lands a short sentence; everything gets a full clause, a qualifier, room to breathe. A writer at the high end drops short sentences constantly, often right after a longer one, which is exactly the mechanism behind burstiness as a concept: length that varies rather than settles. For the underlying statistic, see sentence-length variance.
Clause ordering compounds this. A build-to-point writer delays the payoff, stacking context first. A front-loaded writer states the conclusion, then backfills. Mix build-to-point ordering with high variance and you get prose that keeps readers waiting, then rewards them with a short, flat sentence right when the tension peaks. That's rhythm. It's not decorative. It's structural.
How each metric is measured
Average sentence length and variance are counted directly from sentence boundaries; no modelling involved, just arithmetic across every sentence in a sample. Complexity preference is judged, not counted: a reader classifies whether clauses run simple, compound, complex, or mixed across a sample, which is why it carries a weight of 1 rather than the heavier weights given to length and variance. Clause ordering is judged the same way, also weight 1.
Semicolon frequency and parenthetical rate are straightforward counts per 1,000 and per 300 words respectively, each carrying real weight in a match score: semicolons at 1.5, parentheticals at 1. Sentence length carries the highest weight in a comparison, 2, with tolerance set at whichever is larger: 8 words, or 70% of the profile's own average. Variance carries weight 1.5, tolerance the larger of 25 or 1.2 times the profile's value. That tolerance design matters: a writer with a naturally high variance gets more room to vary before the system calls it a mismatch, and a writer near zero on semicolons or parentheticals still counts as matching another near-zero writer, even though that agreement never becomes a headline finding. You can run your own sample through the style analysis tool and see where it lands against these bands.
What AI drafts do to sentences
The measured pattern is consistent across two separate studies. In "Same Brief, Four Voices" (SGR-001, 2026-07-28), running a banking brief bare versus through a profile on the same model dropped average sentence length from about 17 words to about 11, and the share of sentences under eight words doubled, from 24% to 48%. Ten em dashes in the unprofiled draft became zero once a profile was applied. One run per arm, unedited.
In "ChatGPT vs a measured voice" (SGR-002, 2026-09-20), bare GPT-5 produced 17.5 em dashes per draft. When the same model was fed a one-page measured prompt asking for roughly 14 contractions per 1,000 words and a formality score of 3.7, it produced 0.5 contractions per 1,000 words and a judged formality of 7.2, while still obeying the countable rules: no em dashes, more short sentences. It followed the numbers it could count and ignored the tone it couldn't.
The cliché checker behind these studies flags two structural tells specifically: uniform sentence length, every sentence in a paragraph sitting inside one 20 to 40 word band, and paragraphs containing no sentence under eight words at all. Both are measurable. Both showed up repeatedly in unprofiled drafts across both studies.
Limits
SGR-001 ran one brief per arm, unedited, across three briefs total. SGR-002 used one author, two held-out posts, one run per arm, and wasn't blind. Detector human scores in SGR-002 didn't track voice match at all: the prompt arm scored 0.78 on a detector while matching the profile worst, at 0.62. That's a warning against treating any single metric, including the ones on this page, as the whole story.
Questions
What is syntactic stylometry?
It's the study of sentence-level structure as a fingerprint: length, variance, clause order, and punctuation habits, rather than word choice or topic. Two writers using the same vocabulary can still be told apart by how their sentences are shaped and how much that shape varies from one sentence to the next across a piece.
Why does sentence-length variance matter more than average length?
Because average length flattens rhythm. A paragraph of uniform 17-word sentences and one mixing 4-word fragments with 30-word sentences can share an identical average while reading completely differently. Variance captures that swing; average sentence length, on its own, cannot.
What counts as a front-loaded sentence versus build-to-point?
Front-loaded states the conclusion first, then supplies reasoning after. Build-to-point stacks context and clauses first, delaying the point until the end. Both are valid structures; a writer's habitual preference between them is one of the seven syntactic attributes measured here.
How is the reference corpus built?
It's drawn from 299 recent public pieces across 47 sources: blogs, newsletters, essays, magazine features, and newspapers, totalling 835,655 words. Each piece is fetched, measured with the product's own code, and then discarded; no source text is stored. Full methodology sits on the measurement framework page.
What do the published studies show about AI-drafted sentences?
Unprofiled AI drafts skew toward longer average sentence length, fewer short sentences, uniform length bands within paragraphs, and heavy em dash use, 17.5 per draft in one study. Applying a measured profile roughly halved average length and doubled the share of short sentences in the other, though sample sizes in both studies were small.
Methodology
Attribute definitions from ScriptGrain's profile schema; corpus percentiles computed on 2026-09-20 from 299 public pieces measured with the product's own code; study figures from SGR-001 and SGR-002 as published.