The lexical layer: which vocabulary metrics identify a writer
By Jack Stovell · published 2026-09-24 · checked 2026-09-20
A writer's fingerprint isn't hiding in rare words. It's in the habits: how often they contract "do not" to "don't", whether their sentences lean on short words or long ones, and the small stock of phrases they reach for without noticing. Six lexical attributes measure exactly that.
The six lexical attributes
Here's the list, and each one earns its place.
Vocabulary diversity index runs from 0 to 1 and reports how much of a text's vocabulary isn't repeated. A writer who says "good" fifteen times scores low. One who reaches for "good", "solid", "sound" and "decent" across the same piece scores higher.
Preferred words are the recurring distinctive words a profile catalogues: not "the" or "and", but the specific nouns and adjectives a writer keeps returning to. Someone who writes about software might lean on "brittle" and "clean" far more than base rate predicts.
Word length distribution splits vocabulary into three bands: short (1 to 4 letters), medium (5 to 7), long (8 or more). A writer heavy on short words reads punchy. One who drifts long reads formal, sometimes laboured.
Rare word rate counts genuinely uncommon words per 1,000. It's a thin signal on its own, easy to fake by dropping in three obscure adjectives, but useful alongside the others.
Filler phrases get catalogued too: "honestly", "basically", things people say without registering them. They're small tics with big identifying power.
Contraction frequency counts contractions per 1,000 words. It's the single most-weighted lexical signal in scoring, and for good reason: it tracks formality more reliably than almost anything else on this list.
What the reference corpus shows
The reference corpus behind these numbers is built from 299 recent public pieces across 47 sources: personal blogs, engineering blogs, marketing blogs, newsletters, essays, magazine features, newspapers. 835,655 words total, median piece length 1,728 words. Each piece was fetched, measured, and discarded; nothing is stored.
On contractions, the corpus runs from 8.1 per 1,000 words at the 10th percentile up to 32.2 at the 90th, with the median sitting at 19.5. A writer at the 10th percentile sounds formal almost by default: "it is" rather than "it's", "cannot" rather than "can't", the register of a report rather than a conversation. A writer at the 90th percentile sounds like they're talking to you, because they mostly are: contractions everywhere, sentences that read like speech.
That gap matters more than it looks. Going from 8 contractions per 1,000 words to 32 doesn't just change word count, it changes how trustworthy the voice feels. Formal writers read as careful. Loose, contraction-heavy writers read as direct. Neither is wrong. They're just different corners of the same distribution.
Word length tells a similar story from the other direction. Short words (1 to 4 letters) make up 50% of a piece at the 10th percentile and 62% at the 90th, with 55% as the median. Long words (8+ letters) run from 11% to 20% across the same spread. A writer stacked at the short end reads fast and plain. One stacked at the long end reads considered, maybe a bit academic. Neither band is inherently better; it depends what the piece is for.
| Measure | 10th percentile | 25th | Median | 75th | 90th |
|---|---|---|---|---|---|
| Contractions per 1,000 words | 8.1 | 13.8 | 19.5 | 24.7 | 32.2 |
| Short words (1 to 4 letters), share | 50% | 53% | 55% | 59% | 62% |
| Long words (8+ letters), share | 11% | 13% | 16% | 18% | 20% |
How each metric is measured
Two of these are counted straight from the text in code; the other four (diversity, preferred words, rare words, fillers) are read by the extraction model from the samples and then checked for in drafts.
Contraction frequency runs off a regex: it scans for the apostrophe forms ("don't", "it's", "they're") and counts hits per 1,000 words. No interpretation, no borderline cases. It's a mechanical count.
Word length distribution works off letter-count bands. Every word gets sorted into short, medium or long by counting its letters, then the shares get tallied as percentages of total words. Simple, and because it's simple, it's hard to argue with.
Rare word rate is the extraction's estimate of genuinely uncommon words per 1,000, read from the samples rather than counted against a fixed list.
Signature phrases (preferred words, filler phrases, discourse markers) get checked against the profile's own catalogued list, built earlier from that writer's sample text. If the profile lists "here's the thing" and "fair enough" as recurring phrases, the checker looks for those specific strings in the draft. This one carries weight 2 in scoring, same as contractions, and it's scored as the share of up to four catalogued phrases actually used. Use three of the four and you've hit 75% on that component.
Word-length mix carries a lighter weight of 1. Contraction rate is scored with a tolerance: the larger of 10 per 1,000 words or 70% of the profile's rate. So a profile calling for 20 contractions per 1,000 gives you a tolerance band of 14, not the flat 10; a profile calling for 8 gives you the flat 10, because 70% of 8 is smaller. That tolerance exists because contraction rate is noisy at the sentence level even when a writer is being consistent overall.
Reading your own numbers
Run a sample through the free analyser and it reports these six against the corpus percentiles directly, not as abstract scores. You'll see where your contraction rate sits: 10th percentile, median, 90th, somewhere between. Same for word length shares, same for rare word rate.
That percentile framing matters. A raw number like "contractions: 14.2 per 1,000" tells you nothing on its own. Knowing that 14.2 sits between the 10th percentile (8.1) and the median (19.5) tells you something: this writer leans formal but isn't rigid about it. That's a usable fact.
The signature-phrase check runs against a profile, so it needs one: score a draft against your profile and the deltas say which of your catalogued words and phrases it used. If you're inconsistent, that's where it'll show: a piece with none of your usual phrases reads like someone else wrote it, even if the sentence rhythm matches.
For a fuller picture of how lexical measures sit alongside sentence-level and punctuation attributes, the writing voice attributes reference lays out all 45. And if you want the mechanics behind scoring and tolerance bands generally, the voice measurement framework covers that ground. Function words, which sit just outside this lexical layer but interact with it constantly, get their own entry in the glossary.
Limits
Three limits worth knowing before you trust these numbers too far.
Topic changes vocabulary. A piece about grief and a piece about tax software will produce wildly different rare word rates and preferred words, even from the same writer. That's topic, not voice drift, and it's easy to mistake one for the other.
Short samples make rates noisy. A 200-word sample can swing from 5 contractions per 1,000 to 40 depending on three sentences going one way or the other. The corpus percentiles above are built from pieces with a median length of 1,728 words; judging a 300-word sample against them needs a wide margin of error built in.
Diversity index depends on length too, and in a specific direction: longer pieces naturally show lower diversity, because a writer eventually repeats even their favourite words. Comparing a 3,000-word essay's diversity score against a 500-word post's isn't quite fair to either.
None of this makes the six attributes useless. It means read them as a cluster, against pieces of comparable length and topic, and don't treat any single number as a verdict.
Questions
What is vocabulary diversity index measuring?
It measures how much of a text's vocabulary is not repeated, scored 0 to 1. Higher scores mean a wider range of distinct words relative to total word count. It drops naturally as pieces get longer, since even varied writers repeat their most useful words eventually, so it's most meaningful when comparing pieces of similar length.
Why does contraction frequency carry more weight than rare word rate?
Contraction frequency tracks formality reliably and is cheap to fake convincingly wrong: a writer either contracts naturally or doesn't, and the pattern holds across a piece. Rare word rate is thinner, easy to nudge with a handful of unusual adjectives without changing the underlying voice. Scoring weights reflect that: contractions at weight 2, word-length mix at weight 1.
How are signature phrases actually checked?
The checker takes the preferred words, filler phrases and discourse markers catalogued in a writer's profile and searches the draft text for those specific strings. It's a literal match, not a judgement call. Scoring counts the share of up to four catalogued phrases found, so using three of four gives 75% on that component.
Can two very different writers share the same contraction rate?
Yes, easily. Contraction rate alone is one signal among six lexical attributes, and plenty more beyond that layer. A formal-sounding writer and a chatty one could both land near the corpus median of 19.5 per 1,000 words while differing sharply on word length, preferred words and filler phrases. That's why the metrics get read as a set.
Does a short sample ever give a trustworthy reading?
Not reliably. Rates like contraction frequency and rare word rate get noisy on small samples, since a handful of sentences can swing the count sharply either way. The reference corpus is built from pieces with a median of 1,728 words. Judging anything much shorter against those percentiles needs real caution built into the reading.
Methodology
Attribute definitions from ScriptGrain's profile schema; corpus percentiles computed on 2026-09-20 from 299 public pieces (835,655 words, 47 sources) measured with the product's own code and discarded, no text stored. Sources are the piece URLs' domains listed in the corpus file.