Word length analyser

By Jack Stovell · published 2026-09-24

Word length distribution is the share of short, medium and long words in a piece of writing. ScriptGrain's code bands them as short (1 to 4 letters), medium (5 to 7) and long (8 or more). A voice profile stores the three shares as the extraction model reads them. In ScriptGrain's 299-piece reference corpus, the median long-word share is 16%.

Analyse your word length distribution free

Your free voice profile measures word length distribution among all 45 attributes, from a sample of your own writing. No card needed.

Try the six-signal check first

At a glance

What word length distribution measures

Word length distribution measures how heavy or light your vocabulary reads, word by word. It sorts every word into one of three bands: short (1 to 4 letters), medium (5 to 7) and long (8 or more), then works out what share of your writing falls into each. Someone who writes "use" scores differently to someone who writes "utilise", even though both mean the same thing. That's the whole measurement. No judgement about which is better, just a count.

The bands matter because they're a proxy for something readers feel without naming it: effort. A page full of long words asks more of a reader than a page full of short ones. The attribute sits in the Lexical group alongside vocabulary diversity, rare word rate and preferred words; together they describe the raw material a writer reaches for. The Lexical group page covers all six, and the lexical layer reference has the group's corpus data.

Here's the bit worth being precise about. When ScriptGrain builds a voice profile, the value stored for word length distribution is Claude Sonnet 5's reading of your writing samples, generated across two passes: one per sample, one to synthesise the final profile. That's an extraction model's judgement call, not a raw count from code. But when a draft is scored against that profile, code counts the three shares directly from the draft, letters only, apostrophes stripped out. The free analyser counts only the long-word share, at 7+ letters. Two different processes, same attribute, worth knowing which one you're looking at.

How word length distribution is scored

Word length distribution is one of 28 attributes (out of 45) that can move your voice match score. It carries a weight of 1, and code counts it from the draft itself (measureTextFeatures); no model reads it. Both your profile's shares and the draft's shares get normalised, then similarity is worked out as 1 minus half the summed absolute difference between them.

For example, if your profile says you write 55% short words and a draft comes in at 60%, that's a small gap and the score barely moves. If it comes in at 35%, that's a large gap and it shows. This sits inside a bigger scoring system where code handles sentence, punctuation, pronoun, article and word-length features, and Claude Haiku separately judges things like formality, humour, structure and rhythm. The two streams combine through a weighted comparison and a fixed calibration, and the app translates the result into bands: 90%+ Excellent, 75 to 89% Good, 60 to 74% Fair, under 60% Low. The Voice Match method sets out how the whole framework fits together, and the voice match glossary entry defines the score.

How to analyse word length distribution in your own writing

Analysing word length distribution by hand is mostly counting letters, patiently. Here's how to do it without any tool:

  1. Take a sample of your writing, at least a few hundred words.
  2. Strip out apostrophes from contractions (don't becomes dont for counting purposes).
  3. Go through every word and sort it: 1 to 4 letters is short, 5 to 7 is medium, 8 or more is long.
  4. Total up each band, then divide by the total word count to get a percentage for each.
  5. Compare your short and long shares with the corpus figures in the next section to see where you sit.

That's the manual method. It's slow, but it's exact, and doing it once by hand tends to sharpen your eye for your own habits afterwards. If you'd rather skip the counting, ScriptGrain's free voice profile reads word length distribution straight from a writing sample, as one of all 45 attributes.

Typical word length distribution in published writing

Typical published writing sits in a fairly narrow band for both short and long words, at least according to ScriptGrain's reference corpus: 299 recent public English-language pieces (blogs, newsletters and essay sites) from 47 sources, built on 20 September 2026, each measured with the product's own code and then discarded. For short words (1 to 4 letters), the corpus shows a 10th percentile of 50%, a 25th of 53%, a median of 55%, a 75th of 59% and a 90th of 62%. For long words (8 or more letters), the 10th percentile sits at 11%, the 25th at 13%, the median at 16%, the 75th at 18% and the 90th at 20%.

Read that plainly: most published writing is more than half short words, and long words rarely climb past a fifth of the total. Writing that sits well outside these ranges, heavy on long words especially, tends to read as dense or formal even when the sentences themselves are short. The free writing style analyser reports a long-word share too, but at 7+ letters, so its figure runs higher than the 8+ corpus numbers here. Flesch reading ease covers readability more broadly.

How to make AI use shorter, plainer words

Getting an AI assistant to use shorter, plainer words is mostly a matter of naming the swap you want, directly, rather than asking vaguely for "simpler" writing.

  1. Ask for the everyday word over the Latinate one: use, not utilise; help, not facilitate.
  2. Turn nominalisations back into verbs: decide, not make a decision.
  3. Cut stacked modifiers and let one precise noun do the work, instead of three adjectives propping up a vague one.

When a ScriptGrain voice profile writes for you, your stored word length shares go into the system prompt's voice profile block as a hard constraint, alongside a register anchor: "If you notice yourself elevating diction, nominalizing verbs, or reaching for conceptual metaphor, stop and write the plain version instead." The profile's shares reach ChatGPT, Claude and other tools through the API and the MCP server.

How to make AI use longer, more precise words

Getting an AI assistant to use longer, more precise words means asking for accuracy, not length for its own sake.

  1. Allow the technical term when it's the accurate one, and define it once.
  2. Ask for precise nouns and verbs in place of vague short ones such as thing, stuff and get.
  3. Let a concept take its proper name instead of a paraphrase.

A ScriptGrain profile uses the same mechanism when your stored shares lean toward medium and long words: the shares sit in the system prompt as a constraint, the same register anchor applies, and the profile reaches ChatGPT, Claude and other tools through the API and MCP server. Related attributes worth checking here are contraction frequency and filler phrases, since both interact with how heavy or light a voice reads.

Questions

What is a word length analyser?

A word length analyser is a tool that sorts the words in a piece of writing into short, medium and long bands and reports what share falls into each. ScriptGrain's free analyser reports the long-word share (7+ letters) as a percentage. Draft scoring counts all three shares (short, medium, long) with a different cut-off: 8+ letters for long, not 7+.

How is word length scored?

When a draft is scored, code counts letters with apostrophes removed and sorts every word into short (1 to 4 letters), medium (5 to 7) or long (8 or more), as shares of all words. It compares those shares with the profile's at weight 1, as one of the 28 attributes that can move a voice match score. The profile's own shares are the extraction model's reading of your samples.

How do I check the average word length in my writing?

Count letters, not syllables. Strip apostrophes, sort every word into short (1 to 4 letters), medium (5 to 7) or long (8 or more), then divide each band's total by your overall word count. For an average word length, divide total letters by total words instead. In the middle 80% of ScriptGrain's 299-piece reference corpus, short words run from 50% to 62% and long words from 11% to 20%.

How do I stop ChatGPT using big words?

Name the swap directly rather than asking for "simpler" writing in general. Ask it to use everyday words over Latinate ones (use instead of utilise), turn nominalisations back into verbs (decide instead of make a decision), and cut stacked modifiers so one precise noun carries the sentence. In ScriptGrain, a profile's word length shares go into the generator's system prompt, beside a register anchor that asks for the plain version.

How do I make AI writing plain English?

Plain English means shorter, more common words doing precise work, not vague ones doing padding. Push the model to prefer everyday vocabulary, turn abstract nouns back into verbs, and trim modifier stacks. If you want this held across every draft, a ScriptGrain voice profile stores your own word length shares, and ChatGPT, Claude and other tools can read them through the API and MCP server.

Related

Sources