Lexical density vs lexical diversity
Lexical density is the share of a text's words that are content words (nouns, verbs, adjectives, adverbs). Lexical diversity is the range of different words used. A text can score high on one and low on the other.
The numbers
- Typical lexical density of written English (Ure): above 40% (Wikipedia: Lexical density)
- Typical lexical density of spoken English (Ure): below 40% (Wikipedia: Lexical density)
- Lexical density range Stubbs found in fiction: 40 to 54% (Wikipedia: Lexical density)
- Lexical density range Stubbs found in non-fiction: 40 to 65% (Wikipedia: Lexical density)
Two different measures, routinely confused
Lexical density measures how much of a text is made of content words: nouns, main verbs, adjectives and adverbs. Ure's 1971 formula is the number of lexical words divided by the total number of words, multiplied by 100; Halliday later proposed a variant that divides by the number of clauses instead. Lexical diversity measures something else: how varied the vocabulary is, expressed through the relationship between unique words (types) and total words (tokens) using indices such as TTR, vocd and MTLD.
The two are routinely conflated because both start from word counts and both are reported as single scores. But density asks how information-packed a text is, while diversity asks how repetitive it is, and they can move in opposite directions, as the worked example below shows.
Typical density values
Ure's threshold is the standard reference point: written English usually has a lexical density above 40% and spoken English below 40%, because speech leans harder on pronouns, auxiliaries and other grammatical words. Stubbs found fiction typically in the 40 to 54% range and non-fiction between 40 and 65%.
Density is sometimes treated as a proxy for difficulty, but Text Inspector notes the absence of convincing proof of any strong link between lexical density and readability, and prioritises diversity measures for that reason.
Diversity and text length
The simplest diversity measure, the type-token ratio, falls as texts get longer, because a longer text must reuse words. Raw TTR should therefore only compare equal-length samples; indices such as MTLD, vocd and the moving-average MATTR were designed to reduce this length sensitivity. See the type-token-ratio entry for the mechanics.
Worked example
Compare two versions of one sentence. Version A (10 words): 'The dogs were barking loudly because the dogs were hungry.' Version B (9 words): 'The dogs were barking loudly because they were hungry.'
Version A: content words are dogs, barking, loudly, dogs, hungry = 5 of 10 tokens, so density = 5 / 10 = 50%; distinct types = 7 of 10, so TTR = 70%. Version B: content words are dogs, barking, loudly, hungry = 4 of 9 tokens, so density = 4 / 9 = 44%; distinct types = 8 of 9, so TTR = 89%.
Replacing one repeated noun with a pronoun lowered density (less content per word) but raised diversity (less repetition). The two measures answer different questions and must not be used as synonyms.
Sources
- Wikipedia: Lexical density
- Wikipedia: Lexical diversity
- Text Inspector: Lexical Density vs Lexical Diversity
More terms
- Burstiness
- Perplexity
- Sentence Length Variance
- Flesch Reading Ease
- Type-token ratio (TTR)
- Hapax legomenon
- Function words
- Stylometric fingerprint
- Idiolect
- Voice vs tone vs style
- Passive voice
ScriptGrain · Why ScriptGrain · Sample voice profiles · The Journal · Support