Lexical density vs lexical diversity

Lexical density is the share of a text's words that are content words (nouns, verbs, adjectives, adverbs). Lexical diversity is the range of different words used. A text can score high on one and low on the other.

The numbers

Two different measures, routinely confused

Lexical density measures how much of a text is made of content words: nouns, main verbs, adjectives and adverbs. Ure's 1971 formula is the number of lexical words divided by the total number of words, multiplied by 100; Halliday later proposed a variant that divides by the number of clauses instead. Lexical diversity measures something else: how varied the vocabulary is, expressed through the relationship between unique words (types) and total words (tokens) using indices such as TTR, vocd and MTLD.

The two are routinely conflated because both start from word counts and both are reported as single scores. But density asks how information-packed a text is, while diversity asks how repetitive it is, and they can move in opposite directions, as the worked example below shows.

Typical density values

Ure's threshold is the standard reference point: written English usually has a lexical density above 40% and spoken English below 40%, because speech leans harder on pronouns, auxiliaries and other grammatical words. Stubbs found fiction typically in the 40 to 54% range and non-fiction between 40 and 65%.

Density is sometimes treated as a proxy for difficulty, but Text Inspector notes the absence of convincing proof of any strong link between lexical density and readability, and prioritises diversity measures for that reason.

Diversity and text length

The simplest diversity measure, the type-token ratio, falls as texts get longer, because a longer text must reuse words. Raw TTR should therefore only compare equal-length samples; indices such as MTLD, vocd and the moving-average MATTR were designed to reduce this length sensitivity. See the type-token-ratio entry for the mechanics.

Worked example

Compare two versions of one sentence. Version A (10 words): 'The dogs were barking loudly because the dogs were hungry.' Version B (9 words): 'The dogs were barking loudly because they were hungry.'

Version A: content words are dogs, barking, loudly, dogs, hungry = 5 of 10 tokens, so density = 5 / 10 = 50%; distinct types = 7 of 10, so TTR = 70%. Version B: content words are dogs, barking, loudly, hungry = 4 of 9 tokens, so density = 4 / 9 = 44%; distinct types = 8 of 9, so TTR = 89%.

Replacing one repeated noun with a pronoun lowered density (less content per word) but raised diversity (less repetition). The two measures answer different questions and must not be used as synonyms.

Sources

More terms

ScriptGrain · Why ScriptGrain · Sample voice profiles · The Journal · Support