# Writing & stylometry glossary · ScriptGrain

> Plain-English definitions of the terms behind writing measurement, each with real numbers, cited sources, and a worked example.

Canonical: https://scriptgrain.com/glossary

# The vocabulary of writing measurement

Plain-English definitions of the terms behind writing measurement, each with real numbers, cited sources, and a worked example.

- [Voice Match](https://scriptgrain.com/glossary/voice-match)
Voice Match is ScriptGrain's measure of how closely a piece of writing aligns with a defined writing-voice profile, across the style features that profile measures. It is computed by comparing measured features of the draft against the profile, not by asking a model whether the text sounds right.
- [Burstiness](https://scriptgrain.com/glossary/burstiness)
Burstiness measures how much writing varies across a document, most visibly in sentence length and structure. Human writing tends to alternate short and long sentences; AI-generated text is typically more uniform, producing low burstiness.
- [Perplexity](https://scriptgrain.com/glossary/perplexity)
Perplexity measures how predictable a text is to a language model: the average surprise per word. Low perplexity means the model finds each word easy to predict; AI-generated text typically scores lower than human writing.
- [Sentence Length Variance](https://scriptgrain.com/glossary/sentence-length-variance)
Sentence length variance measures how widely sentence lengths spread around their average in a passage, reported as variance or standard deviation in words. High variance signals varied rhythm; low variance signals uniform, metronomic prose.
- [Flesch Reading Ease](https://scriptgrain.com/glossary/flesch-reading-ease)
The Flesch reading ease score rates text readability from 0 to 100 using average sentence length and average syllables per word. Higher scores mean easier reading; 60-69 is standard, readable by high school students.
- [Type-token ratio (TTR)](https://scriptgrain.com/glossary/type-token-ratio)
The type-token ratio (TTR) measures vocabulary diversity by dividing the number of distinct words (types) by the total number of words (tokens) in a text. Higher values mean less repetition; scores are only comparable between equal-length texts.
- [Hapax legomenon](https://scriptgrain.com/glossary/hapax-legomenon)
A hapax legomenon is a word that occurs exactly once within a given context: a single text, an author's complete works, or a language's entire written record. From Greek, 'said once'; plural: hapax legomena.
- [Function words](https://scriptgrain.com/glossary/function-words)
Function words (the, of, and, to, was) carry grammatical structure rather than content meaning. They form a small closed class, used at high frequency and largely unconsciously, which makes their usage rates powerful evidence in authorship attribution.
- [Lexical density vs lexical diversity](https://scriptgrain.com/glossary/lexical-density-diversity)
Lexical density is the share of a text's words that are content words (nouns, verbs, adjectives, adverbs). Lexical diversity is the range of different words used. A text can score high on one and low on the other.
- [Stylometric fingerprint](https://scriptgrain.com/glossary/stylometric-fingerprint)
A stylometric fingerprint, also called a writeprint, is the measurable pattern of a writer's unconscious habits, from word choice to punctuation, that stays stable enough across texts to identify authorship, much as a physical fingerprint identifies a person.
- [Idiolect](https://scriptgrain.com/glossary/idiolect)
An idiolect is one person's individual version of a language: the vocabulary, grammar, spelling and phrasing habits unique to them. In writing, your idiolect is effectively your fingerprint; no two people use language identically.
- [Voice vs tone vs style](https://scriptgrain.com/glossary/voice-tone-style)
Voice is a writer's stable personality on the page; tone is how that personality flexes for context and audience; style is the concrete, countable choices, such as sentence length and punctuation, that carry both.
- [Passive voice](https://scriptgrain.com/glossary/passive-voice)
The passive voice is a grammatical construction in which the subject receives the action rather than performing it, formed with a form of 'be' plus a past participle: 'the report was submitted'. It is a choice, not an error.
- [Humanised AI content](https://scriptgrain.com/glossary/humanised-ai-content)
Humanised AI content is machine-drafted text revised so that its measurable habits (sentence rhythm, contractions, phrasing, punctuation) fall inside the range of human writing, as distinct from text rewritten only to fool a detector. The revision can be measured: before-and-after counts of the tells, and a voice-match score against a real writer.
- [Tone matching](https://scriptgrain.com/glossary/tone-matching)
Tone matching is making a piece of writing carry the same register as a reference: the same formality, warmth, humour and confidence. It is one layer of voice matching, not the whole of it; a draft can match a writer's tone and still miss their sentence rhythm, punctuation and vocabulary, which is why tone alone is a weak test.
- [Brand voice](https://scriptgrain.com/glossary/brand-voice)
Brand voice is the consistent set of writing habits a brand's published copy shares: its register, sentence rhythm, vocabulary, punctuation and structure. Described, it is a page of adjectives; measured, it is a set of attributes extracted from the brand's own writing that any new draft can be scored against.
- [Voice drift](https://scriptgrain.com/glossary/voice-drift)
Voice drift is the gradual movement of a writer's, brand's or client's published writing away from its own established habits: sentences lengthen or flatten, contractions disappear, the register formalises, new fillers creep in. It is measurable as the change in voice-match score between a baseline and later pieces, and it is what multi-writer teams, ghostwriters and AI assistants most reliably introduce.
