# Lexical score analyser: every attribute, explained

> A lexical score analyser measures word choice: how varied your vocabulary is, which words you keep reaching for, how long and rare they run, and how often you u

Canonical: https://scriptgrain.com/attributes/lexical

# Lexical score analyser

*By Jack Stovell · published 2026-09-24*

A lexical score analyser measures word choice: how varied your vocabulary is, which words you keep reaching for, how long and rare they run, and how often you use contractions such as "it's" and "don't". ScriptGrain reads six of these attributes as part of a 45-attribute voice profile, and four of the six move your voice match score.

## Analyse your lexical attributes free

Your free voice profile measures every lexical attribute among all 45, from a sample of your own writing. No card needed.

[Try the six-signal check first](https://scriptgrain.com/#sg-measure)

## At a glance

- Attributes: **6 of 45** (ScriptGrain voice profile)
- Counted in code: **3 (when a draft is scored, or in the free tools)** (ScriptGrain voice-match engine and free tools)
- Judged by a model: **1** (ScriptGrain extraction and voice-match engine)
- Kept as word lists: **2** (ScriptGrain extraction)
- Move the voice match score: **4 of 6** (ScriptGrain voice-match engine (published method))

## What the lexical group measures

The lexical group captures word choice: how varied it is, which words recur, how long and rare they are, and how often you contract. It is one of eight groups in the 45 voice attributes, alongside syntax, tone and five others; the lexical layer reference holds the corpus data for this one.

Vocabulary diversity scores how varied your word choice is, on a scale from 0 (highly repetitive) to 1 (maximally varied). The free in-browser tools count it as a [type-token ratio](https://scriptgrain.com/glossary/type-token-ratio) (distinct words divided by total words); the value in your profile is the extraction model's 0 to 1 estimate, with no formula behind it. Draft scoring doesn't measure it, so it doesn't move your voice match score.

Preferred words are the most-recurring distinctive words across your samples, up to five per analysis. The extraction model catalogues them verbatim, and they feed the shared signature-phrase check in voice match, at weight 2. Example: a writer who leans on "genuinely" and "distinctive" in every third paragraph will see both surface here.

Word length distribution is the share of short, medium and long words in your writing. Code counts it from the draft when a draft is scored, and it carries a voice match weight of 1. A profile heavy on short words reads differently to a reader than one stacked with long ones, even before you notice why.

Rare word rate tracks how often uncommon words appear, expressed as rare words per 1,000 words. A model judges it, no code counts it, and it doesn't move your voice match score.

Filler phrases are the recurring filler expressions you reach for, catalogued up to five per analysis, carrying a voice match weight of 2. Think "to be fair" or "at the end of the day": small, repeated, and more revealing than most writers expect.

[Contraction frequency](https://scriptgrain.com/attributes/contraction-frequency) measures how often contractions like "it's", "don't" and "we're" appear, expressed as contractions per 1,000 words. Code counts it when a draft is scored, and it carries a voice match weight of 2.

So: code counts three of the six (word length distribution and contraction frequency when a draft is scored, vocabulary diversity only in the free tools), a model judges one (rare word rate), and two are catalogued as lists (preferred words, filler phrases). Four of the six move your voice match score. Every value stored in your profile, whichever bucket it falls into, comes from the same source: the extraction model, Claude Sonnet 5, reading your samples. There's no dial you turn to say "I use a lot of rare words." The model reads the text and decides.

For the other seven groups, the [writing voice attributes](https://scriptgrain.com/reference/writing-voice-attributes) reference lays out all 45.

## How lexical features are scored

Lexical features are scored in three ways: code counts two of them in the draft, code looks for the two word lists in the draft, and two never reach the score at all. That split matters more than it sounds like it should.

Word length distribution and contraction frequency are code counts. When a draft is scored, code runs across it, measures word lengths, spots contractions and produces a number. No interpretation, just arithmetic against text. Vocabulary diversity is counted only in the free in-browser tools, as a type-token ratio; draft scoring doesn't measure it.

Rare word rate is judged: the extraction model decides what counts as "rare", and no code counts it.

Preferred words and filler phrases are catalogued: the model reads your samples and pulls out the words and phrases that recur, up to five each per analysis. There's no 0 to 1 scale for either, just a list. When a draft is scored, code counts how many entries from those lists turn up in it.

When it comes to whether any of this moves your voice match score, four of the six do: preferred words, word length distribution, filler phrases, and contraction frequency. Vocabulary diversity and rare word rate get recorded but sit outside the score. Across the full 45-attribute profile, 28 attributes can move a voice match score, so lexical supplies 4 of the 28. The [Voice Match method](https://scriptgrain.com/reference/voice-measurement-framework) works the same way for every group: code counts sentence, punctuation, pronoun, article and word-length features, Claude Haiku judges formality, humour, structure and rhythm, and a weighted comparison runs through a fixed calibration to land on a [voice match](https://scriptgrain.com/glossary/voice-match) percentage.

If you want the theory behind why diversity and density get treated as separate ideas rather than one number, the [lexical density and diversity](https://scriptgrain.com/glossary/lexical-density-diversity) glossary entry covers the distinction.

## How to analyse the lexical of your own writing

You can analyse the lexical patterns in your own writing by hand, with nothing more than a sample and a bit of patience.

1. Pull a sample of your writing, at least a few hundred words, ideally more.
2. Count total words, then count unique words. Divide unique by total for a rough diversity ratio.
3. Scan for words you use more than twice that aren't common function words. Those are your preferred words.
4. Sort every word by length (short, medium, long) and estimate the share each bucket takes.
5. Flag anything you'd call an uncommon or technical word, then work out how many appear per 1,000 words.
6. Search for phrases you repeat as verbal filler: "to be fair", "at the end of the day", "so", that kind of thing.
7. Count contractions like "it's" and "don't", then scale that count to a per-1,000-word rate.

Do this across a few different pieces and patterns start to show up whether you're looking for them or not. ScriptGrain's free voice profile reads all six of these lexical attributes from a writing sample, alongside the other 39.

## How to make AI writing match your lexical

You make AI writing match your lexical patterns by giving it a structured read of your vocabulary instead of a vague style note.

1. Gather writing samples, ideally several, each over 50 words; the app suggests 3,000+ words total across up to 20 samples.
2. Run the samples through analysis so the model can extract vocabulary diversity, preferred words, word length shares, rare word rate, filler phrases and contraction frequency.
3. Check the resulting profile against pieces you know are "you". Preferred words and fillers are usually the fastest gut check.
4. Feed that profile into a drafting tool, rather than typing "sound more casual" into a blank prompt.
5. Compare generated drafts back against the profile using a voice match score, and revise where the lexical attributes drift.

A ScriptGrain voice profile carries the lexical group, and the other seven groups alongside it, into any draft generated through the platform. The same profile reaches ChatGPT, Claude and other tools through the API and the MCP server: `GET /v1/profiles/{id}` and the MCP `get_profile` tool both return all 45 attributes as JSON, lexical included, so whichever tool is doing the writing reads the same preferred words and contraction frequency every time.

## Attributes in this group

- [Vocabulary diversity](https://scriptgrain.com/attributes/vocabulary-diversity): How varied the word choice is: 0 means highly repetitive, 1 means maximally varied.
- [Preferred words](https://scriptgrain.com/attributes/preferred-words): The most-recurring distinctive words across your samples.
- [Word length distribution](https://scriptgrain.com/attributes/word-length-distribution): The share of short, medium and long words in your writing.
- [Rare word rate](https://scriptgrain.com/attributes/rare-word-rate): How often uncommon words appear.
- [Filler phrases](https://scriptgrain.com/attributes/filler-phrases): Filler expressions you reach for repeatedly.
- [Contraction frequency](https://scriptgrain.com/attributes/contraction-frequency): How often contractions (it's, don't, we're) appear.

## Questions

### What is a lexical score analyser?

It's a tool that reads your vocabulary and turns it into measurable attributes: how varied your word choice is, which words you repeat, how long and rare they run, your filler phrases, and how often you contract. ScriptGrain reads six such attributes as part of a 45-attribute voice profile, and four of the six feed into the voice match score.

### How is vocabulary scored in writing?

When a draft is scored, code counts word length distribution and contraction frequency directly, and counts hits from the two word lists (preferred words, filler phrases). Vocabulary diversity is counted only in the free tools, and a model judges rare word rate. Every value stored in a profile, whatever the method, is the extraction model's reading of your samples.

### How do I analyse the vocabulary in my own writing?

Take a sample, count unique words against total words for a diversity ratio, note words you repeat distinctively, estimate short, medium and long word shares, flag uncommon words per 1,000 words, and log any filler phrases or contractions you lean on. Doing this across several pieces reveals patterns you'd otherwise never notice about your own habits.

### What is lexical stylometry?

Lexical stylometry is the study of word-level patterns (vocabulary diversity, word length, rare word use) as a fingerprint for identifying or describing a writer's style. It overlaps with ideas like type-token ratio and lexical density, and in ScriptGrain it is one of the eight groups in a 45-attribute voice profile.

### How do I make AI match my vocabulary?

Feed it samples of your actual writing, not instructions about your writing. A model reads the samples and extracts your vocabulary diversity, preferred words, word length distribution, rare word rate, filler phrases and contraction frequency into a profile. That profile then travels with your drafts and into other tools through ScriptGrain's API and MCP server, so every tool reads the same vocabulary.

## Related

- [All 45 voice attributes](https://scriptgrain.com/attributes)
- [Lexical reference data](https://scriptgrain.com/reference/lexical-layer)
- [How voice match is scored](https://scriptgrain.com/reference/voice-measurement-framework)

## Sources

- [How ScriptGrain scores voice match](https://scriptgrain.com/reference/voice-measurement-framework)
