Rare word rate analyser

By Jack Stovell · published 2026-09-24

Rare word rate is how often uncommon words turn up in your writing, expressed as rare words per 1,000 words. It's one of six lexical attributes in a ScriptGrain voice profile. The extraction model judges it from your samples with no definition of "rare"; no code counts it, and it never moves your voice match score.

Analyse your rare word rate free

Your free voice profile measures rare word rate among all 45 attributes, from a sample of your own writing. No card needed.

Try the six-signal check first

At a glance

What rare word rate measures

Rare word rate measures how often you reach for the uncommon word over the obvious one. The app's own definition is blunt: "how often uncommon words appear." That's it. No formula, no dictionary tier, no frequency list.

So what counts as "uncommon"? The honest answer is that it's a judgement call. The extraction model (Claude Sonnet 5, running two passes: one per sample, one to synthesise) reads your writing and forms an impression. It's asked for a float per 1,000 words. It isn't given a definition of "rare." That leaves room for context: "escarpment" is rare in a marketing newsletter and unremarkable in a geology textbook. A model can weigh that. A fixed word list can't.

This sits in the Lexical group alongside vocabulary diversity, preferred words, and word length distribution. Together they sketch the shape of your vocabulary: how varied it is, which words you lean on, how long your words tend to run, and how often you go rare. See the full lexical group for how they fit together.

How rare word rate is scored

Rare word rate doesn't move your voice match score. Full stop. There's no comparator built for it in the voice-match engine, which means when the system checks a draft against your profile, this attribute sits out. Polish passes can't target it either, because there's nothing to aim at.

That doesn't make it decorative. It still goes to the generator with the rest of the profile (more on that further down), and it still describes how you write. It just doesn't feed the voice match number the app shows you. If you want the mechanics of how that score does work (which features are counted by code, which are judged by a model, how the weighting lands), the voice measurement framework covers it. And if "voice match" itself is a fuzzy term to you, the voice match glossary entry defines it.

How to analyse rare word rate in your own writing

You don't need software to get a rough read on this. Grab a sample of your own writing, at least a few hundred words, and work through it by hand.

  1. Read the passage once for content, ignoring style entirely.
  2. Read it again, this time circling every word you wouldn't expect a general reader to know without pausing.
  3. Count the circled words.
  4. Divide by the total word count, then multiply by 1,000, to get a rough rate per 1,000 words.
  5. Repeat across two or three samples and compare. Consistency across samples tells you more than any single count.

That gives you an estimate, not a verdict: "would a general reader pause here" is a judgement, and yours will drift on a tired Tuesday. ScriptGrain's free voice profile reads this attribute from a writing sample as one of all 45, no manual circling required.

One thing to watch for: the free writing style analyser shows a "Words used exactly once" figure under this attribute's name, and it is a different measurement entirely. It's the hapax legomenon share, the percentage of words in a specific piece that appear only once in it. That's about repetition within one text. Rare word rate is about how uncommon your vocabulary is against the language generally. Different unit, different idea, easy to conflate.

Examples of rare word rate in real writing

Example 1: "The negotiation stalled." Plain, ordinary, low rare word rate.

Example 2: "The negotiation foundered on a single obdurate clause." Same event, higher rare word rate: "foundered" and "obdurate" are doing work that "stalled" and "difficult" would do more plainly.

Example 3: "She read the report twice before she believed it." No rare words at all, and it doesn't need any: the sentence's power sits in rhythm, not vocabulary.

How to make AI use more unusual words

If your AI assistant sounds flat, give it explicit permission and a clear target rather than a vague instruction to "sound better."

  1. Tell it directly: use the exact, less common word when it's more precise than the familiar one. Precision, not decoration, is the test.
  2. Show it a passage with a rare word rate you like and ask it to match the frequency, never to lift the actual words.
  3. Ask for concrete, specific nouns instead of generic category words: "spaniel" instead of "dog," "ledger" instead of "document."

A ScriptGrain profile sends rare word rate in the voice profile JSON block with every other attribute, under "Treat every attribute as a hard constraint, not a suggestion." A register anchor adds: "The profile is the CEILING for stylistic elevation, not the floor." So a ScriptGrain draft is told not to run fancier than your samples. The profile reaches ChatGPT, Claude and other tools through the API and MCP server, so the same stored rate travels with you.

How to make AI use fewer rare words

If your AI keeps reaching for words your reader would need a dictionary for, the fix is a direct constraint, not a polite request.

  1. Ask explicitly for words a general reader knows without a dictionary.
  2. Tell it your samples are the ceiling: no diction more elevated than yours, ever.
  3. Ban the specific AI vocabulary it keeps reaching for, by name, such as "constellation of".

The ceiling anchor is built for this case: ScriptGrain's generator is told the profile is the ceiling for stylistic elevation, so a plain writer's profile asks for plain drafts even when the model leans grander. Through the API and MCP server, ChatGPT, Claude and other tools read the same stored rate.

Questions

What is a hapax legomenon?

A hapax legomenon is a word that appears exactly once in a given text. It's a term from corpus linguistics, and it's the idea behind ScriptGrain's free "Words used exactly once" figure, the hapax share as a percentage of words in that specific piece. It's a different measurement from rare word rate: hapax share is about repetition within one text, rare word rate is about uncommonness in the language generally.

How is rare word rate scored?

It isn't scored in the sense of moving a number. Rare word rate has no comparator in ScriptGrain's voice-match engine, so it never affects your voice match score and polish tools can't target it. It's still extracted and stored (a float per 1,000 words, judged by the extraction model reading your samples), and it still goes to the generator with the rest of the profile, just not into the score you see on a comparison.

How do I measure rare words in my own writing?

Read a sample twice: once for content, once to circle words a general reader might not know without pausing. Count the circled words, divide by total words, multiply by 1,000. That gives a rough per-1,000 rate. For consistency, do this across a few samples rather than one, since a single passage can mislead you either way.

How do I make AI use simpler vocabulary?

Ask for words a general reader knows without a dictionary, tell the model your own writing is the ceiling for elevation, and name the specific words you want banned, such as "constellation of." A named list of words to avoid gives the model a clearer target than a vague instruction like "write simply".

How do I make AI use more interesting words?

Give explicit permission to use a precise, less common word over a vague familiar one, and ask for concrete nouns rather than generic categories. Showing the model a sample passage and asking it to match your rate of unusual words (not copy the words themselves) gives it a clearer target than a general instruction to "sound smarter".

Related

Sources