# How a Voice Match score is calculated: method, weights, limits · ScriptGrain

> The published method behind ScriptGrain's Voice Match score: which features are counted in code, which are judged, the weights and tolerances, the calibration from raw agreement to the number you see, and where the score is unreliable.

Canonical: https://scriptgrain.com/reference/voice-measurement-framework

# How writing voice is measured: the Voice Match method

*By Jack Stovell · published 2026-09-21 · checked 2026-09-20*

Voice Match is a weighted comparison of a draft's measured features against a profile's, calibrated onto a 0 to 1 scale. It measures distance from a voice, not quality. A brilliant paragraph in the wrong voice scores low; a mediocre one in the right voice scores high. That's the whole premise.

## What Voice Match measures

Here's the thing: some of what makes a voice a voice can be counted, and some of it can't. Sentence length, contraction rate, comma density: these are arithmetic. Humour register, argument structure, whether a sentence feels confident or hedged: these need judgement.

So the method splits down the middle. One set of features gets measured straight from the text, no model involved, no interpretation required. The other gets read by a small model trained to notice things a word-counter can't. Both feed into the same score, weighted by how much each one tends to carry a voice.

Full detail on every attribute the system tracks, including the ones covered below, sits at [/reference/writing-voice-attributes](https://scriptgrain.com/reference/writing-voice-attributes).

## Measured in code

No model reads these. They come straight out of the text: average sentence length, sentence-length variance, contractions per 1,000 words, commas per sentence, exclamation marks per 1,000, semicolons per 1,000, parentheticals per 300 words, whether the body contains a question mark, ellipses per 1,000, pronoun mix across I, we and you, article balance across the, a and an, and the short-to-medium-to-long word mix.

Two more sit alongside these: how many of the profile's catalogued signature words and phrases turn up in the draft, and how many author-banned terms appear. The second one only ever costs points. A clean draft earns nothing from avoiding banned terms; it just doesn't lose anything either.

None of this requires interpretation. A sentence either has 14 words or it doesn't. That's the appeal, and also the limit: code can tell you a sentence is short, but not whether it's short because the writer is being punchy or just being terse.

## Judged by a model

For the qualities that resist counting, a small model (Claude Haiku) reads the draft directly. It scores formality on a 0 to 10 scale, confidence versus hedging on 0 to 1, and then works through a longer list: metaphor rate, humour register, emotional expressiveness, opening style, argument structure, rhythm, sentence complexity, clause ordering, specificity, repetition as emphasis, analogy use.

These are the features a rule-based system genuinely can't get to. You can count semicolons. You can't count "dry humour" without something that understands what dry humour sounds like. That's the trade-off: judgement buys you access to real stylistic texture, at the cost of a small amount of run-to-run variation. More on that in the reliability section below.

## Tolerances and weights

Every feature, whether measured or judged, gets converted into a similarity score between 0 and 1. For numeric features the formula is simple: similarity equals 1 minus the distance from the profile value, divided by a tolerance, floored at zero. Get further away than the tolerance allows and the similarity just sits at nothing.

The tolerances themselves are not fixed numbers; most scale with the profile. Sentence length tolerance is the larger of 8 words or 70% of the profile's average. Variance tolerance is the larger of 25 or 1.2 times the profile's own variance. Contractions: the larger of 10 per 1,000 words or 70% of the profile's rate. Comma density: the larger of 0.9 or 70% of the profile's. Exclamation marks get a flat 2.5 per 1,000. Semicolons: the larger of 2 per 1,000 or the profile's own rate. Parentheticals: the larger of 1.5 per 300 words or the profile's rate. Formality gets a flat 3 points on the 0 to 10 scale. Confidence versus hedging gets 0.35. Metaphor rate: the larger of 3 per 1,000 or the profile's rate.

Enumerated attributes work differently, because there's no distance to divide. An exact match scores 1. "Mixed" or "varied" on either side scores 0.7. If the profile's value contains the draft's label as a substring, that scores 0.9. Adjacent registers, dry and none, dry and sarcastic, warm and self-deprecating, score 0.5. On ordered scales like low/medium/high or abstract/balanced/data-driven, one step away scores 0.5 and two steps scores 0.1. Yes/no attributes score either 1 or 0.2; there's no partial credit for a coin flip.

Then there's weighting, because not every habit carries a voice equally. Average sentence length, contraction rate, formality, signature phrases and banned terms each carry a weight of 2. That's the top tier. Sentence-length variance, comma density, exclamation rate, semicolon use, pronoun mix, confidence versus hedging, humour register, argument structure and rhythm sit at 1.5. Parentheticals, questions in body, ellipsis habit, word-length mix, emotional expressiveness, sentence complexity, clause ordering and specificity sit at 1. Article balance, opening style and metaphor rate sit at 0.75. Repetition as emphasis and analogy use sit at 0.5, the lightest touch in the system.

One quiet rule worth naming: rare habits like exclamations, semicolons and parentheticals, where both the draft and the profile sit near zero, still count as agreement. But they're never held up as the reason for a match. Plain text has no semicolons. That's not a discovery, it's just arithmetic agreeing with itself.

And a floor: at least six features need to be comparable, or the system refuses to produce a score at all. Below that, there isn't enough signal to mean anything.

Feature · How it is obtained · Weight · Tolerance (distance at which similarity reaches 0)
Average sentence length · counted · 2 · larger of 8 words or 70% of the profile value
Contraction rate · counted · 2 · larger of 10 per 1,000 or 70% of the profile rate
Formality (0 to 10) · judged · 2 · 3 points
Signature phrases · counted against the profile's list · 2 · share of up to four catalogued phrases used
Author-banned terms · counted · 2 (penalty only) · each hit removes 0.5 of similarity
Sentence-length variance · counted · 1.5 · larger of 25 or 1.2 times the profile value
Comma density · counted · 1.5 · larger of 0.9 per sentence or 70% of the profile value
Exclamation rate · counted · 1.5 · 2.5 per 1,000 words
Semicolon use · counted · 1.5 · larger of 2 per 1,000 or the profile rate
Pronoun mix (I / we / you) · counted · 1.5 · distribution distance
Confidence versus hedging · judged · 1.5 · 0.35 on the 0 to 1 scale
Humour register · judged · 1.5 · exact 1, adjacent 0.5, mixed 0.7
Argument structure · judged · 1.5 · exact 1, mixed 0.7
Rhythm · judged · 1.5 · exact 1, mixed 0.7
Parenthetical asides · counted · 1 · larger of 1.5 per 300 words or the profile rate
Questions in body copy · counted · 1 · match 1, mismatch 0.2
Ellipsis habit · counted, then banded · 1 · one band away 0.5, two away 0.1
Word-length mix · counted · 1 · distribution distance
Emotional expressiveness · judged · 1 · one step away 0.5, two steps 0.1
Sentence complexity · judged · 1 · exact 1, mixed 0.7
Clause ordering · judged · 1 · exact 1, mixed 0.7
Specificity · judged · 1 · one step away 0.5, two steps 0.1
Article balance (the / a / an) · counted · 0.75 · distribution distance
Opening style · judged · 0.75 · exact 1
Metaphor rate · judged · 0.75 · larger of 3 per 1,000 or the profile rate
Repetition as emphasis · judged · 0.5 · match 1, mismatch 0.2
Analogy use · judged · 0.5 · match 1, mismatch 0.2

## From raw agreement to the score you see

The raw score is a weighted mean of all those similarities. Simple enough. Except raw scores have a problem: a genuinely on-voice draft typically lands around 0.55 to 0.7 raw, which reads like a poor grade even when the match is strong. Nobody wants to see "0.6" and be told that's good.

So the raw value passes through a fixed, monotone calibration before it's shown to you. Raw 0 displays as 0.05. Raw 0.3 becomes 0.35. Raw 0.45 becomes 0.60. Raw 0.6 becomes 0.85. Raw 0.7 becomes 0.92. Raw 0.85 becomes 0.97. Raw 1 becomes 0.99.

Here's the important bit: the ordering never changes. If draft A scores higher than draft B on the raw scale, it scores higher after calibration too. The calibration reshapes what the number looks like. It never decides which draft is closer to the voice.

Raw weighted agreement · Score shown
0.00 · 0.05
0.30 · 0.35
0.45 · 0.60
0.60 · 0.85
0.70 · 0.92
0.85 · 0.97
1.00 · 0.99

## Reading a score

On-voice drafts typically land between 0.85 and 0.95. Partial matches sit from 0.5 to 0.8. Anything under 0.45 reads as a clearly different voice.

Polish, the revision tool, works toward a target score (0.9 by default) and returns the best version it produces along the way. Humanize uses a lower default target of 0.85, reflecting that its job is different: sounding human, not matching a specific catalogued profile down to the decimal.

## Comparing two pieces without a profile

Not every comparison starts with a profile. The free off-voice check works from two pieces of text alone: the main piece becomes a pseudo-profile, and the second piece is scored against it.

Because there's no real profile underneath, this runs through a separate pairwise calibration. Raw 0.45 shows as 0.45. Raw 0.6 becomes 0.75. Raw 0.7 becomes 0.88. Raw 0.85 becomes 0.96. The bands read differently too: 0.80 and above counts as on voice, 0.55 to 0.79 counts as drifting, and under 0.55 counts as off voice. You can try the checker itself at [/tools/brand-voice-consistency-checker](https://scriptgrain.com/tools/brand-voice-consistency-checker).

## Limits: where the score is unreliable

Short text is the biggest one. The API will accept drafts of 20 words or more, and the free checker needs at least 120, but rates measured per 1,000 words get noisy under roughly 200 words. A 40-word sample can't tell you much about someone's semicolon habit; there just isn't enough of it to count.

Older profiles present a different problem. Some had enumerated fields filled with descriptive text rather than a clean label, a pattern from brand-mimic profiles built before 17 August 2026. Those got capped near 0.62 until the fix landed, because the enumerated-matching logic had nothing clean to compare against.

The judged subset carries a small amount of run-to-run variation too, since a model is doing the reading rather than a fixed rule. And the deepest limitation isn't a bug at all: the score measures distance from a profile, never quality. A brilliant paragraph written in the wrong voice will still score low, and it should. That's not a flaw in the method. That's the method doing its job.

For a fuller look at how a measured voice compares against an unguided model's version of "sounding human," see [/research/chatgpt-vs-a-measured-voice](https://scriptgrain.com/research/chatgpt-vs-a-measured-voice). And for the shorthand definition, [/glossary/voice-match](https://scriptgrain.com/glossary/voice-match) has it in a sentence.

## Worked example

Take a profile with an average sentence length of 16 words, 20 contractions per 1,000 words and a formality score of 4. Now score a draft against it: 21.6 words average, 6 contractions per 1,000, formality 6.

Sentence length first. Tolerance is the larger of 8 or 70% of 16, which is 11.2, so the tolerance is 11.2. The distance between 21.6 and 16 is 5.6. Similarity is 1 minus 5.6 divided by 11.2, which comes to 0.5.

Contractions next. Tolerance is the larger of 10 or 70% of 20, which is 14. The distance between 6 and 20 is 14. Similarity is 1 minus 14 divided by 14, which is 0.

Formality last of the three. Tolerance is a flat 3. The distance between 6 and 4 is 2. Similarity is 1 minus 2 divided by 3, which comes to 0.33.

Each of these three carries a weight of 2. Weighted mean across just these three: (0.5 + 0 + 0.33) divided by 3, giving a raw agreement of 0.28. Run that through the calibration and it lands around 0.33. An off-voice reading, and correctly so; the draft runs notably longer per sentence, drops contractions hard, and reads more formally than the profile calls for.

But that's three features out of roughly twenty. A real comparison pulls in sentence variance, comma density, pronoun mix, confidence versus hedging, humour register, rhythm and more, each with its own tolerance and weight. No single habit decides the outcome. That's deliberate: a writer who happens to skip contractions in one draft shouldn't get flagged as a different voice entirely. The system is built to average across enough signals that one quirk doesn't sink the whole reading, and every score comes back with its full set of deltas, feature by feature, so you can see exactly which ones pulled the number down.

## Questions

### How is a voice match score calculated?

Every measurable and judged feature of a draft gets compared against the same feature in a profile, converted into a similarity between 0 and 1, then combined into a weighted mean. That raw figure passes through a fixed calibration curve before display, which changes how the number looks without changing which draft scored closer to the voice.

### What counts as a good voice match score?

On-voice drafts typically land between 0.85 and 0.95 after calibration. Scores from 0.5 to 0.8 suggest a partial match, close on some features and off on others. Anything under 0.45 signals a clearly different voice. These bands apply to profile-based comparisons; the free pairwise checker uses different thresholds.

### Why does a formal draft score low against a casual profile?

Formality carries a weight of 2 and a flat tolerance of 3 points on its 0 to 10 scale. A two-point gap already costs a third of the available similarity on that feature alone. Combined with related features like contraction rate, which often move in the same direction as formality, the effect compounds rather than cancels out.

### Can a well-written draft still score badly?

Yes, and that's by design. The score measures distance from a specific profile, not writing quality. A sharp, well-argued paragraph written in an unfamiliar voice will score low if its sentence rhythm, formality and habits differ enough from the profile, regardless of how good the writing is on its own terms.

### Why do scores shift slightly between runs?

Part of the comparison relies on a small model judging qualities like humour register, rhythm and argument structure, and model-based judgement carries a small amount of natural variation run to run. The features measured directly from the text (sentence length, contractions, punctuation counts) never change between runs; only the judged subset does.

## Methodology

Every rule on this page is read from the scoring engine (supabase/functions/_shared/voice-match.ts in the ScriptGrain repository, unit-tested) as it stood on 2026-09-20. The same code scores drafts in the app, the API, the MCP server and the free tools, so the published method is the shipped method. The worked example was computed by hand from those rules.

## Sources

- [ScriptGrain glossary: Voice Match](https://scriptgrain.com/glossary/voice-match)
- [ScriptGrain: Writing voice can be measured (the 45 attributes)](https://scriptgrain.com/writing-voice)
- [ScriptGrain API reference (v1)](https://scriptgrain.com/docs/api)
- [ScriptGrain research index (SGR studies)](https://scriptgrain.com/research)
