Tone matching

Tone matching is making a piece of writing carry the same register as a reference: the same formality, warmth, humour and confidence. It is one layer of voice matching, not the whole of it; a draft can match a writer's tone and still miss their sentence rhythm, punctuation and vocabulary, which is why tone alone is a weak test.

The numbers

What tone matching is

Tone matching means checking whether a piece of writing sounds formal or casual, warm or dry, confident or hedging, in the way that matches a target voice. It sits alongside vocabulary, sentence shape, rhetoric, punctuation, function words, content patterns and cadence as one of eight layers in the full voice, tone and style picture. Tone gets the most attention because it's the easiest thing to notice and the easiest thing to fake with a slider. That's exactly why it's a weak test on its own, which the next two sections cover properly.

How tone is measured

Six attributes make up the tone and register layer. Formality score runs 0 to 10, judged, carrying a weight of 2 and a tolerance of 3. Humour register sorts into dry, sarcastic, self-deprecating, warm or none, judged, weighted at 1.5, with adjacent registers scoring half credit rather than zero. Emotional expressiveness is low, medium or high, judged, weight 1. Contraction rate is counted, not judged: per 1,000 words, weight 2. Confidence versus hedging scores 0 to 1, judged, weight 1.5, tolerance 0.35. Audience adaptation gets recorded on the profile rather than scored against a fixed target.

So one of the six is counted outright, one is recorded on the profile, and four rest on judgement calls with built-in tolerance. That mix matters. A counted number like contraction rate doesn't argue with you. A judged score like humour register needs a human or model call, and the tolerance band exists because reasonable readers won't always land on the exact same number.

Why tone alone is a weak test

Here's the thing: two pieces of writing can share a tone label and still read nothing alike. SGR-002 found exactly this. GPT-5, working from the free prompt builder's one-page brief, was asked for about 14 contractions per 1,000 words and a formality score of 3.7. It obeyed the countable rules faithfully, more short sentences, no em dashes, and still produced 0.5 contractions per 1,000 words against a judged formality of 7.2. Wrong tone, correct mechanics.

That gap is the whole argument. A tone slider set to "casual" tells a model roughly what register to aim for. It says nothing about sentence length distribution, punctuation habits, rhetorical structure or the other thirty-nine attributes that actually carry a voice. Two drafts can both claim "casual" and differ on twenty other measurements without either one being wrong about its label.

What it cannot tell you

Tone tells you the register a piece is aiming for. It doesn't tell you whether the sentences are long or short, whether the punctuation is restrained or excitable, whether the vocabulary runs plain or ornate, or whether the argument builds gradually or lands its point up front. Those live in other layers entirely. A brand voice consistency checker that only checks tone will wave through a draft that gets the mood right and everything else wrong.

Worked example

take one request, "explain why the invoice is late," and two invented replies. Reply A: "We regret to inform you that the delay in processing your invoice was due to an unforeseen administrative oversight, for which we sincerely apologise." Reply B: "Sorry, your invoice is late. We messed up the paperwork. Fixing it now."

Reply A scores formality 8, zero contractions, and reads as hedging rather than declarative, all that "regret to inform" and "sincerely apologise" softening. Reply B scores formality 3, roughly 20 contractions per 1,000 words at that density, and states things plainly: late, messed up, fixing it. A tone slider set to "casual" would nudge a model toward something closer to Reply B.

what the slider captures is register. What it misses is everything else that makes B feel like a different voice from A rather than just a different mood, sentence rhythm, word choice, how the apology is structured, whether it hedges or commits. Tone matching gets you in the right neighbourhood. It doesn't get you to the right house.

Sources

More terms