# Tone and register, measured: formality, humour, hedging · ScriptGrain

> The six tone attributes ScriptGrain measures (formality, humour register, expressiveness, contraction rate, confidence versus hedging, audience adaptation), how a model judges them, and why a prompt can describe register but not hold it.

Canonical: https://scriptgrain.com/reference/tone-and-register

# Tone and register, measured: the attributes behind 'that sounds like us'

*By Jack Stovell · published 2026-09-24 · checked 2026-09-20*

"That sounds like us" is a judgement about register, and register isn't as fuzzy as it feels. Six attributes cover it: two get counted, four get judged by a model reading the text. That split matters more than it sounds like it should, and it's the reason prompts keep failing where measurement doesn't.

## The six tone and register attributes

Out of the full set of 45 voice attributes tracked in the [voice measurement framework](https://scriptgrain.com/reference/voice-measurement-framework), six deal specifically with tone and register. They are formality_score (0 to 10), humour_register (dry, sarcastic, self-deprecating, warm, or none), emotional_expressiveness (low, medium, high), contraction_rate, confidence_vs_hedging (0 for a heavy hedger, 1 for fully declarative), and audience_adaptation, which just asks whether the register shifts depending on who's reading it.

Two of these get counted directly from the text. Contraction rate is one. Audience adaptation, when it's tested across samples, is closer to a structural check than a judgement call. The other four need a reader, human or model, to make a call. Formality, humour, expressiveness and the confidence-hedging balance all involve interpretation. That's not a flaw in the system. It's just how tone works: some of it is arithmetic, most of it is taste.

The full list of these attributes and how they interact with the other 39 lives on the [writing voice attributes](https://scriptgrain.com/reference/writing-voice-attributes) reference page, if you want the complete picture.

Attribute · Values · How obtained · Weight in scoring · Tolerance
formality_score · 0 to 10 · judged · 2 · 3 points
humour_register · dry / sarcastic / self-deprecating / warm / none · judged · 1.5 · adjacent registers 0.5
emotional_expressiveness · low / medium / high · judged · 1 · one step 0.5, two steps 0.1
contraction_rate · per 1,000 words · counted · 2 · larger of 10 or 70% of profile
confidence_vs_hedging · 0 to 1 · judged · 1.5 · 0.35
audience_adaptation · yes / no · extracted · n/a · recorded on the profile

## Counted versus judged

Here's the distinction that does most of the work in this whole page: a prompt describes, a score checks.

When you write a prompt asking for "about 14 contractions per 1,000 words," you're describing a target. The model reads that instruction and can, in principle, obey it mechanically, counting contractions as it writes, stopping when it hits the number. Countable attributes are obedient this way. Contraction rate is counted. Sentence length is counted. Em dash frequency is counted.

Judged attributes don't work like that. Formality is scored by a small model (Claude Haiku in this framework) actually reading the draft and rating it, then that rating gets compared against a target with a tolerance of 3 points and a weight of 2. Confidence versus hedging carries a tolerance of 0.35 and a weight of 1.5. Humour gets a weight of 1.5 too, but with a twist: adjacent pairs score partial credit. Dry and none, dry and sarcastic, warm and self-deprecating all score 0.5 against each other rather than zero, because a model calling something "dry" when a human would call it "sarcastic" isn't as wrong as calling it "warm." Expressiveness runs on its ordered scale at weight 1. Contraction rate, being counted rather than judged, carries weight 2.

The judged subset will vary slightly run to run, because a model is reading it each time and models aren't perfectly consistent readers. That's expected. It's also exactly why a prompt (which only steers the countable stuff) can drift wildly on the judged stuff while looking, on paper, like it followed instructions.

## Formality on a 0 to 10 scale

Formality runs from 0 to 10, and the gap between a 2, a 5 and an 8 is bigger than most people expect.

A 2 reads like this: "Yeah, that's the one that broke last time, just swap it out." Loose, contracted, no ceremony.

A 5 sits in the middle: "That component failed previously; replacing it should resolve the issue." Still plain, but tidied up. Fewer contractions, more complete clauses.

An 8 sounds like this: "The component in question has been identified as the source of prior failure and should be replaced accordingly." No contractions, longer clause structure, formal connective tissue.

Most business writing that "sounds like us" for a genuinely conversational brand sits somewhere between 2 and 4. Most writing that sounds stiff and corporate, even when the writer didn't intend it, drifts up past 6 without anyone noticing. That drift is invisible to the writer and obvious to the reader. It's also, as it turns out, exactly what happens when you ask a general-purpose model to sound informal.

## Confidence versus hedging

Confidence versus hedging runs 0 to 1. Zero is a heavy hedger: "it might be the case that," "this could potentially suggest," "there's some evidence that perhaps." One is fully declarative: the writer just says the thing.

Tolerance here is 0.35, weight 1.5, which is a fairly forgiving band. But the direction of the error matters. A voice that's supposed to sit near 0.8 (direct, plain claims, minimal qualification) reads as a different person entirely if it drifts to 0.4. Readers notice confidence shifts faster than they notice formality shifts, because confidence changes what the sentence is doing, not just how it's dressed.

## Humour and warmth as labels

Humour isn't a number, it's a label: dry, sarcastic, self-deprecating, warm, or none. That's a harder thing to score than formality, because humour depends on context the reader might not have.

The adjacent-pair scoring exists precisely because humour labelling is imprecise even for humans. Dry and sarcastic get confused constantly; a flat, understated line can read as either depending on the reader's mood. Warm and self-deprecating overlap too, since a lot of self-deprecating humour is warm humour aimed inward. Scoring these as 0.5-adjacent rather than flat wrong-or-right is an admission that the label is doing its best, not stating a fact.

Emotional expressiveness (low, medium, high) is a gentler scale by comparison, weight 1, and tends to move in step with formality: high formality nearly always drags expressiveness down towards low, whether the writer meant it to or not.

## Why prompts lose the register

Here's the finding that makes this whole page worth writing. SGR-002, "ChatGPT vs a measured voice" (2026-09-20), took one author's five blog posts, held two back, built a study profile from the other three (2,447 words), then turned each held-out post into a brief. Each brief ran three ways: GPT-5 bare, GPT-5 with the free prompt builder's one-page measured prompt, and generation against the profile itself.

The prompt asked for about 14 contractions per 1,000 words and a formality of 3.7. GPT-5 produced 0.5 contractions per 1,000 words and a judged formality of 7.2, while still obeying the countable rules: no em dashes, more short sentences. It followed the instructions it could follow mechanically and ignored, or simply failed to hit, the instructions that needed judgement to satisfy.

That's the whole argument in one result. A prompt describes intent in words; the model interprets those words however it interprets them. A score checks the output against a target and can be re-run, re-measured, corrected. Mean voice match against the study profile came out 0.83 for the measured generation, 0.74 for bare GPT-5, 0.62 for GPT-5 with the prompt, worse than bare. Against the held-out post itself: 0.91, 0.82, 0.80. The prompt-guided version undercut the plain version. Bare GPT-5 also wrote 2,309 words against a 1,393 target, with 17.5 em dashes per draft; the measured version wrote 1,302 words with none. Full detail, including the detector-score comparison, sits in the [study itself](https://scriptgrain.com/research/chatgpt-vs-a-measured-voice).

## Limits

Judged values carry model variance; the same draft read twice by the same judge won't always score identically. Sarcasm is the hardest label of the five humour categories to pin down, and it's the one most likely to get swapped for "dry" or missed entirely. Register also shifts by audience on purpose in plenty of genuinely good voices, which means audience_adaptation being "true" isn't a flaw to fix, it's sometimes the whole point. None of this is exotic. The [voice, tone and style glossary](https://scriptgrain.com/glossary/voice-tone-style) and the [style analysis tool](https://scriptgrain.com/tools/writing-style-analysis) both work from the same six attributes described here, so the definitions stay consistent wherever you run into them.

## Questions

### What counts as a good formality score for business writing?

There's no universal target; it depends entirely on the brand. Somewhere between 2 and 4 tends to read as approachable and plain-spoken, while anything past 6 starts to feel stiff even to readers who can't articulate why. The point isn't hitting a number, it's matching whatever number your actual voice already sits at.

### Why does contraction rate get counted rather than judged?

Contractions are objectively countable: either "don't" appears or "do not" does. There's no interpretation needed, so a script can tally them directly from the text. That's why contraction rate carries weight 2 and no tolerance debate, unlike formality or humour, which need a reader's judgement and therefore carry variance.

### Can a model reliably detect sarcasm?

Not consistently. Sarcasm is named directly as the hardest of the five humour labels, and the scoring system accounts for this by giving partial credit (0.5) to adjacent pairs like dry and sarcastic. A model, or a human, reading a single passage without wider context will sometimes call sarcasm dry, and sometimes call dry sarcastic.

### Does audience adaptation mean a voice is inconsistent?

No. Audience adaptation asks whether register shifts by audience, and for a lot of genuinely strong voices, it does, on purpose. A voice that sounds slightly more formal in a client email and looser in a blog post isn't broken. The measurement exists to check whether that shift is deliberate and controlled, not to flatten it out.

### Why did GPT-5 miss formality and contractions despite a detailed prompt?

Because the prompt described a target in language, and GPT-5 satisfied the parts it could follow mechanically (no em dashes, shorter sentences) while missing the judged parts entirely. It asked for 14 contractions per 1,000 words and formality 3.7, and produced 0.5 contractions and a judged 7.2. Description isn't the same as measurement against a target.

## Methodology

Attribute definitions and scoring rules from ScriptGrain's profile schema and scoring engine as of 2026-09-20; the prompt-versus-profile figures are SGR-002's as published, two briefs, one author, one run per arm.

## Sources

- [SGR-002: ChatGPT vs a measured voice](https://scriptgrain.com/research/chatgpt-vs-a-measured-voice)
- [ScriptGrain: Writing voice can be measured (the 45 attributes)](https://scriptgrain.com/writing-voice)
- [ScriptGrain API reference (v1)](https://scriptgrain.com/docs/api)
- [ScriptGrain research index (SGR studies)](https://scriptgrain.com/research)
