# Keeping brand voice consistent across several writers: a measured rollout in four steps · ScriptGrain

> How to keep one brand voice across several writers when a style guide has stopped working: build the profile, set a threshold, score every draft and read the spread across writers as the consistency number. Four steps, with the numbers.

Canonical: https://scriptgrain.com/blog/keep-brand-voice-consistent-across-several-writers

# Keeping brand voice consistent across several writers: a measured rollout in four steps

*By Jack Stovell · 2026-09-24 · Guides*

Brand voice stays consistent across several writers when every draft gets scored against a measured [profile](https://scriptgrain.com/glossary/brand-voice), and the team watches the spread of scores across writers rather than trusting individual judgement. A written guide can't do that. It can only describe a voice. It can't tell you when someone's drifted from it.

*Method: 45 attributes across 8 layers (lexical, syntactic, tone, rhetorical, punctuation and format, among others), extracted from the profile schema at scriptgrain.com/writing-voice. Full scoring method with tolerances published at [/reference/voice-measurement-framework](https://scriptgrain.com/reference/voice-measurement-framework).*

## Why the style guide stopped working

Here's the thing about style guides: they're written once, read twice, and then ignored. Not out of malice. Out of drift. A guide says "keep sentences short and confident." Fine. But how short? Confident compared to what? Nobody agrees, and nobody checks.

So writers interpret the guide through their own instincts. One tightens everything into fragments. Another hedges every claim because that's how they write emails. Six months in, you've got five people producing five voices, each one convinced they're following the same document. And to be fair, they are. The document just doesn't say anything measurable.

That's the gap. A guide describes tone in adjectives: warm, punchy, direct. A profile scores it in numbers: sentence length, contraction rate, formality, confidence versus hedging. One of these can tell you a draft has drifted before it publishes. The other tells you after a reader notices.

## Step one: build the profile from the writing, not the deck

Start with the actual writing, not the brand deck. Pull the best existing copy, the pieces everyone agrees sound right, and extract the profile from that corpus directly. The deck describes intentions. The writing shows what's actually happening across all 45 attributes: lexical choices, sentence rhythm, punctuation habits, tone, argument structure.

This matters because decks lie by omission. A deck might say "we're conversational," but it won't tell you the brand uses contractions at a rate of 21 per thousand words, or that semicolons basically never appear, or that parentheticals show up roughly twice every 300 words. Writing has all of that baked in already. You just have to measure it.

Build the profile once you've got a real sample, ideally several thousand words minimum. Too little writing and the profile picks up noise instead of signal (more on that later, under where this breaks). Once it exists, it travels by API and MCP, read live, so any tool checking a draft against it gets the current version, not a snapshot from last year.

## Step two: agree the threshold

Pick a number and commit to it before anyone starts writing against the profile. The default is 0.85 on voice-match. That's the calibrated score where a draft counts as genuinely on-voice: not close, not "basically fine," actually matching across the weighted attributes.

Some brands drop this to 0.80. That's for a brand whose own writing varies naturally, where forcing 0.85 would mean rejecting perfectly good copy because the source material itself wasn't tightly consistent. Fair enough. The threshold should reflect the brand's real range, not an aspirational one.

Below 0.45, a draft is off-voice, full stop. Between 0.5 and 0.8, it's partial: recognisable but not there yet. The weighting behind this isn't flat either. Sentence length, contractions, formality, signature phrases and banned terms each carry a weight of 2. Variance, commas, pronoun mix, confidence, humour, argument structure and rhythm carry 1.5. Agree these numbers with the team once, in writing, and stop relitigating them every time a draft scores 0.79 and someone feels precious about it.

## Step three: score every draft and name what moved

Every draft gets scored before it goes anywhere near a publish button. Not spot-checked. Every one. This is the step that actually replaces the style guide, because a score doesn't argue and doesn't have a bad day.

When a draft comes back under threshold, the scoring names what moved. Maybe sentence length crept up from 17 words average to 24. Maybe contractions dropped because a writer was tired and defaulted to formal habits. Maybe the confidence-versus-hedging score fell because three paragraphs in a row started with "it could be argued that." You get the specific attribute, not a vague "this feels off" note from an editor with limited time.

That specificity is the whole point. SGR-001, "Same Brief, Four Voices" (28 July 2026), ran the same banking brief bare and through a profile on the same model. Unedited, one run per arm. Average sentence length fell from about 17 words to about 11. Sentences under eight words doubled, from 24% to 48%. Contractions rose by roughly 70%. Ten em dashes became none. That's not a vibe shift. That's a set of measurable, nameable changes, and it's exactly the kind of drift a style guide would never have caught, because a style guide doesn't count anything.

Onboarding a new writer works the same way. Hand them the profile's narrative and its attributes rather than a slide deck, and score their first three drafts. You'll know within three pieces whether they've picked up the voice or whether they need a second look. That's faster than months of editorial back-and-forth guessing at "does this sound right."

## Step four: read consistency as a spread

Here's where most teams stop too early. They look at one writer's average score, see 0.87, and call it done. But consistency across a team isn't measured by one person's mean. It's measured by the spread across all of them.

Score each writer's last few pieces against the brand profile. The mean per writer tells you who's closest to the brand voice individually. The spread across writers, the gap between your tightest and loosest performer, is the actual consistency figure the team should be watching. A brand can have five writers all individually scoring above 0.85 and still have a consistency problem if one of them sits at 0.86 and another sits at 0.95. That's not "everyone's fine." That's a nine-point spread, and readers notice spreads even when they can't name why.

This is also why the profile isn't static. It can be rebuilt as the brand's writing evolves, and a monthly monitor scores live pages so drift gets caught after publication too, not just before. Watching the spread monthly, rather than the mean once, is what actually keeps five writers sounding like one brand instead of five people loosely inspired by the same brief.

## Where this does not work

None of this is universal, and pretending otherwise would be dishonest.

Short formats under 120 words don't give the scoring enough to work with. A tweet or a one-line meta description doesn't contain enough sentence structure, punctuation variety, or rhythm for the attributes to mean anything. You're measuring noise at that length, not voice.

A brand with several deliberate voices, say, a playful social account and a formal investor-relations arm, needs multiple profiles, not one. Forcing both into a single measured voice will flatten the deliberate distinction the brand actually wants. That's not drift. That's design, and the scoring should respect it by running separate profiles rather than averaging them into mush.

Translated copy breaks the model too. The attributes here (sentence length, contraction rate, article balance) are calibrated to the source language's mechanics. Translate a profile-matched English draft into French and the numbers won't transfer, because French doesn't use contractions the same way English does. You'd need a profile built natively from French writing, not a translated one.

And a profile built from too little writing just won't hold. If the source corpus is a few hundred words, the "voice" you extract is mostly coincidence, not pattern. [Voice drift](https://scriptgrain.com/glossary/voice-drift) creeps in fast when the baseline itself was never stable to begin with.

The [brand voice analyser](https://scriptgrain.com/tools/brand-voice-analyser) is free for prospects and compares any page to a reference piece with no account needed, which is a reasonable way to sanity-check whether a format or a language pair is even a good candidate before committing to a full profile. Teams building this out properly, especially agencies managing several client voices at once, tend to start from the [template](https://scriptgrain.com/reference/brand-voice-guidelines-template) and read the underlying method at [/reference/brand-voice-measurement](https://scriptgrain.com/reference/brand-voice-measurement) before rolling anything out across writers. Agencies specifically have their own considerations, covered separately at [/for-agencies](https://scriptgrain.com/for-agencies).

## Questions

### How many samples do you need to build a profile?

Enough that the pattern isn't coincidence: a few thousand words of representative writing is the practical floor. Less than that and the extraction picks up quirks from one piece rather than genuine habits. More matters less than variety: several different formats beat one long document of the same type.

### What threshold should we use?

0.85 by default. Drop to 0.80 only if the brand's own source writing genuinely varies enough that 0.85 would reject good copy. Agree the number with the team before scoring starts, and don't renegotiate it every time a single draft lands just under it.

### Does this replace editors?

No. It replaces the guesswork editors currently do around tone, which a score handles faster and more consistently. Editors still judge accuracy, argument quality and whether the thing actually says something. The score just tells everyone, editor included, whether the voice matches before that conversation starts.

### What happens when a draft scores in the partial range?

Between 0.5 and 0.8, the draft is recognisable but off in specific, named ways: sentence length too long, confidence too hedged, contractions too low. That naming is the useful part. A writer fixing three named attributes gets back on-voice faster than one told vaguely to "make it sound more like us."

### Can the profile change over time?

Yes, and it should. Brand writing evolves, new formats appear, and a profile frozen at launch will eventually mismatch the writing it's meant to measure. Rebuild it periodically from current best-performing copy, and the monthly monitor will keep flagging live pages against whichever version is current.

[More from the ScriptGrain Journal](https://scriptgrain.com/blog)
