Brand voice guidelines for marketing agencies: a template with measurable criteria
By Jack Stovell · published 2026-09-21 · checked 2026-09-20
A guideline is only worth writing if a draft can fail it. Most brand voice documents contain nothing a reviewer or a machine could fail a draft on: adjectives like "friendly" and "bold" that mean whatever the reader wants them to mean. This page gives you a template with rows a script can actually score.
Why most guidelines cannot be checked
Here's the thing about "friendly but professional": it's not a criterion, it's a mood board. You can't run a draft against it and get a number out the other end. Nobody fails a review because their copy wasn't "authentic enough". The word doesn't point at anything measurable.
Compare that to "contraction frequency: 21 per 1,000 words" or "average sentence length: 16.9 words, high variance". Those are checkable. A junior writer, a freelancer, or a language model can look at a draft and see whether it landed near the number or drifted away from it. That's the whole point of a guideline: it has to be failable, in the sense that a specific draft can be shown to miss it, with the specific feature named.
Most brand documents skip this because nobody measured the client's actual writing in the first place. They wrote down aspirations instead. This template starts from measurement, using the same 45 attributes across 8 layers that the voice measurement framework defines: lexical, syntactic, tone and register, rhetorical, punctuation and format, function words, content patterns, quirks and cadence. Extraction reads samples in two passes, one per sample and one synthesising across all of them, and it takes roughly one to three minutes once you've got the writing pasted in.
The template
The table below has one row per checkable attribute, grouped by layer, with a column for the measured value and a column for the threshold you're holding drafts to. Some rows come straight from the client's own writing. Others (audience, banned terms) the agency has to write in by hand, because no amount of measurement tells you a client's audience or their list of forbidden words. More on that split in the next section.
Use the table as a working document, not a museum piece. It's meant to be edited every time you get new client writing, and reviewed every time a draft goes out.
| Section | What to write | The measurable criterion | How it is checked |
|---|---|---|---|
| Who the client is | Audience, offer, three things the client would never say | Author-banned terms list | Each hit removes 0.5 similarity; the score names them |
| Register | Formal or conversational, and how sure the voice sounds | formality 0 to 10; confidence 0 to 1 | Judged per draft; tolerance 3 points and 0.35 |
| Sentences | Typical length and how much it varies | avg_sentence_length; sentence_length_variance | Counted per draft; tolerance 70% of the profile average |
| Contractions | Contract or write the words out | contraction_frequency per 1,000 | Counted; tolerance 10 per 1,000 or 70% |
| Punctuation | Commas, semicolons, exclamation marks, dashes, brackets | comma_density; semicolon, exclamation, parenthetical, dash rates | Counted; rare habits at zero on both sides never headlined |
| Person | I, we or you | pronoun_distribution | Counted as shares of I / we / you |
| Words the client reaches for | The recurring vocabulary and fillers | preferred_words; filler_phrases; discourse_markers | Signature-phrase hits, up to four expected per draft |
| How a piece opens and argues | Hook, context or direct; evidence-first or conclusion-first | opening_style; argument_structure; rhythm_pattern | Judged per draft against the profile label |
| Humour and warmth | None, dry, warm or self-deprecating | humour_register; emotional_expressiveness | Judged; adjacent registers score 0.5 |
| Threshold | The score below which a draft goes back | Voice Match | 0.85 on voice; 0.55 to 0.79 drifting; under 0.55 off |
| English variant | British or American | english_variant on the profile | Enforced at generation |
Filling it from a client's writing
So which rows fill themselves in, and which need a human?
The lexical, syntactic, tone, rhetorical, punctuation, function word, content pattern, and quirks rows all come from measurement. Paste in the client's writing, three or more pieces and about 3,000 words is the recommended minimum, and the extraction gives you vocabulary diversity, sentence length and variance, formality score, humour register, comma density, pronoun distribution, the lot. You don't write these rows. You copy them in from the profile and move on.
A single sample of 50 or more words is technically enough to get a profile out. It won't be a good one. Confidence score on a single short sample tends to be low, and the number travels with the profile precisely so you know how much to trust it. Three pieces spread across contexts (a blog post, an email, a product page) gives the extraction enough variance to be useful.
What measurement can't give you: who the client is trying to reach, and what they never want to say. Audience isn't a stylometric property; it's a business decision, and it belongs to the agency and the client together. Banned terms work the same way: a legal team's list of words that can't appear near a product claim, or a founder's pet hate ("we don't say 'solutions'"), doesn't show up in a writing sample no matter how much of it you feed the extraction. Someone has to sit down and ask the client directly, then write the answer into the template by hand.
That's the honest split. Everything measurable, the profile fills in. Everything that's a judgement call about the business, the agency fills in. Mixing the two up, guessing at audience from writing samples, or trying to measure banned terms statistically, produces a template that looks complete and isn't.
Setting the threshold
0.85 is the default. That's the score a draft needs to hit against the client's profile before it counts as "on voice", on the same 0 to 1 scale the scoring runs on: 0.85 and above is on voice, 0.55 to 0.79 is drifting, under 0.55 is off.
But 0.85 isn't gospel, and it shouldn't be treated as one. If a client's own writing varies a lot (different writers, different contexts, a founder who writes loose emails and a marketing team who writes tight landing pages) then their own baseline pieces might not score 0.85 against each other. Holding external drafts to a threshold their own writing wouldn't clear is a good way to fail everything, forever, for no reason.
That's why the threshold is a conversation with the client, not a number you set once and forget. Ask them directly: how consistent do you actually want this to be? A client who wants every touchpoint sounding identical needs a high threshold and probably a tighter set of source samples. A client who's fine with some range between their blog and their sales emails can run lower, maybe 0.75, and treat "drifting" as acceptable rather than a fail state.
Get this wrong in either direction and the guideline stops being useful. Too high, and every draft fails, reviewers stop trusting the score, and everyone quietly goes back to eyeballing it. Too low, and the guideline passes drafts that don't actually sound like the client, which defeats the point of having one.
Reviewing drafts against it
The review loop has four steps, and it always runs in this order: score, read, fix, rescore.
Score the draft against the client's profile first. Don't read it for "voice" with your own ear before you've got a number; that's how you end up disagreeing with your own team about something a script could have settled. The brand voice measurement approach compares the draft to the profile feature by feature, so the output isn't just a single 0.85 or 0.62. It's a list of which specific features moved.
Read the features that moved, not the whole draft again from scratch. If the score dropped because sentence length variance collapsed (every sentence between 20 and 30 words, nothing shorter, nothing punchier) that's specific and fixable. If contraction frequency dropped because the writer got formal under deadline pressure, that's specific too. Reading the whole draft for a vague sense of wrongness wastes time the score already saved you.
Fix those specific features. Not the draft in general. If parenthetical rate is the thing that moved, add or cut parentheticals; don't rewrite paragraphs that were already fine.
Rescore. If it clears the threshold, ship it. If it doesn't, you'll now see exactly which feature is still off, and you fix that one too. This is the loop, every time, and it's faster than it sounds once a team has done it a handful of times.
Keeping it current
Client writing changes. New hires write differently to the person they replaced; a rebrand shifts formality up or down; six months of a founder writing every newsletter personally will drift the whole voice toward however that founder writes. A profile built in January doesn't necessarily describe the client in July.
The fix is to rebuild the profile when the writing changes, not on some fixed calendar you picked at random. Drift is measurable month to month, the same way a draft's drift from the profile is measurable. Agencies handling more than one client tend to want this checked in one place rather than profile by profile, which is what the multi-client voice drift check is for.
Rebuilding isn't a big job. It's the same extraction step you ran the first time: paste in the client's more recent writing, three or more pieces, let it run its one to three minutes, and compare the new profile to the old one. If the numbers haven't moved much, you're done, and the old thresholds still hold. If they have, the glossary entry on voice drift is worth having on hand when you're explaining to a client why the guideline they signed off on six months ago needs an update now. Agencies running this across a client roster are the exact case the agency plan is built for: Studio holds 10 client profiles across 3 seats at £99 a month, Agency holds unlimited profiles across 10 seats at £299.
Limits
The template measures how a client writes, never what they should say: audience, positioning and banned terms are business decisions the agency writes in by hand. A profile built from press releases or a previous agency"s copy is a profile of that agency. Short formats under about 120 words do not carry enough sentence structure to score, and a client with several deliberate voices needs several profiles rather than one averaged target."
Questions
What counts as a measurable brand voice attribute?
Anything with a number or a fixed category attached: sentence length, contraction frequency, formality score, humour register, comma density, pronoun distribution. If two different reviewers (or a script) can independently check a draft against it and get the same answer, it's measurable. "Friendly" and "on brand" fail this test; "0.85 or above" doesn't.
Can a brand voice guideline be built from one piece of client writing?
Technically, yes: the extraction needs only one sample of 50 or more words to produce a profile. In practice it won't be reliable. Confidence score tends to be low on a single short sample, which is exactly why three or more pieces, roughly 3,000 words total, is the recommended minimum for a guideline you intend to hold drafts to.
Why isn't audience part of the measured template rows?
Audience is a business decision, not a property of the writing itself. Two clients could write in near-identical styles while targeting completely different readers. The extraction measures how someone writes; it has no way of knowing who they're trying to reach. That row has to be filled in by the agency, in conversation with the client, every time.
Should every client use the same 0.85 threshold?
No. 0.85 is the sensible default, but a client whose own writing varies a lot across contexts or writers might not clear 0.85 against their own baseline pieces. Setting the threshold is a conversation with the client about how consistent they actually want to be, not a fixed rule applied identically everywhere.
How often should a client's voice profile be rebuilt?
Whenever their writing has actually changed, not on a fixed schedule. New hires, rebrands, and shifts in who's writing the copy all move a profile over time, and that drift is measurable month to month. Agencies managing several clients tend to check this on a recurring basis rather than waiting for a draft to fail unexpectedly.
Methodology
The template maps each section of a conventional brand voice document to the attribute ScriptGrain measures and the rule the scoring engine applies (as of 2026-09-20). It was written by ScriptGrain; no third-party template is quoted.