# How to evaluate a brand voice AI tool: 12-point checklist · ScriptGrain

> Twelve questions to put to any brand voice AI tool before buying, each with a pass condition, plus three test prompts to run in every tool on the shortlist and a way to score the result that does not depend on a demo.

Canonical: https://scriptgrain.com/reference/how-to-evaluate-brand-voice-ai-tool

# How to evaluate a brand voice AI tool: a twelve-point checklist with test prompts

*By Jack Stovell · published 2026-09-24 · checked 2026-09-20*

Every brand voice tool claims consistency. That word does no work on its own. The twelve questions below are the ones a demo cannot answer with adjectives, and each one has a condition a tool either meets on its own pages or does not.

## Why most demos cannot answer these

A demo is a performance. Someone picks a good brief, runs it once, and the output sounds plausible. That's not evidence, that's staging.

Here's the thing: a real evaluation needs a number, a published method, or a named list of attributes. Adjectives like "authentic" or "on-brand" cost nothing to say and mean nothing on their own. So the twelve points below all ask the same underlying question in different clothes: can you check this claim on the vendor's own pages, without booking a call?

Across the 2026-09-20 harvest of 61 tools and 14,792 pages (13,581 classified), the finding was blunt. No competitor page publishes a scoring method for voice, an attribute list, or before-and-after numbers on one brief. Not one, out of thousands of pages. That absence is the whole reason this checklist exists.

## The twelve points

Run through these against any tool's own site before you touch a demo. A vendor that meets a point will usually say so plainly, in text you can screenshot.

Does it publish a scoring method, not just a score? Does it score every draft, or only a one-off site diagnostic? Does it publish a named attribute list (and how many attributes)? Does it show before-and-after numbers on a single brief? Can it hold more than one brand profile at once? Does it support British and American English as a profile setting? Is there a public API? Is there an MCP server, and on which plans? Does pricing include the scoring feature, or is it an upsell? What's the minimum sample size it asks for, and does it recommend more? Are samples stored, and can you delete them? Are samples used to train models, or just to build your profile?

Most tools answer two or three of these on their own pages. VoiceMoat, for instance, publishes a 0 to 100 voice score, but no published method behind it: point 1 fails even though a score exists. Noren skips scoring entirely and hands you a Markdown profile to paste into any LLM: no score, no API, no MCP, three points gone at once. Jasper has a Brand Score, but it's a site diagnostic, not a per-draft method: point 2 fails there too.

To be fair, some of these twelve are genuinely hard to publish. An attribute list means admitting what you actually measure, and that's a commitment most vendors would rather avoid making in public.

## Three test prompts to run in every tool

Skip the sales call. Run these three prompts yourself, in the tool's own interface, and write down what happens.

First, take the same 200-word brief and run it twice: once with the brand voice applied, once without. Count what actually changed. Word choice? Sentence length? Structure? If the two outputs read almost identically, the voice feature isn't doing anything measurable.

Second, paste in a real published piece from the brand and ask the tool how close it is. Expect a number back. "Pretty close" is not an answer; a tool with a genuine scoring method will give you a figure, because that's what scoring means.

Third, ask the tool for the list of attributes it measured to produce that number. If it can't name them, it didn't measure anything. It guessed, and dressed the guess in a percentage.

These three prompts take about ten minutes each. That's the whole test. Anyone selling you a longer onboarding process before you can run them is selling you delay.

## Scoring the result

Count the points met from the vendor's own pages, not the sales call. That distinction matters more than it sounds like it should, because a salesperson can promise anything verbally with zero comeback if it turns out untrue.

Any tool meeting fewer than six of the twelve points is selling a description, not a measurement. Six is the line, not because it's a round number, but because below it you're relying on the tool's own confidence rather than anything it can show you.

The [reference table comparing tools](https://scriptgrain.com/reference/ai-writing-voice-tools-matrix) lays out prices, scores, published methods and API and MCP access for each named competitor, pulled from their own pages on the dates given. Reading it, a pattern shows up fast: scoring and MCP support cluster together, and most tools that skip one skip both. ContentIn, Taplio and Pressmaster all ship MCP servers but no score. DoppelWriter and Rytr have neither. Acrolinx scores content against a style guide and publishes an API, but no MCP page turned up on its site.

If you want the underlying logic behind the twelve points rather than just the checklist, that's laid out at [the voice measurement framework](https://scriptgrain.com/reference/voice-measurement-framework), and the attribute categories themselves get defined at [writing voice attributes](https://scriptgrain.com/reference/writing-voice-attributes).

## What ScriptGrain answers

So, plainly, here's what we publish. ScriptGrain scores every draft from 0 to 1 across 45 attributes organised into 8 layers, and the method behind that scoring is published, not held back for a sales call. That's points 1, 2 and 3 answered on the page, not in a meeting.

Profiles are per-client: 10 included on Studio, unlimited on Agency, with extra seats at £15 a month on Studio and £25 a month on Agency beyond those included in the plan. British or American English gets set directly on the profile. A public API and MCP server ship on every plan, including Free. That covers points 5 through 8 without exception across the pricing tiers.

On samples: we recommend three or more pieces and around 3,000 words for a strong profile, though the minimum is one sample of 50 words. Samples are stored to build the profile, and they're deletable along with it. They are not used to train models. Points 10, 11 and 12, answered directly.

Where it doesn't fit: there's no scheduling built in, and no enterprise governance layer. If you need either of those, this isn't that tool, and we'd rather say so here than have you find out three weeks in.

Pricing runs in pounds and dollars: Writer at £12 a month, Studio at £99 for 10 profiles and 3 seats, Agency at £299 for unlimited profiles and 10 seats. Full detail sits on the [pricing page](https://scriptgrain.com/pricing), and the scoring itself can be tried directly through the [brand voice analyser](https://scriptgrain.com/tools/brand-voice-analyser).

## Limits

A checklist built from published pages rewards vendors who publish and penalises those who sell by demo, which is a bias worth knowing: an enterprise tool may meet a point in a contract it never puts on a web page. The three test prompts measure what you can see in ten minutes, not integration depth, governance or support. ScriptGrain wrote the checklist, and it is the tool that meets most of it; read the matrix and the method pages and decide for yourself.

## Questions

### What counts as a published scoring method?

A method is published when the vendor's own pages describe how a score gets calculated, not just that a score exists. VoiceMoat shows a 0 to 100 score with no published method behind it, which fails this test. A published method names the attributes measured and explains, in text anyone can read, how a draft ends up with its number.

### Why does the sample size a tool asks for matter?

A tool built on one short sample is guessing at patterns that need repetition to prove themselves. Rytr's My Voice page runs to 173 words with no detail on sample requirements. ScriptGrain accepts one sample of 50 words as a technical minimum but recommends three or more pieces, roughly 3,000 words, before treating the profile as reliable.

### Is an MCP server the same thing as an API?

No. An API is a general connection point other software can call. MCP is a specific protocol that lets AI assistants query a tool directly and pull live data, rather than working from a static export. ScriptGrain ships both on every plan including Free; several competitors, including Taplio and Pressmaster, ship MCP without a documented API.

### Does a high voice score mean the writing is good?

No, and that distinction matters. A voice score measures closeness to a defined brand profile across named attributes, nothing about quality in general. A technically strong piece can still score low if it drifts from the brand's usual sentence length, vocabulary or structure. Scoring measures fit, not merit; the two overlap sometimes but aren't the same thing.

### Why run the three test prompts before a demo call?

Because a demo shows you one curated example, chosen to look good. The three prompts (a before-and-after brief, a real published sample scored against the brand, and a request for the attribute list) take about ten minutes total and can't be staged the same way. They force the tool to show its working, rather than its best day.

## Methodology

The checklist was written by ScriptGrain on 2026-09-20 from the gaps found across 13,581 competitor pages; the competitor facts cited are from those vendors' own pages on the dates given in the tools matrix.

## Sources

- [ScriptGrain: AI writing voice tools compared (the matrix)](https://scriptgrain.com/reference/ai-writing-voice-tools-matrix)
- [ScriptGrain: Writing voice can be measured (the 45 attributes)](https://scriptgrain.com/writing-voice)
- [ScriptGrain API reference (v1)](https://scriptgrain.com/docs/api)
- [ScriptGrain research index (SGR studies)](https://scriptgrain.com/research)
