Voice profile JSON schema: 45 attributes across 8 layers, as the API returns them

By Jack Stovell · published 2026-09-21 · checked 2026-09-20

A voice profile is one JSON object of 45 named attributes across eight layers, returned by the API and the MCP server. Fields are typed as numbers, strings, string arrays, share objects and booleans. Here's the schema, field by field, plus how to read it and use it downstream.

The response shape

Here's the thing about the response: it's one envelope, and the shape doesn't change whether you hit the API directly or call it through MCP. `GET /v1/profiles/{id}` returns a `data` object and a `meta` object. Inside `data` you get `id`, `name`, and `status`, which is one of "processing", "complete" or "failed". While it's running, `progress_percent` sits somewhere between 5 and 95. Once it's done, you get `confidence_score` (0 to 1), `sample_count`, `total_word_count`, and `english_variant`, which is either "uk" or "us".

There's also `created_at`, a `narrative` field (a prose description of the voice, written for humans), and `attributes`, which holds the 45 fields this page is mostly about. `meta` just carries a `request_id`. That's the whole envelope.

Types are straightforward once you see them. Numbers are JSON numbers. Labels are strings. Lists are arrays of strings. Shares are objects of numbers. `pronoun_distribution` might read `{ "I": 0.48, "we": 0.22, "you": 0.30 }`. `the_a_an_ratio` might read `{ "the": 0.48, "a": 0.38, "an": 0.13 }`. Yes/no fields are booleans, plain and simple.

One wrinkle worth flagging: profiles extracted before 2026-08-17 sometimes hold descriptive text in fields that are meant to be labels. The scorer tolerates that, so nothing breaks. But new extractions write canonical labels, so if you're building against fresh profiles you won't hit this at all. The MCP tool `get_profile` returns the identical object, so anything you build against the REST response works unchanged through MCP.

The attributes object, field by field

The 45 fields sit in eight layers, and the grouping isn't arbitrary. It reflects how the extraction actually works: two passes, one per sample and one synthesis across all samples, each layer capturing a different kind of pattern.

Lexical covers six fields: `vocabulary_diversity_index` (0 to 1), `preferred_words`, `word_length_distribution` (short 1 to 4 letters, medium 5 to 7, long 8+), `rare_word_rate`, `filler_phrases`, and `contraction_frequency`.

Syntactic covers seven: `avg_sentence_length`, `sentence_length_variance`, `complexity_preference`, `clause_ordering`, `parenthetical_rate`, `dash_frequency`, `semicolon_frequency`.

Tone and register covers six: `formality_score`, `humour_register`, `emotional_expressiveness`, `contraction_rate`, `confidence_vs_hedging`, `audience_adaptation`.

Rhetorical covers six: `opening_style`, `closing_pattern`, `metaphor_usage_rate`, `repetition_as_emphasis`, `transition_style`, `argument_structure`.

Punctuation and format covers six: `comma_density`, `exclamation_rate`, `ellipsis_usage`, `question_mark_in_body`, `capitalisation_quirks`, `list_preference`.

Function words covers five: `the_a_an_ratio`, `but_however_yet_preference`, `pronoun_distribution`, `discourse_markers`, `paragraph_opener_words`.

Content patterns covers four: `specificity_level`, `anecdote_usage_rate`, `claim_density`, `analogy_preference`.

Quirks and cadence covers five: `consistent_misspellings`, `rhythm_pattern`, `paragraph_length_preference`, `power_sentence_position`, `structural_signatures`.

That's 45. The table attached to this section lists each field against its layer, type and unit, so you don't have to hold all of it in your head. For the plain-English meaning of each attribute, and what a high or low value actually implies about someone's writing, see Writing voice attributes.

FieldJSON typeLayerUnit or values
vocabulary_diversity_indexnumberLexical0 to 1
preferred_wordsstring[]Lexicalrecurring distinctive words
word_length_distributionobject of number sharesLexicalshort (1 to 4 letters) / medium (5 to 7) / long (8+)
rare_word_ratenumberLexicalrare words per 1,000
filler_phrasesstring[]Lexicalcatalogued fillers
contraction_frequencynumberLexicalcontractions per 1,000 words
avg_sentence_lengthnumberSyntacticwords
sentence_length_variancenumberSyntacticvariance of sentence lengths
complexity_preferencestringSyntacticsimple / compound / complex / mixed
clause_orderingstringSyntacticfront-loaded / build-to-point / mixed
parenthetical_ratenumberSyntacticper 300 words
dash_frequencynumberSyntacticper 1,000 words
semicolon_frequencynumberSyntacticper 1,000 words
formality_scorenumberTone and register0 to 10
humour_registerstringTone and registerdry / sarcastic / self-deprecating / warm / none
emotional_expressivenessstringTone and registerlow / medium / high
contraction_ratenumberTone and registerrate
confidence_vs_hedgingnumberTone and register0 (heavy hedger) to 1 (declarative)
audience_adaptationbooleanTone and registerregister shifts by audience
opening_stylestringRhetoricalhook / context / direct / anecdote / question
closing_patternstringRhetoricalhow pieces end
metaphor_usage_ratenumberRhetoricalper 1,000 words
repetition_as_emphasisbooleanRhetoricalrepeats for effect
transition_stylestringRhetoricalhow paragraphs connect
argument_structurestringRhetoricalevidence-first / conclusion-first / narrative
comma_densitynumberPunctuation and formatper sentence
exclamation_ratenumberPunctuation and formatper 1,000 words
ellipsis_usagestringPunctuation and formatfrequent / occasional / rare / never
question_mark_in_bodybooleanPunctuation and formatquestions in body copy
capitalisation_quirksstringPunctuation and formatany noted
list_preferencestringPunctuation and formatbullets / numbered / inline / mixed / avoids
the_a_an_ratioobject of number sharesFunction wordsthe / a / an
but_however_yet_preferencestringFunction wordsbut / however / yet / mixed
pronoun_distributionobject of number sharesFunction wordsI / we / you
discourse_markersstring[]Function wordsconnecting phrases
paragraph_opener_wordsstring[]Function wordsrecurring first words
specificity_levelstringContent patternsabstract / balanced / data-driven
anecdote_usage_ratenumberContent patternsper 1,000 words
claim_densitynumberContent patternsclaims per paragraph
analogy_preferencebooleanContent patternsreaches for analogies
consistent_misspellingsstring[]Quirks and cadencehabitual spellings
rhythm_patternstringQuirks and cadencepunchy / flowing / mixed
paragraph_length_preferencestringQuirks and cadenceshort / medium / long / varied
power_sentence_positionstringQuirks and cadencestart / end / both / varied
structural_signaturesstringQuirks and cadencenotable patterns

Types and units

So, the units. Rates are per 1,000 words unless the field says otherwise. `rare_word_rate`, `metaphor_usage_rate`, `anecdote_usage_rate`, `exclamation_rate`, `dash_frequency`, `semicolon_frequency`, `contraction_frequency`: all per 1,000. `parenthetical_rate` is the one exception in that family, measured per 300 words rather than per 1,000, so watch that when you're comparing it against the others.

`formality_score` runs 0 to 10. `confidence_vs_hedging` runs 0 to 1, where 0 is a heavy hedger and 1 is fully declarative. `vocabulary_diversity_index` also runs 0 to 1. `comma_density` is per sentence, not per thousand words, which trips people up if they're skimming.

Labels like `humour_register`, `opening_style`, `complexity_preference` and `clause_ordering` are enumerated strings. `humour_register` takes dry, sarcastic, self-deprecating, warm or none. `opening_style` takes hook, context, direct, anecdote or question. Booleans like `repetition_as_emphasis`, `analogy_preference` and `audience_adaptation` are exactly that: true or false, no middle setting.

Share objects deserve their own mention. `pronoun_distribution` and `the_a_an_ratio` are the two you'll see most, and their values sum to roughly 1 across the object's keys. Treat them as proportions, not counts. If `pronoun_distribution` shows `we` at 0.22, that's 22% of the counted pronouns, not 22 uses per thousand words.

Reading confidence_score

A low `confidence_score` almost always means one of two things: too few samples, or samples that are too short, or samples pulled from mixed genres that don't share a consistent voice. Sometimes it's all three at once. The product recommends three or more pieces and roughly 3,000 words total for extraction, and that recommendation exists because of what the two-pass process needs to work with.

Pass one reads each sample separately and measures it. Pass two synthesises one profile from everything pass one found. A single 300-word blog excerpt will still produce a profile (the minimum is one sample of 50+ words), but the synthesis has almost nothing to reconcile, so confidence sits low and the numbers can be noisy. Feed it three pieces spanning 3,000 words and pass two has actual disagreement to resolve, patterns that hold across samples versus patterns that were just one piece having an odd day. That's what confidence is really measuring: consistency across samples, not correctness of any single one.

Extraction itself takes roughly one to three minutes regardless of sample count, so there's no time cost to doing this properly. If your `confidence_score` comes back low, the fix is almost never to retry the same samples. It's to add more of them, and to try to keep them within one genre if you can.

Using the attributes downstream

Here's the thing about consuming this data: you shouldn't reimplement scoring yourself. `/v1/voice-match` exists precisely so you don't have to write your own distance function across 45 heterogeneous fields, some of which are shares, some booleans, some enumerated strings with no natural ordering. Send it a profile id and the text to score, and let it do the comparison. Full detail on that endpoint, including request and response shape, lives at Brand voice API.

If you're building a system prompt from a profile rather than scoring against one, the counted fields are your raw material. `avg_sentence_length`, `contraction_rate`, `formality_score`, `humour_register`, `opening_style`, `closing_pattern`: these translate almost directly into instructions. The narrative field is worth including too. It's already prose, already synthesised, and it often captures interactions between attributes that a flat list of numbers won't show on its own.

If you're running this across an agency or a team, store profiles per client rather than per project. Voice tends to belong to the client, not the individual piece of content, and a per-client profile lets you keep sample counts climbing over time, which only helps confidence. For the tool-level mechanics of calling any of this through an agent or assistant, see the MCP tool reference; for the broader picture of how the API and MCP server fit together, see AI writing API and MCP server.

Limits

The schema describes the fields, not their accuracy: the labelled attributes are a model"s reading of the samples and can differ between extractions of the same text, and every rate is only as stable as the sample size behind it. Older profiles may carry descriptive text in label fields. The example values on this page are illustrative shares, not a real person's profile."

Questions

What's the minimum data needed to generate a voice profile?

One sample of 50 or more words will produce a profile. It won't be a strong one. Confidence depends on volume and consistency, and a single short sample gives pass two almost nothing to synthesise against. The product recommends three or more pieces and about 3,000 words total, which gives the extraction enough material to separate genuine patterns from one-off quirks.

Why does parenthetical_rate use a different unit from the other rate fields?

Most rate fields in the schema are counted per 1,000 words, but `parenthetical_rate` is measured per 300 words instead. It's a deliberate exception, not an inconsistency. Parentheticals are rarer events than, say, commas or contractions, so a per-300 window keeps the resulting numbers legible rather than producing tiny decimals across most real samples.

Can I compare two profiles' share objects directly?

Yes, and that's exactly what `/v1/voice-match` does for you. Share objects like `pronoun_distribution` and `the_a_an_ratio` sum to roughly 1, so comparing them means comparing distributions, not raw counts. You can eyeball two profiles' share objects side by side, but for anything scored or automated, use the match endpoint with the text rather than writing your own comparison logic.

Does the schema change between the REST API and MCP?

No. `get_profile` through MCP returns the same object as `GET /v1/profiles/{id}` through the REST API: same envelope, same 45 attributes, same types. If you've built logic against one, it works unchanged against the other. That consistency is deliberate, so you can prototype through MCP and move to direct API calls later without touching your parsing code.

What should I do if confidence_score stays low after adding samples?

Check genre consistency first. Mixed sources (a technical document alongside a casual social post) can keep confidence low even with plenty of words, because pass two is finding real disagreement between samples rather than noise. Narrow the samples to one genre or register, keep the total around 3,000 words or more, and confidence should climb. If it doesn't, the voice itself may genuinely vary by context, which the `audience_adaptation` field will usually confirm.

Methodology

Field names, types and values are read from the profile schema and the api-v1 response as of 2026-09-20 (ScriptGrain repository); the example values are illustrative shares, not a real person's profile.

Sources