What we measured ourselves
Original ScriptGrain studies: what we measured ourselves, with the question, sample, method, findings and limitations stated in full.
- SGR-001: Same Brief, Four Voices
Does a measured voice profile materially change the output of the same underlying model?
Three content briefs, each run twice: once through a bare frontier model and once through the same model with a ScriptGrain voice profile applied. The facts stayed identical. The measured style did not.
SGR-001: Same Brief, Four Voices
Does a measured voice profile materially change the output of the same underlying model?
Three content briefs, each run twice: once through a bare frontier model and once through the same model with a ScriptGrain voice profile applied. The facts stayed identical. The measured style did not.
Conducted 2026-07-27, published 2026-07-28.
Sample
Three original content briefs (a banking feature announcement, an opinion piece on productivity, and a B2B piece on replacing status meetings), each run through two arms. A fourth run reused the banking brief against a profile built from a Victorian novelist's prose. (n = 8)
Method
- Arm one: the same frontier model ScriptGrain runs on, called bare, with no system prompt and nothing but the brief.
- Arm two: the ScriptGrain pipeline with a voice profile applied as a hard constraint on generation.
- Outputs were compared on measured style features rather than on preference, so the comparison does not rest on which version anyone liked.
What was measured
- Average sentence length
- Share of sentences under eight words
- Contraction rate
- Em-dash count
- Formatting shape (headings, bolded topic sentences, list use)
- Output length against the brief's word target
Findings
- Sentence length collapsed on the challenger-bank brief: 17 to 11 words. Average sentence length in the profiled arm fell to roughly 11 words from roughly 17 in the bare arm, and sentences under eight words doubled from 24% to 48%.
- Voice lives in the distribution, not the average: 16% to 36%. On the opinion brief the average sentence length was effectively identical across arms at about 15 words, yet short punch sentences went from 16% to 36%. Averaging the two arms together would have shown no difference at all.
- Em dashes vanished: 10 to 0. The bare arm used ten em dashes on the banking brief. The profiled arm used none, which is the single most visible marker separating default model prose from written-by-a-person prose.
- Formatting is part of identity: 708 to 605 words. The bare arm produced a scannable listicle with four headings and fourteen bolded topic sentences, and overran a 600-word brief by 108 words. The profiled arm produced continuous prose with no headings and landed at 605.
- An extreme profile holds under the same facts: 0 contractions in 626 words. The banking brief run against a Victorian-novelist profile carried the identical feature, mechanic and price point, with no contractions, no headings and sums of money spelled out as words.
Limitations
- One run per arm. This is a demonstration of effect direction and size on a handful of briefs, not a statistical result, and no significance test was performed.
- The briefs were written by the study's author, who also builds the product. Neither arm was blind.
- Both arms used the same underlying model, so the study says nothing about how profiles behave across different models.
- A voice profile changes identity, not quality. Nothing here shows that profiled output is better, only that it is measurably different and consistently so.