Do brands in the same sector sound the same? 89 UK brands' websites, measured
SGR-009, conducted 2026-10-04. Analyst: Jack Stovell, founder, ScriptGrain.
Question: Measured with a pairwise voice score, are two pages from the same brand's website more alike than a page from that brand and a page from a rival in the same sector?
Summary
Background. Brand guidelines promise a distinctive voice, and many readers say company websites in the same sector all sound alike. ScriptGrain's pairwise voice score treats one page as the baseline and scores others against it. A pilot on 101 pages from 26 brands found it told a brand's own pages from a rival's barely better than chance (AUC 0.54), and that two rivals' homepages scored closer than a homepage did to its own inner pages. This study tests both findings on new brands chosen before any page was read.
What we did. We fixed three hypotheses, a panel of 96 UK brands in 10 sectors and the rules for picking five page types (homepage, about, product, pricing, careers) and published them before fetching anything. 89 brands had at least two readable pages, giving 299 pages and 89,102 ordered page pairs. Every pair was scored with the production computation; each page's qualitative attributes were judged by Claude Haiku 4.5 at temperature 0, as in the product, and again by Claude Sonnet 5.5 in three separate runs.
What we found. Two pages from the same brand scored higher than a page pair from rivals in the same sector 55% of the time (AUC 0.551, 95% interval 0.521 to 0.582), where 50% is a coin toss. 26% of a brand's own page pairs read "on voice" and 24% read "off voice"; for rivals in the same sector, 18% read on voice. The pilot's second finding did not hold: same-type pages from rival brands scored no closer than different pages from one brand (median 0.657 against 0.654). Every judge run and the counted features alone gave the same answer.
What it means. Measured this way, a UK brand's website pages are only slightly more like each other than like a rival's, in every sector but one. Either brands in a sector write very alike, or this kind of measure reads little brand voice in website copy, or both; the study cannot separate the two. It describes one measure on one set of pages; it does not rate any brand's writing.
Key numbers
- Scored with a pairwise voice measure, two pages from the same UK brand's website were more alike than two pages from rival brands in the same sector only 55% of the time, against 50% for a coin toss (ScriptGrain, 2026; 89 brands, 299 pages).
- 24% of pairs of pages from the same brand's website scored "off voice" against each other, and only 26% scored "on voice" (ScriptGrain, 2026).
- A page from a rival brand in the same sector scored "on voice" against a brand's own page in 18% of pairs, and a page from a brand in another sector in 12%.
- Pages of the same type from rival brands, such as two banks' pricing pages, scored no closer than two different pages from one brand: a median of 0.657 against 0.654 on the raw scale.
- By sector, energy brands were the easiest to tell apart from their rivals (AUC 0.66) and banks the hardest (0.49, no better than chance).
- A larger AI judge (Claude Sonnet 5.5, three runs) disagreed with the production judge on 45% of its readings of tone and structure, yet gave the same answer: AUC 0.558, 0.564, 0.558.
- Counted features alone (sentence length, punctuation, pronouns, contractions), with no AI judge, told brands apart as well as the full score: AUC 0.553 against 0.551.
Sample
A panel of 96 UK consumer brands in 10 sectors (10 per sector, 6 in energy), frozen before fetching, none from the pilot and no two in a sector sharing a parent company. For each brand: the homepage plus the shortest URL matching fixed rules for an about, product, pricing and careers page, fetched on 4 October 2026 from a UK location and reduced to text with the product's own page reader, cut to the first 700 words. 89 brands had at least two readable pages.
- Page found by the frozen URL rules for its type on the brand's own domain or a subdomain
- At least 150 words after the product's page reader removed scripts, styles and markup
- A brand is included if at least two of its pages qualify
Method
Pre-registration: hypotheses H1 to H3, the brand panel, the page-type rules and both judges were committed to the public repository before any page was fetched (docs/research/sgr-009-preregistration.md and scripts/experiments/sgr-009/panel.json).
Pages: Firecrawl's site map of each brand, the page-type rules applied to its links (shortest URL first, up to three tries per type), Firecrawl's raw HTML from a UK location, and ScriptGrain's own page reader, exactly what the product reads when given a URL.
Judges: every page judged with the production judge prompt. Primary: Claude Haiku 4.5 at temperature 0 with the production limits, run twice. Robustness: Claude Sonnet 5.5, run three times (its temperature cannot be set). No judge call failed.
Scoring: every ordered pair of included pages through the production computation, once per judge run and once with counted features only.
Tests: AUC and difference in medians with 95% intervals from 2,000 bootstrap resamples of brands within sectors (fixed seed). Analyses beyond H1 to H3 are labelled exploratory.
Variables measured
- Pairwise voice score (raw and band)
- AUC, same brand against rival brands
- Counted features per page
- Judged attributes per page, five runs
- Judge agreement between runs
What each measure means
- Pairwise voice score
- ScriptGrain's comparison of one text against another: the first text's counted features and judged attributes become a baseline profile, the second is compared with it attribute by attribute, and a weighted agreement (the raw score, 0 to 1) is mapped to a display score with bands: on voice (0.80 and above), drifting (0.55 to 0.80), off voice (below 0.55). This is the score behind the off-voice check and the compare-voice API.
- Counted features
- Measured in code, the same way every time: sentence length and its spread, contractions, commas, exclamation marks, semicolons, brackets, question marks, ellipses, the share of I, we and you, the share of the, a and an, and word lengths.
- Judged attributes
- Thirteen qualitative readings made by an AI model with ScriptGrain's production judge prompt, such as formality (0 to 10), humour, opening style, argument structure and rhythm. The model sees only the page's text.
- AUC
- The probability that a randomly chosen pair of a brand's own pages scores higher than a randomly chosen pair of rival pages (ties count half). 0.5 means the score cannot tell them apart; 1.0 means it always can.
- Pair kinds
- Same brand: two different pages from one brand's site. Rival, same sector: a page from one brand and a page from another brand in the same sector. Other sector: a page from one brand and a page from a brand in another sector. Each ordered pair is scored once, with the first page as the baseline.
- 95% interval
- From 2,000 bootstrap resamples of brands, drawn within each sector, so a brand with many pages cannot dominate the estimate.
Findings
- H1 (supported): a brand's own pages were barely more alike than a rival's (AUC 0.551 (95% interval 0.521 to 0.582)): We predicted the score would tell a brand's own page pairs from rival pairs in the same sector less than 60% of the time. It did so 55% of the time, in line with the pilot (0.54). The interval stays above 0.50, so there is a little brand signal, but it is small: the middle of the range for a brand's own pairs (median raw 0.654) sits close to rivals' (0.629). Against brands in other sectors the figure was 0.610.
- H2 (not supported): page type did not beat brand (median 0.657 (rivals, same page type) against 0.654 (same brand, different page type); difference 0.003 (95% interval -0.012 to 0.028)): We predicted, from the pilot, that two rivals' pages of the same type would score closer than two different pages from one brand. They scored the same. The pilot's version of the test also reversed: rival homepages scored a median 0.649 against each other and a homepage scored 0.657 against its own inner pages (pilot: 0.673 and 0.653). Page type does carry some signal, about as much as brand: among pairs from different brands, same-type pairs scored higher than different-type pairs 58% of the time.
- H3 (supported): sector carried a little signal (AUC 0.563 (95% interval 0.542 to 0.586)): A rival's page in the same sector scored higher than a page from another sector 56% of the time. 33% of other-sector pairs read off voice, against 26% of same-sector pairs and 24% of a brand's own.
- Robustness: every judge gave the same answer (H1 AUC 0.558, 0.564, 0.558 on three Sonnet 5.5 runs; 0.553 with no judge): As pre-registered, the result counts as robust to the judge only if H1 held on all three Sonnet runs and H2 and H3 matched on at least two. All three runs matched on all three hypotheses. The two models read the pages differently (they agreed exactly on 55% of attribute readings), yet the conclusions did not move. Haiku at temperature 0 was close to repeatable (98.4% of readings identical across two runs; median change in a pair's raw score 0, largest 0.114); Sonnet's three runs agreed on 89%.
- Exploratory: sectors differ, but no sector stands out clearly (AUC from 0.66 (energy) to 0.49 (banks)): Energy and travel brands were the easiest to tell from their rivals; banks, software and insurance the hardest. Each sector has six to ten brands, so these are point estimates without intervals and small differences between sectors should not be read as rankings. By page type the AUC ranged from 0.53 to 0.57; using only pages of 600 words or more gave 0.555.
What this study does not show
- That brands in a sector really do sound the same to readers. The study measures one score; it does not ask people, and a low AUC can mean the measure misses what makes a brand distinct.
- Which brands write well or badly. No page or brand is rated, and the per-brand numbers are not a league table.
- That a page, or a website, was written by AI, or that voice measurement, ScriptGrain's or anyone's, can tell whether a text was written by AI. Nothing in this study tests that.
- How the score behaves on other kinds of writing, such as articles, emails or a single author's work. Website copy is written by many hands and shares a lot with every site in its sector.
How a brand's pages scored
Five judge runs, one answer
Sector by sector
Pre-registered hypotheses and results
| Hypothesis | Result | Primary judge (Haiku 4.5) | Sonnet 5.5 runs A, B, C |
|---|---|---|---|
| H1: the score tells a brand's own page pairs from rival pairs in the same sector less than 60% of the time (AUC below 0.60) | Supported | 0.551 (95% interval 0.521 to 0.582) | 0.558, 0.564, 0.558 |
| H2: same-type pages from rival brands score higher than different-type pages from the same brand (difference in medians above zero) | Not supported | 0.003 (95% interval -0.012 to 0.028) | 0.002, 0.000, 0.000 |
| H3: a rival's page in the same sector scores higher than a page from another sector (AUC above 0.50) | Supported | 0.563 (95% interval 0.542 to 0.586) | 0.558, 0.559, 0.557 |
Scores by kind of pair
Raw scores (0 to 1) and band shares under the primary judge. Each ordered pair is counted once.
| Pair | Pairs | Raw 10th percentile | Median | 90th percentile | On voice | Drifting | Off voice |
|---|---|---|---|---|---|---|---|
| A brand's own pages | 782 | 0.496 | 0.654 | 0.770 | 26% | 50% | 24% |
| Rival brand, same sector | 8,336 | 0.490 | 0.629 | 0.750 | 18% | 56% | 26% |
| Brand in another sector | 79,984 | 0.472 | 0.605 | 0.727 | 12% | 55% | 33% |
Brand against page type
Median raw score for each combination. The same brand can only be paired with a different page type, because each brand has one page of each type.
| Same page type | Different page type | |
|---|---|---|
| Same brand | not possible | 0.654 |
| Rival brand, same sector | 0.657 | 0.620 |
| Brand in another sector | 0.626 | 0.600 |
By sector
| Sector | Brands analysed | Pages | AUC, own pages against rivals |
|---|---|---|---|
| energy | 6 | 18 | 0.66 |
| travel | 8 | 30 | 0.61 |
| fashion | 7 | 19 | 0.58 |
| telecoms | 10 | 31 | 0.58 |
| food and drink | 10 | 25 | 0.56 |
| charities | 9 | 32 | 0.55 |
| supermarkets | 10 | 33 | 0.55 |
| insurance | 9 | 33 | 0.53 |
| software | 10 | 40 | 0.52 |
| banks | 10 | 38 | 0.49 |
AUC by the type of the baseline page
| Baseline page | Pages | AUC |
|---|---|---|
| homepage | 76 | 0.57 |
| about | 66 | 0.53 |
| product | 83 | 0.56 |
| pricing | 41 | 0.54 |
| careers | 33 | 0.56 |
How much the judges agreed
Exact agreement: the share of attribute readings that were identical between two runs. Pair score change: how far a page pair's raw score moved between the runs.
| Runs compared | Exact agreement | Median pair score change | 95th percentile change |
|---|---|---|---|
| Haiku run 1 and Haiku run 2 | 98% | 0.000 | 0.038 |
| Sonnet run A and Sonnet run B | 89% | 0.011 | 0.051 |
| Haiku run 1 and Sonnet run A | 55% | 0.038 | 0.119 |
Coverage
How many brands in each sector had each page type missing under the frozen rules. Careers pages often live on separate job sites, and many homepages build their text in the browser, which the page reader does not run.
| Sector | Brands | Included | No homepage | No about | No product | No pricing | No careers |
|---|---|---|---|---|---|---|---|
| banks | 10 | 10 | 0 | 3 | 0 | 4 | 5 |
| energy | 6 | 6 | 1 | 1 | 1 | 5 | 4 |
| telecoms | 10 | 10 | 3 | 4 | 1 | 5 | 6 |
| software | 10 | 10 | 1 | 3 | 0 | 0 | 6 |
| supermarkets | 10 | 10 | 1 | 4 | 0 | 5 | 7 |
| travel | 10 | 8 | 2 | 3 | 2 | 4 | 8 |
| insurance | 10 | 9 | 2 | 2 | 1 | 3 | 8 |
| fashion | 10 | 7 | 4 | 6 | 1 | 9 | 8 |
| charities | 10 | 9 | 1 | 1 | 2 | 10 | 3 |
| food and drink | 10 | 10 | 1 | 3 | 3 | 10 | 8 |
The brand panel
| Sector | Brands |
|---|---|
| banks | metrobankonline.co.uk, tsb.co.uk, virginmoney.com, co-operativebank.co.uk, chase.co.uk, zopa.com, atombank.co.uk, kroo.com, skipton.co.uk, coventrybuildingsociety.co.uk |
| energy | scottishpower.co.uk, utilita.co.uk, goodenergy.co.uk, ecotricity.co.uk, outfoxthemarket.co.uk, fuseenergy.com |
| telecoms | ee.co.uk, talktalk.co.uk, hyperoptic.com, communityfibre.co.uk, smarty.co.uk, idmobile.co.uk, lebara.co.uk, lycamobile.co.uk, o2.co.uk, zen.co.uk |
| software | freeagent.com, gocardless.com, xero.com, sage.com, intercom.com, hubspot.com, canva.com, figma.com, clickup.com, miro.com |
| supermarkets | tesco.com, sainsburys.co.uk, asda.com, morrisons.com, aldi.co.uk, lidl.co.uk, ocado.com, coop.co.uk, iceland.co.uk, booths.co.uk |
| travel | britishairways.com, easyjet.com, ryanair.com, jet2.com, virginatlantic.com, tui.co.uk, loganair.co.uk, thetrainline.com, nationalexpress.com, eurostar.com |
| insurance | aviva.co.uk, directline.com, admiral.com, lv.com, axa.co.uk, hastingsdirect.com, esure.com, legalandgeneral.com, zurich.co.uk, bupa.co.uk |
| fashion | next.co.uk, marksandspencer.com, johnlewis.com, asos.com, boohoo.com, primark.com, superdry.com, riverisland.com, newlook.com, fatface.com |
| charities | cancerresearchuk.org, oxfam.org.uk, redcross.org.uk, macmillan.org.uk, rnli.org, nspcc.org.uk, barnardos.org.uk, mind.org.uk, savethechildren.org.uk, rspca.org.uk |
| food and drink | innocentdrinks.co.uk, pret.com, greggs.co.uk, nandos.co.uk, leon.co, wagamama.com, itsu.com, cookfood.net, gailsbread.co.uk, deliciouslyella.com |
Limitations
- 299 pages from 89 brands rather than the 480 the design allowed: careers and pricing pages were often missing, and some homepages build their text in the browser. The rules were applied as frozen and every gap is shown above.
- The page reader keeps navigation, footer and cookie text, which every page of one site shares. That raises same-brand scores, so the true separation of brand voice by this measure is, if anything, lower.
- Website pages are written by many people and agencies over years. 'Same brand' is a label, not a single writer.
- UK consumer brands in one week of October 2026, in English. Business-to-business, public-sector and personal sites are not covered.
- Both judges are Claude models; another vendor's model would read the pages differently. The counted-features-only result does not depend on any model.
- Per-sector figures rest on six to ten brands each and have no intervals.
Competing interests
ScriptGrain sells the score this study measures, and the analyst builds it. The study looks for a weakness in it. The hypotheses, brand panel and page rules were committed to the public repository before any page was fetched, every hypothesis is reported whether it passed or not, and the data and scripts are public.
Reproducing this
- Pre-registration, brand panel and every script: docs/research/sgr-009-preregistration.md and scripts/experiments/sgr-009/ in the ScriptGrain repository.
- Every page is named by URL and fetch time, so any page can be fetched again and measured; pages change, so expect small differences.
- Cost of the judges: 6.01 US dollars of API calls across five runs.
Data
- Every page: brand, sector, type, URL, counted features and the primary judge's readings (CSV) (csv)
- Every ordered page pair: kind and raw score under Haiku, Sonnet and counted features only (CSV, 89,102 rows) (csv)
- Every judge reading, all five runs (JSON) (json)
- Summary: hypotheses, intervals, bands, coverage, judge agreement (JSON) (json)
Released under CC BY 4.0: free to reuse, including commercially, with credit to ScriptGrain and a link to this page.
Citations
- Hanley, J. A. and McNeil, B. J. (1982). The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology
- Firecrawl (site map and page fetching)
How to cite
ScriptGrain (2026). Do brands in the same sector sound the same? 89 UK brands' websites, measured (Study SGR-009, conducted 4 October 2026). Dataset licensed CC BY 4.0. https://scriptgrain.com/research/do-brands-in-the-same-sector-sound-the-same
About the analyst
Jack Stovell has worked in finance and data for more than twelve years, building management reporting, forecasting and profitability models for advertising agencies and tech scale-ups, from SQL and Power BI reporting to board-level analysis. Since 2016 he has run Adapt Progress Evolve, an applied AI studio, where he builds and operates AI systems and data products: ScriptGrain's measurement of writing voice across 45 attributes, UK Spend, which brings 16.7 million rows of UK council spending into one queryable dataset, and more than thirty AI agents running in production. He designs ScriptGrain's studies and is accountable for every number in them.