Where did the AI-isms come from? 17,000 pieces of writing from before ChatGPT, measured
SGR-007, conducted 2026-10-04. Analyst: Jack Stovell, founder, ScriptGrain.
Question: Before ChatGPT, which kinds of human writing already used the words and habits now called AI-isms most, and when did each become common?
Summary
Background. Words such as "delve", "pivotal" and "seamless", and shapes such as "not X but Y", are now widely read as signs of AI writing. Models learned to write from human text, so these habits must have had human homes first. Popular explanations exist, most famously that "delve" came from the formal English of outsourced workers in Africa, especially Nigeria, who rated model answers (The Guardian, April 2024), but they have rarely been measured.
What we did. We fixed four hypotheses, the sources and the word lists before measuring anything, and published them. We then sampled 6,600 pages from Common Crawl's October 2021 snapshot of the web, the kind of text language models were trained on, across 16 kinds of writing, plus 103 company homepages from the same crawl, 10,000 arXiv abstracts written before ChatGPT, 600 passages from pre-1928 books in 8 genres, and traced every listed word through Google Books from 1800 to 2019. Every text was measured with the same deterministic code ScriptGrain uses, and only the numbers were kept.
What we found. Press releases were the densest home of AI-isms before ChatGPT: 1.61 uses per 1,000 words, against 0.44 in news and 0.24 in Wikipedia. Of the 44 listed words common enough to place, 14 were most at home there. The sentence shapes associated with AI writing had a different home: pre-1928 speeches, sermons and conduct books. In books, 40 of 43 listed words were already more common in 2019 than in 1980. The "delve came from Nigerian English" claim was not supported: Nigerian and Kenyan news used "delve" a little more than UK and US news, but the word was rare everywhere and the difference was within chance.
What it means. The "AI voice" looks less like an invention than an inheritance: the vocabulary of corporate announcements laid over the rhetorical shapes of older speeches and sermons, both well established in human writing long before language models. This study shows where the closest human matches were, not where any model learned anything.
Key numbers
- Before ChatGPT, press releases used AI-isms such as "innovative", "pivotal" and "seamless" 3.7 times as often as news articles: 1.61 against 0.44 uses per 1,000 words in pages from 2021 and earlier (ScriptGrain, 2026).
- Words that studies of scientific writing found language models over-use, including "innovative", "advancements", "pivotal", "delve", "meticulous", "underscore", "comprehensive", "potential", were most common in press releases in writing from before ChatGPT.
- The sentence shapes associated with AI writing, such as "not X but Y" and reflexive lists of three, were most frequent in pre-1928 speeches (2.07 per 1,000 words) and sermons and religion (1.94), against 0.45 in UK news.
- "Delve" was rare in 2021 news everywhere: it appeared on 3 of 600 Nigerian news pages, 1 of 600 Kenyan, and none of 1,000 UK and US pages.
- In English-language books, "innovative" went from 0.14 to 12.12 uses per million words between 1950 and 1980; "transformative" rose 31 times between 1980 and 2019, long before ChatGPT.
- Before ChatGPT, "delve" appeared in about 1 in 5,287 arXiv abstracts (1,057,471 abstracts from 2015 to November 2022 with no later revision); only 3.5% contained any of the twelve words later identified as 2024 AI tells.
- "Delve" was already 3.0 times as common in English-language books in 2019 as in 1980 (Google Books).
- 40 of 43 words now called AI-isms were more common in English-language books in 2019 than in 1980.
- Of 26 kinds of writing measured, a Q&A forum (English Stack Exchange) used AI-isms least: 0.04 per 1,000 words, 37 times less than press releases.
Sample
33 web sources in 16 kinds of writing, read from Common Crawl's October 2021 crawl (CC-MAIN-2021-43): up to 200 pages per source, in a seeded random order, kept if the product's article reader found at least 200 words. The 103 company homepages are the SGR-004 panel as Common Crawl saw them in October 2021. Scientific abstracts: 10,000 arXiv abstracts, a seeded random sample of the 1,057,471 from 2015 to November 2022 with no revision after 30 November 2022 (SGR-008's census). Books: 25 public-domain books in each of 8 genres from Project Gutenberg, three passages of about 800 words from each. Google Books Ngram (English 2019 corpus) for every listed word, 1800 to 2019.
- Written before ChatGPT launched (30 November 2022): the web pages were crawled in October 2021, and the books are older
- English pages with status 200, one capture per URL, no index, tag, category, author or search pages
- At least 200 words of article text for web pages, 100 for homepages
Method
Pre-registration: hypotheses H1 to H4, sources, sampling and word lists were committed to the public repository before any page was measured (docs/research/sgr-007-preregistration.md). Two clarifications made before measuring are listed there with reasons.
Scientific abstracts: a seeded random sample of 10,000 from the 1,057,471 arXiv abstracts dated 2015 to November 2022 with no revision after ChatGPT's launch, from SGR-008's census of arXiv, cleaned of LaTeX and measured like every other text; the share of all eligible abstracts containing each word is reported alongside.
Web pages: looked up in Common Crawl's own index files (read directly, not through its busy query server), fetched from its WARC archives and reduced to article text with the product's reader.
Measurement: each page or passage is scored with the same deterministic code ScriptGrain uses, for the word lists, the sentence-shape scan, em dashes, sentence length and contractions. No language model reads or scores any text.
Word lists: SGR-008's lists of 2024 tells and publicly mocked ChatGPT words (built from published studies of scientific writing) and SGR-004's AI-era marketing words, used unchanged.
Tests: differences in mean rate per page with 95% bootstrap intervals (10,000 resamples, seed 20261004). Exploratory analyses are labelled as such.
Variables measured
- AI-ism rate per 1,000 words
- Rate for each word list
- Delve rate
- AI sentence shapes per 1,000 words
- Em dashes per 1,000 words
- Average sentence length
- Contractions per 1,000 words
- Google Books frequency per million words, 1800 to 2019
What each measure means
- AI-ism rate (the primary measure)
- Uses per 1,000 words of three published word lists combined, each occurrence counted once: SGR-008's 2024 tells (innovative, utilized, advancements, pivotal, facilitates, firstly, tackle, showcasing, underscores, delve, delves, intricate), SGR-008's list of words publicly mocked as ChatGPT tells (delve, intricate, meticulous, tapestry, commendable, showcas, pivotal, realm), and SGR-004's AI-era marketing words.
- AI sentence shapes
- ScriptGrain's published scan for AI-associated sentence shapes and transitions per 1,000 words, the same scan as the free AI cliché checker: 'not X but Y' balancing, reflexive lists of three, chains of semicolons and stock connectives.
- Delve rate
- Uses of delve, delves, delved and delving per 1,000 words (for hypothesis H2).
- Home
- For each word, the kind of writing where its rate is highest relative to its rate across everything measured. Counted only for words with at least 20 uses.
- Take-off year
- In Google Books, the first year after which the word's 5-year average stays above twice its 1950 to 1980 average.
- 95% interval
- From 10,000 bootstrap resamples of pages. A pre-registered hypothesis counts as supported only if the interval for the difference excludes zero in the predicted direction.
Findings
- H1 (supported): press releases were the densest pre-AI home of AI-isms (1.61 per 1,000 words, against 0.44 in news and 0.24 in Wikipedia): Press releases led every kind of writing measured, followed by company homepages (1.36). As pre-registered, press releases, marketing and SEO blogs (0.69) and essay-mill essays (0.53) each used more than news, all countries pooled, and more than Wikipedia. Press releases against news: difference 1.17 (95% interval 0.99 to 1.35). Essay mills only just cleared news: 0.09 (95% interval 0.02 to 0.17).
- Exploratory: the classic ChatGPT words were most at home in press releases (14 of 44 words): For each word with at least 20 uses, we found the kind of writing where its rate was highest relative to everything measured. Of 44 words common enough to place, press releases were home to 14: "innovative", "advancements", "pivotal", "delve", "meticulous", "underscore", "comprehensive", "potential", "seamless", "empower", "streamline", "revolutionize", "reimagine", "redefine". Essay mills were home to "utilized", "firstly", the vocabulary of the student essay. The word lists were picked from what models over-use, so this is a description of where they lived, not a test.
- Exploratory: AI sentence shapes had a different, older home (2.07 per 1,000 words in pre-1928 speeches, 0.45 in UK news): The shapes the AI cliché scan looks for ("not X but Y" balancing, reflexive lists of three, stock connectives) were densest not on the modern web but in pre-1928 speeches, sermons and religion, conduct of life and self-help, science. 7 of 8 book genres used them more often than UK news, US news and Wikipedia; the exception was fiction. Vocabulary and shape point to different ancestors: the announcement and the pulpit.
- Exploratory: scientific abstracts had the quieter words, not the famous ones (0.36 AI-isms per 1,000 words (13 of 26); first of 26 for the quieter words): Pre-ChatGPT arXiv abstracts were mid-table on the main AI-ism rate but used the quieter LLM-favoured words more than any other writing measured (2.24 per 1,000 words of words such as "significant", "enhance", "additionally" and "potential"), and were home to "facilitates", "enhance", "additionally", "significant". Their AI sentence-shape rate (1.25) was the highest of any modern writing measured. The famous words were rare: across all 1,057,471 eligible abstracts, "delve" appeared in about 1 in 5,287. Studies of scientific writing since 2023 report large rises in exactly those words; see the citations.
- H2 (not supported): no clear 'delve' gap between Nigerian or Kenyan news and UK or US news (difference 0.005 per 1,000 words, 95% interval 0.000 to 0.013): "Delve" was used a little more in Nigerian and Kenyan news (pooled) than in UK and US news, in the predicted direction, but the interval reaches zero: it appeared on 3 of 600 Nigerian pages, 1 of 600 Kenyan pages and none of 1,000 UK and US pages. The wider 2024 tell list showed no difference either: 0.016 (95% interval -0.032 to 0.060). In public news writing from 2021, "delve" was too rare anywhere to carry the explanation on its own. This does not test the raters themselves.
- H4 (supported): the words were rising in books for decades before ChatGPT (40 of 43 words more common in 2019 than in 1980): In Google Books, "innovative" took off around 1975 (0.14 uses per million words in 1950, 12.12 in 1980), "transformative" rose 31 times between 1980 and 2019, "reimagine" 106 times, and "delve" took off around 2007. The exceptions, "utilized", "commendable", "streamline", had peaked earlier.
- H3 (not supported): fiction was not the least AI-like writing; casual Q&A was (Q&A forum 0.04, pre-1928 fiction 0.19 per 1,000 words): We predicted that pre-1928 fiction would use AI-isms least. It was low, but several book genres were lower, and the lowest of all was English Stack Exchange, where people answer each other's questions in plain words. In this sample, writing that announces or persuades used the most AI-isms, and plain answers to questions used the fewest.
What this study does not show
- Where any language model learned anything. The study finds the human writing that most resembles today's AI-isms; it does not trace a model's training.
- Anything about the people who rated model answers. H2 measures public news writing in Nigeria and Kenya, not the raters, so it neither confirms nor rules out the rater explanation.
- That a text containing these words or shapes was written by AI, or that voice measurement, ScriptGrain's or anyone's, can tell whether a text was written by AI. Nothing in this study tests that.
- What the words meant. Lists count spellings, so 'harness' in a pre-1928 commerce book is usually a horse's, not 'harness the power of'.
AI-isms by kind of writing
Two kinds of AI tell, two homes
Each point is one kind of writing. Left to right: AI-ism vocabulary. Bottom to top: AI sentence shapes. Pre-1928 books sit high and to the left; press releases and company homepages sit far to the right.
Three AI-isms in books, 1900 to 2019
Pre-registered hypotheses and results
| Hypothesis | Result | Test |
|---|---|---|
| H1: press releases, marketing blogs and essay mills each use more AI-isms than news and than Wikipedia | Supported | Press releases: vs news 1.17 (95% interval 0.99 to 1.35); Marketing and SEO blogs: vs news 0.25 (95% interval 0.16 to 0.34); Essay-mill essays: vs news 0.09 (95% interval 0.02 to 0.17) |
| H2: 'delve' and the 2024 tells are more common in Nigerian and Kenyan news than in UK and US news | Not supported | delve 0.005 (0.000 to 0.013); 2024 tells 0.016 (-0.032 to 0.060) |
| H3: pre-1928 fiction uses AI-isms least of every kind of writing | Not supported | lowest was Q&A forum |
| H4: most listed words are more common in books in 2019 than in 1980 | Supported | 40 of 43 words |
Every kind of writing measured
| Kind of writing | Pages or passages | AI-isms per 1,000 | 95% interval | AI sentence shapes per 1,000 | Em dashes per 1,000 | Pages with 'delve' |
|---|---|---|---|---|---|---|
| Press releases | 600 | 1.61 | 1.44 to 1.78 | 0.61 | 0.39 | 0.3% |
| Company homepages | 103 | 1.36 | 0.96 to 1.82 | 0.11 | 1.07 | 0.0% |
| News, Philippines | 200 | 0.87 | 0.70 to 1.04 | 0.19 | 3.73 | 0.0% |
| Medium posts | 200 | 0.83 | 0.64 to 1.05 | 0.49 | 4.30 | 0.5% |
| Self-help and productivity | 600 | 0.75 | 0.66 to 0.85 | 0.94 | 2.49 | 0.3% |
| Government news | 200 | 0.72 | 0.52 to 0.92 | 0.31 | 0.00 | 0.0% |
| Marketing and SEO blogs | 600 | 0.69 | 0.61 to 0.78 | 0.26 | 1.00 | 0.8% |
| News, Nigeria | 600 | 0.61 | 0.49 to 0.74 | 0.47 | 0.54 | 0.5% |
| Essay-mill essays | 600 | 0.53 | 0.47 to 0.60 | 0.92 | 0.04 | 1.2% |
| News, India | 400 | 0.43 | 0.36 to 0.52 | 0.21 | 0.73 | 0.3% |
| Careers advice | 200 | 0.41 | 0.28 to 0.55 | 0.48 | 0.93 | 0.0% |
| How-to content | 400 | 0.39 | 0.31 to 0.48 | 0.31 | 0.95 | 0.5% |
| Scientific abstracts (arXiv, before 2023) | 10,000 | 0.36 | 0.32 to 0.39 | 1.25 | 0.09 | 0.0% |
| News, Kenya | 600 | 0.35 | 0.29 to 0.42 | 0.28 | 0.30 | 0.2% |
| News, US | 600 | 0.33 | 0.24 to 0.42 | 0.34 | 2.45 | 0.0% |
| Books before 1928: commerce | 75 | 0.31 | 0.12 to 0.50 | 1.19 | 1.23 | 0.0% |
| Books before 1928: history | 75 | 0.29 | 0.16 to 0.43 | 1.17 | 0.65 | 0.0% |
| Books before 1928: conduct of life and self-help | 75 | 0.27 | 0.14 to 0.42 | 1.80 | 0.98 | 0.0% |
| News, UK | 400 | 0.26 | 0.19 to 0.34 | 0.45 | 0.12 | 0.0% |
| Wikipedia | 200 | 0.24 | 0.15 to 0.38 | 0.71 | 0.40 | 0.0% |
| Books before 1928: essays | 75 | 0.24 | 0.12 to 0.38 | 1.37 | 0.94 | 0.0% |
| Books before 1928: fiction | 75 | 0.19 | 0.08 to 0.30 | 0.59 | 0.65 | 0.0% |
| Books before 1928: sermons and religion | 75 | 0.14 | 0.07 to 0.23 | 1.94 | 3.53 | 0.0% |
| Books before 1928: speeches | 75 | 0.12 | 0.05 to 0.21 | 2.07 | 0.89 | 0.0% |
| Books before 1928: science | 75 | 0.09 | 0.02 to 0.20 | 1.73 | 2.19 | 0.0% |
| Q&A forum | 200 | 0.04 | 0.02 to 0.07 | 0.52 | 0.42 | 0.0% |
Where each word was most at home
Words with at least 20 uses across everything measured. 'Times its overall rate' compares the word's rate in its home with its rate across all the writing measured.
| Word | Uses | Home | Times its overall rate | Runner-up |
|---|---|---|---|---|
| game changer | 30 | News, Philippines | 35.8 | News, UK |
| next-level | 173 | News, Philippines | 32.4 | Company homepages |
| game-changer | 42 | News, Philippines | 25.6 | News, UK |
| supercharge | 20 | Company homepages | 18.1 | Self-help and productivity |
| at scale | 35 | Company homepages | 13.8 | Press releases |
| unleash | 91 | News, India | 13.2 | Medium posts |
| cutting-edge | 81 | Government news | 12.2 | Press releases |
| landscape | 355 | Company homepages | 11.2 | Press releases |
| transformative | 54 | News, Philippines | 11.1 | Press releases |
| harness | 119 | Books before 1928: commerce | 10.8 | Press releases |
| meticulous | 49 | Press releases | 9.5 | Careers advice |
| tackle | 340 | Government news | 7.8 | News, Nigeria |
| reimagine | 28 | Press releases | 7.6 | Company homepages |
| embark | 132 | News, Nigeria | 7.4 | Government news |
| innovative | 510 | Press releases | 7.2 | Essay-mill essays |
| delve | 38 | Press releases | 7.1 | Essay-mill essays |
| realm | 114 | Books before 1928: commerce | 7.0 | Books before 1928: conduct of life and self-help |
| revolutionize | 56 | Press releases | 6.6 | Medium posts |
| insights | 643 | Company homepages | 5.8 | Press releases |
| seamless | 178 | Press releases | 5.3 | Company homepages |
| unlock | 396 | Marketing and SEO blogs | 5.3 | How-to content |
| advancements | 67 | Press releases | 5.2 | Essay-mill essays |
| elevate | 164 | Books before 1928: essays | 5.1 | Books before 1928: conduct of life and self-help |
| navigate the | 20 | Careers advice | 5.1 | Government news |
| showcase | 243 | Company homepages | 5.0 | Marketing and SEO blogs |
| intricate | 39 | Books before 1928: fiction | 4.1 | Books before 1928: speeches |
| effortless | 47 | How-to content | 4.0 | Self-help and productivity |
| underscore | 48 | Press releases | 4.0 | News, US |
| journey | 690 | Self-help and productivity | 4.0 | Books before 1928: history |
| intricate | 41 | Books before 1928: fiction | 3.9 | Books before 1928: speeches |
| notably | 351 | Wikipedia | 3.9 | Scientific abstracts (arXiv, before 2023) |
| comprehensive | 740 | Press releases | 3.9 | How-to content |
| streamline | 125 | Press releases | 3.9 | Marketing and SEO blogs |
| pivotal | 82 | Press releases | 3.8 | Government news |
| utilized | 194 | Essay-mill essays | 3.5 | Books before 1928: commerce |
| redefine | 72 | Press releases | 3.5 | Essay-mill essays |
| showcasing | 35 | Company homepages | 3.4 | Press releases |
| empower | 490 | Press releases | 3.2 | Company homepages |
| firstly | 105 | Essay-mill essays | 3.1 | Scientific abstracts (arXiv, before 2023) |
| facilitates | 94 | Scientific abstracts (arXiv, before 2023) | 2.9 | Essay-mill essays |
| significant | 2,917 | Scientific abstracts (arXiv, before 2023) | 2.7 | Press releases |
| potential | 2,738 | Press releases | 2.4 | Scientific abstracts (arXiv, before 2023) |
| enhance | 1,681 | Scientific abstracts (arXiv, before 2023) | 2.4 | News, Philippines |
| additionally | 577 | Scientific abstracts (arXiv, before 2023) | 2.2 | Press releases |
| crucial | 697 | Government news | 2.1 | Scientific abstracts (arXiv, before 2023) |
Each word in Google Books
| Word | 1950 | 1980 | 2019 | 2019 vs 1980 | Take-off year |
|---|---|---|---|---|---|
| innovative | 0.14 | 12.12 | 15.21 | 1.3x | 1975 |
| next-level | 0.00 | 0.00 | 0.04 | 17.4x | 1977 |
| showcasing | 0.00 | 0.03 | 0.72 | 21.9x | 1978 |
| transformative | 0.02 | 0.17 | 5.28 | 31.0x | 1979 |
| game changer | 0.00 | 0.00 | 0.31 | new | 1979 |
| reimagine | 0.00 | 0.02 | 1.57 | 106.0x | 1980 |
| underscores | 0.14 | 1.03 | 1.89 | 1.8x | 1982 |
| showcase | 0.45 | 1.15 | 4.07 | 3.5x | 1986 |
| redefine | 0.85 | 2.76 | 4.09 | 1.5x | 1988 |
| cutting-edge | 0.06 | 0.02 | 0.84 | 40.5x | 1990 |
| unleash | 0.71 | 1.38 | 5.20 | 3.8x | 1994 |
| landscape | 10.52 | 17.67 | 29.99 | 1.7x | 1996 |
| game-changer | 0.00 | 0.00 | 0.10 | 341.0x | 1997 |
| navigate the | 0.14 | 0.11 | 1.39 | 13.2x | 1997 |
| pivotal | 1.47 | 1.71 | 3.95 | 2.3x | 1998 |
| empower | 10.09 | 6.26 | 19.42 | 3.1x | 1999 |
| unlock | 1.60 | 1.70 | 8.91 | 5.3x | 2001 |
| effortless | 0.73 | 0.81 | 2.72 | 3.4x | 2002 |
| delve | 0.45 | 0.50 | 1.51 | 3.0x | 2007 |
| delve | 0.80 | 0.84 | 2.49 | 3.0x | 2008 |
| journey | 19.27 | 14.61 | 50.14 | 3.4x | 2010 |
| at scale | 0.03 | 0.06 | 0.40 | 7.0x | 2014 |
| delves | 0.13 | 0.17 | 0.37 | 2.2x | 2017 |
| realm | 9.42 | 8.44 | 28.23 | 3.3x | 2019 |
| utilized | 17.74 | 25.37 | 11.36 | 0.4x | none |
| advancements | 1.27 | 0.78 | 1.86 | 2.4x | none |
| facilitates | 2.26 | 3.29 | 5.16 | 1.6x | none |
| firstly | 3.00 | 5.17 | 6.61 | 1.3x | none |
| tackle | 4.72 | 3.65 | 6.47 | 1.8x | none |
| intricate | 4.70 | 3.42 | 4.43 | 1.3x | none |
| intricate | 5.06 | 3.94 | 5.27 | 1.3x | none |
| meticulous | 2.11 | 2.59 | 3.95 | 1.5x | none |
| tapestry | 1.35 | 1.28 | 2.12 | 1.7x | none |
| commendable | 1.94 | 1.43 | 0.88 | 0.6x | none |
| seamless | 1.97 | 1.14 | 3.05 | 2.7x | none |
| elevate | 14.42 | 18.45 | 19.49 | 1.1x | none |
| streamline | 5.38 | 4.37 | 2.86 | 0.7x | none |
| supercharge | 1.12 | 0.24 | 0.29 | 1.2x | none |
| revolutionise | 0.24 | 0.21 | 0.33 | 1.5x | none |
| revolutionize | 1.21 | 1.18 | 1.65 | 1.4x | none |
| harness | 5.33 | 4.95 | 7.57 | 1.5x | none |
| embark | 6.80 | 6.19 | 6.98 | 1.1x | none |
| in seconds | 0.71 | 0.85 | 1.33 | 1.6x | none |
Coverage
What happened to every page or passage read, by kind of writing.
| Kind of writing | Measured | Too short | No capture |
|---|---|---|---|
| Government news | 200 | 15 | 0 |
| News, US | 600 | 81 | 0 |
| News, UK | 400 | 125 | 0 |
| News, Nigeria | 600 | 78 | 0 |
| News, Kenya | 600 | 124 | 0 |
| News, India | 400 | 11 | 0 |
| News, Philippines | 200 | 2 | 0 |
| Press releases | 600 | 22 | 0 |
| Essay-mill essays | 600 | 2 | 0 |
| Marketing and SEO blogs | 600 | 25 | 0 |
| How-to content | 400 | 9 | 0 |
| Self-help and productivity | 600 | 48 | 0 |
| Careers advice | 200 | 10 | 0 |
| Medium posts | 200 | 158 | 0 |
| Wikipedia | 200 | 5 | 0 |
| Q&A forum | 200 | 0 | 0 |
| Books before 1928: fiction | 75 | 0 | 0 |
| Books before 1928: essays | 75 | 1 | 0 |
| Books before 1928: sermons and religion | 75 | 2 | 0 |
| Books before 1928: conduct of life and self-help | 75 | 3 | 0 |
| Books before 1928: speeches | 75 | 1 | 0 |
| Books before 1928: commerce | 75 | 2 | 0 |
| Books before 1928: history | 75 | 1 | 0 |
| Books before 1928: science | 75 | 1 | 0 |
| Company homepages | 103 | 26 | 32 |
| Scientific abstracts (arXiv, before 2023) | 10,000 | 0 | 0 |
Limitations
- One snapshot of the web (October 2021) and a few sources per kind of writing. A different month or different outlets would move individual numbers; the large gaps, such as press releases against news, are the finding, not the decimals.
- Word lists were taken from studies of what models over-use, so finding them in press releases describes where those words lived; it does not show that press releases were over-represented in any model's training data.
- Pre-1928 books are a different era as well as a different genre, so their high rate of AI sentence shapes may reflect period style as much as genre.
- The sentence-shape scan is ScriptGrain's own instrument; it over-fires on some ordinary sentences, and another checker would count differently.
- Some 'homes' rest on a few dozen uses. The table shows the counts.
Competing interests
ScriptGrain sells writing-voice measurement and generation, and the analyst builds it. The hypotheses, sources and word lists were fixed and published before anything was measured, nothing was re-run or selected, and the scripts are public, so the numbers can be checked without taking our word for them.
Reproducing this
- Pre-registration, sources, word lists and every script: docs/research/sgr-007-preregistration.md and scripts/experiments/sgr-007/ in the ScriptGrain repository.
- Every web page in the data is named by URL and Common Crawl capture time, so any page can be fetched again from the CC-MAIN-2021-43 archive and measured.
Data
- Every page and passage measured (CSV) (csv)
- Summary: groups, hypotheses, word homes, Google Books (JSON) (json)
- arXiv: share of all eligible abstracts containing each 2024 tell word (JSON) (json)
Released under CC BY 4.0: free to reuse, including commercially, with credit to ScriptGrain and a link to this page.
Citations
- Hern, A. (2024). TechScape: How cheap, outsourced labour in Africa is shaping AI English. The Guardian
- Kobak, D. et al. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances
- Kousha, K. et al. (2025). How much are LLMs changing the language of academic papers after ChatGPT? arXiv:2509.09596
- Common Crawl, CC-MAIN-2021-43
- arXiv (abstracts via the OAI-PMH interface)
- Google Books Ngram Viewer
How to cite
ScriptGrain (2026). Where did the AI-isms come from? 17,000 pieces of writing from before ChatGPT, measured (Study SGR-007, conducted 4 October 2026). Dataset licensed CC BY 4.0. https://scriptgrain.com/research/where-did-ai-isms-come-from
About the analyst
Jack Stovell has worked in finance and data for more than twelve years, building management reporting, forecasting and profitability models for advertising agencies and tech scale-ups, from SQL and Power BI reporting to board-level analysis. Since 2016 he has run Adapt Progress Evolve, an applied AI studio, where he builds and operates AI systems and data products: ScriptGrain's measurement of writing voice across 45 attributes, UK Spend, which brings 16.7 million rows of UK council spending into one queryable dataset, and more than thirty AI agents running in production. He designs ScriptGrain's studies and is accountable for every number in them.