Where did the AI-isms come from? 17,000 pieces of writing from before ChatGPT, measured

SGR-007, conducted 2026-10-04. Analyst: Jack Stovell, founder, ScriptGrain.

Question: Before ChatGPT, which kinds of human writing already used the words and habits now called AI-isms most, and when did each become common?

Summary

Background. Words such as "delve", "pivotal" and "seamless", and shapes such as "not X but Y", are now widely read as signs of AI writing. Models learned to write from human text, so these habits must have had human homes first. Popular explanations exist, most famously that "delve" came from the formal English of outsourced workers in Africa, especially Nigeria, who rated model answers (The Guardian, April 2024), but they have rarely been measured.

What we did. We fixed four hypotheses, the sources and the word lists before measuring anything, and published them. We then sampled 6,600 pages from Common Crawl's October 2021 snapshot of the web, the kind of text language models were trained on, across 16 kinds of writing, plus 103 company homepages from the same crawl, 10,000 arXiv abstracts written before ChatGPT, 600 passages from pre-1928 books in 8 genres, and traced every listed word through Google Books from 1800 to 2019. Every text was measured with the same deterministic code ScriptGrain uses, and only the numbers were kept.

What we found. Press releases were the densest home of AI-isms before ChatGPT: 1.61 uses per 1,000 words, against 0.44 in news and 0.24 in Wikipedia. Of the 44 listed words common enough to place, 14 were most at home there. The sentence shapes associated with AI writing had a different home: pre-1928 speeches, sermons and conduct books. In books, 40 of 43 listed words were already more common in 2019 than in 1980. The "delve came from Nigerian English" claim was not supported: Nigerian and Kenyan news used "delve" a little more than UK and US news, but the word was rare everywhere and the difference was within chance.

What it means. The "AI voice" looks less like an invention than an inheritance: the vocabulary of corporate announcements laid over the rhetorical shapes of older speeches and sermons, both well established in human writing long before language models. This study shows where the closest human matches were, not where any model learned anything.

Key numbers

Sample

33 web sources in 16 kinds of writing, read from Common Crawl's October 2021 crawl (CC-MAIN-2021-43): up to 200 pages per source, in a seeded random order, kept if the product's article reader found at least 200 words. The 103 company homepages are the SGR-004 panel as Common Crawl saw them in October 2021. Scientific abstracts: 10,000 arXiv abstracts, a seeded random sample of the 1,057,471 from 2015 to November 2022 with no revision after 30 November 2022 (SGR-008's census). Books: 25 public-domain books in each of 8 genres from Project Gutenberg, three passages of about 800 words from each. Google Books Ngram (English 2019 corpus) for every listed word, 1800 to 2019.

Method

Pre-registration: hypotheses H1 to H4, sources, sampling and word lists were committed to the public repository before any page was measured (docs/research/sgr-007-preregistration.md). Two clarifications made before measuring are listed there with reasons.

Scientific abstracts: a seeded random sample of 10,000 from the 1,057,471 arXiv abstracts dated 2015 to November 2022 with no revision after ChatGPT's launch, from SGR-008's census of arXiv, cleaned of LaTeX and measured like every other text; the share of all eligible abstracts containing each word is reported alongside.

Web pages: looked up in Common Crawl's own index files (read directly, not through its busy query server), fetched from its WARC archives and reduced to article text with the product's reader.

Measurement: each page or passage is scored with the same deterministic code ScriptGrain uses, for the word lists, the sentence-shape scan, em dashes, sentence length and contractions. No language model reads or scores any text.

Word lists: SGR-008's lists of 2024 tells and publicly mocked ChatGPT words (built from published studies of scientific writing) and SGR-004's AI-era marketing words, used unchanged.

Tests: differences in mean rate per page with 95% bootstrap intervals (10,000 resamples, seed 20261004). Exploratory analyses are labelled as such.

Variables measured

What each measure means

AI-ism rate (the primary measure)
Uses per 1,000 words of three published word lists combined, each occurrence counted once: SGR-008's 2024 tells (innovative, utilized, advancements, pivotal, facilitates, firstly, tackle, showcasing, underscores, delve, delves, intricate), SGR-008's list of words publicly mocked as ChatGPT tells (delve, intricate, meticulous, tapestry, commendable, showcas, pivotal, realm), and SGR-004's AI-era marketing words.
AI sentence shapes
ScriptGrain's published scan for AI-associated sentence shapes and transitions per 1,000 words, the same scan as the free AI cliché checker: 'not X but Y' balancing, reflexive lists of three, chains of semicolons and stock connectives.
Delve rate
Uses of delve, delves, delved and delving per 1,000 words (for hypothesis H2).
Home
For each word, the kind of writing where its rate is highest relative to its rate across everything measured. Counted only for words with at least 20 uses.
Take-off year
In Google Books, the first year after which the word's 5-year average stays above twice its 1950 to 1980 average.
95% interval
From 10,000 bootstrap resamples of pages. A pre-registered hypothesis counts as supported only if the interval for the difference excludes zero in the predicted direction.

Findings

What this study does not show

AI-isms by kind of writing

Bar chart of AI-ism rate by kind of writing: press releases highest at 1.61 per 1,000 words, Q&A forum lowest at 0.04.
Source: ScriptGrain SGR-007, CC BY 4.0.

Two kinds of AI tell, two homes

Each point is one kind of writing. Left to right: AI-ism vocabulary. Bottom to top: AI sentence shapes. Pre-1928 books sit high and to the left; press releases and company homepages sit far to the right.

Scatter chart: pre-1928 books have the most AI sentence shapes and few AI-ism words; press releases and company homepages have the most AI-ism words and few AI sentence shapes.
Source: ScriptGrain SGR-007, CC BY 4.0.

Three AI-isms in books, 1900 to 2019

Line chart from Google Books: "innovative" rises sharply from the 1960s, "transformative" climbs from the 1990s and "delve" rises slowly after 2000.
Source: Google Books Ngram (English 2019), compiled by ScriptGrain SGR-007.

Pre-registered hypotheses and results

HypothesisResultTest
H1: press releases, marketing blogs and essay mills each use more AI-isms than news and than WikipediaSupportedPress releases: vs news 1.17 (95% interval 0.99 to 1.35); Marketing and SEO blogs: vs news 0.25 (95% interval 0.16 to 0.34); Essay-mill essays: vs news 0.09 (95% interval 0.02 to 0.17)
H2: 'delve' and the 2024 tells are more common in Nigerian and Kenyan news than in UK and US newsNot supporteddelve 0.005 (0.000 to 0.013); 2024 tells 0.016 (-0.032 to 0.060)
H3: pre-1928 fiction uses AI-isms least of every kind of writingNot supportedlowest was Q&A forum
H4: most listed words are more common in books in 2019 than in 1980Supported40 of 43 words

Every kind of writing measured

Kind of writingPages or passagesAI-isms per 1,00095% intervalAI sentence shapes per 1,000Em dashes per 1,000Pages with 'delve'
Press releases6001.611.44 to 1.780.610.390.3%
Company homepages1031.360.96 to 1.820.111.070.0%
News, Philippines2000.870.70 to 1.040.193.730.0%
Medium posts2000.830.64 to 1.050.494.300.5%
Self-help and productivity6000.750.66 to 0.850.942.490.3%
Government news2000.720.52 to 0.920.310.000.0%
Marketing and SEO blogs6000.690.61 to 0.780.261.000.8%
News, Nigeria6000.610.49 to 0.740.470.540.5%
Essay-mill essays6000.530.47 to 0.600.920.041.2%
News, India4000.430.36 to 0.520.210.730.3%
Careers advice2000.410.28 to 0.550.480.930.0%
How-to content4000.390.31 to 0.480.310.950.5%
Scientific abstracts (arXiv, before 2023)10,0000.360.32 to 0.391.250.090.0%
News, Kenya6000.350.29 to 0.420.280.300.2%
News, US6000.330.24 to 0.420.342.450.0%
Books before 1928: commerce750.310.12 to 0.501.191.230.0%
Books before 1928: history750.290.16 to 0.431.170.650.0%
Books before 1928: conduct of life and self-help750.270.14 to 0.421.800.980.0%
News, UK4000.260.19 to 0.340.450.120.0%
Wikipedia2000.240.15 to 0.380.710.400.0%
Books before 1928: essays750.240.12 to 0.381.370.940.0%
Books before 1928: fiction750.190.08 to 0.300.590.650.0%
Books before 1928: sermons and religion750.140.07 to 0.231.943.530.0%
Books before 1928: speeches750.120.05 to 0.212.070.890.0%
Books before 1928: science750.090.02 to 0.201.732.190.0%
Q&A forum2000.040.02 to 0.070.520.420.0%

Where each word was most at home

Words with at least 20 uses across everything measured. 'Times its overall rate' compares the word's rate in its home with its rate across all the writing measured.

WordUsesHomeTimes its overall rateRunner-up
game changer30News, Philippines35.8News, UK
next-level173News, Philippines32.4Company homepages
game-changer42News, Philippines25.6News, UK
supercharge20Company homepages18.1Self-help and productivity
at scale35Company homepages13.8Press releases
unleash91News, India13.2Medium posts
cutting-edge81Government news12.2Press releases
landscape355Company homepages11.2Press releases
transformative54News, Philippines11.1Press releases
harness119Books before 1928: commerce10.8Press releases
meticulous49Press releases9.5Careers advice
tackle340Government news7.8News, Nigeria
reimagine28Press releases7.6Company homepages
embark132News, Nigeria7.4Government news
innovative510Press releases7.2Essay-mill essays
delve38Press releases7.1Essay-mill essays
realm114Books before 1928: commerce7.0Books before 1928: conduct of life and self-help
revolutionize56Press releases6.6Medium posts
insights643Company homepages5.8Press releases
seamless178Press releases5.3Company homepages
unlock396Marketing and SEO blogs5.3How-to content
advancements67Press releases5.2Essay-mill essays
elevate164Books before 1928: essays5.1Books before 1928: conduct of life and self-help
navigate the20Careers advice5.1Government news
showcase243Company homepages5.0Marketing and SEO blogs
intricate39Books before 1928: fiction4.1Books before 1928: speeches
effortless47How-to content4.0Self-help and productivity
underscore48Press releases4.0News, US
journey690Self-help and productivity4.0Books before 1928: history
intricate41Books before 1928: fiction3.9Books before 1928: speeches
notably351Wikipedia3.9Scientific abstracts (arXiv, before 2023)
comprehensive740Press releases3.9How-to content
streamline125Press releases3.9Marketing and SEO blogs
pivotal82Press releases3.8Government news
utilized194Essay-mill essays3.5Books before 1928: commerce
redefine72Press releases3.5Essay-mill essays
showcasing35Company homepages3.4Press releases
empower490Press releases3.2Company homepages
firstly105Essay-mill essays3.1Scientific abstracts (arXiv, before 2023)
facilitates94Scientific abstracts (arXiv, before 2023)2.9Essay-mill essays
significant2,917Scientific abstracts (arXiv, before 2023)2.7Press releases
potential2,738Press releases2.4Scientific abstracts (arXiv, before 2023)
enhance1,681Scientific abstracts (arXiv, before 2023)2.4News, Philippines
additionally577Scientific abstracts (arXiv, before 2023)2.2Press releases
crucial697Government news2.1Scientific abstracts (arXiv, before 2023)

Each word in Google Books

Word1950198020192019 vs 1980Take-off year
innovative0.1412.1215.211.3x1975
next-level0.000.000.0417.4x1977
showcasing0.000.030.7221.9x1978
transformative0.020.175.2831.0x1979
game changer0.000.000.31new1979
reimagine0.000.021.57106.0x1980
underscores0.141.031.891.8x1982
showcase0.451.154.073.5x1986
redefine0.852.764.091.5x1988
cutting-edge0.060.020.8440.5x1990
unleash0.711.385.203.8x1994
landscape10.5217.6729.991.7x1996
game-changer0.000.000.10341.0x1997
navigate the0.140.111.3913.2x1997
pivotal1.471.713.952.3x1998
empower10.096.2619.423.1x1999
unlock1.601.708.915.3x2001
effortless0.730.812.723.4x2002
delve0.450.501.513.0x2007
delve0.800.842.493.0x2008
journey19.2714.6150.143.4x2010
at scale0.030.060.407.0x2014
delves0.130.170.372.2x2017
realm9.428.4428.233.3x2019
utilized17.7425.3711.360.4xnone
advancements1.270.781.862.4xnone
facilitates2.263.295.161.6xnone
firstly3.005.176.611.3xnone
tackle4.723.656.471.8xnone
intricate4.703.424.431.3xnone
intricate5.063.945.271.3xnone
meticulous2.112.593.951.5xnone
tapestry1.351.282.121.7xnone
commendable1.941.430.880.6xnone
seamless1.971.143.052.7xnone
elevate14.4218.4519.491.1xnone
streamline5.384.372.860.7xnone
supercharge1.120.240.291.2xnone
revolutionise0.240.210.331.5xnone
revolutionize1.211.181.651.4xnone
harness5.334.957.571.5xnone
embark6.806.196.981.1xnone
in seconds0.710.851.331.6xnone

Coverage

What happened to every page or passage read, by kind of writing.

Kind of writingMeasuredToo shortNo capture
Government news200150
News, US600810
News, UK4001250
News, Nigeria600780
News, Kenya6001240
News, India400110
News, Philippines20020
Press releases600220
Essay-mill essays60020
Marketing and SEO blogs600250
How-to content40090
Self-help and productivity600480
Careers advice200100
Medium posts2001580
Wikipedia20050
Q&A forum20000
Books before 1928: fiction7500
Books before 1928: essays7510
Books before 1928: sermons and religion7520
Books before 1928: conduct of life and self-help7530
Books before 1928: speeches7510
Books before 1928: commerce7520
Books before 1928: history7510
Books before 1928: science7510
Company homepages1032632
Scientific abstracts (arXiv, before 2023)10,00000

Limitations

Competing interests

ScriptGrain sells writing-voice measurement and generation, and the analyst builds it. The hypotheses, sources and word lists were fixed and published before anything was measured, nothing was re-run or selected, and the scripts are public, so the numbers can be checked without taking our word for them.

Reproducing this

Data

Released under CC BY 4.0: free to reuse, including commercially, with credit to ScriptGrain and a link to this page.

Citations

How to cite

ScriptGrain (2026). Where did the AI-isms come from? 17,000 pieces of writing from before ChatGPT, measured (Study SGR-007, conducted 4 October 2026). Dataset licensed CC BY 4.0. https://scriptgrain.com/research/where-did-ai-isms-come-from

About the analyst

Jack Stovell has worked in finance and data for more than twelve years, building management reporting, forecasting and profitability models for advertising agencies and tech scale-ups, from SQL and Power BI reporting to board-level analysis. Since 2016 he has run Adapt Progress Evolve, an applied AI studio, where he builds and operates AI systems and data products: ScriptGrain's measurement of writing voice across 45 attributes, UK Spend, which brings 16.7 million rows of UK council spending into one queryable dataset, and more than thirty AI agents running in production. He designs ScriptGrain's studies and is accountable for every number in them.