The AI tells moved: 2.1 million arXiv abstracts, 2015 to 2026

SGR-008, conducted 2026-10-04. Analyst: Jack Stovell, founder, ScriptGrain.

Question: When words like "delve" became a joke, did language models stop shaping scientific writing, or did the tells change?

Summary

Background. Since 2024, studies have shown that words such as "delve", "intricate" and "underscore" spiked in scientific writing after ChatGPT, and that "delve" fell once it became a public joke. Estimates of how many papers are written with language-model help keep rising. One study found other model-favoured words still rising as "delve" fell (Geng and Trotta, 2025). We found no study among the 444 on this topic in our literature search that compared the full set of words rising in 2026 with 2024's, or tested whether the word lists behind the prevalence estimates still work.

What we did. We took every arXiv paper first submitted between January 2015 and September 2026 from arXiv's own metadata service, with the date of every revision, and kept 2,122,337 English abstracts of 50 words or more. A pilot on a 41,657-abstract sample suggested the result; we then wrote down the hypotheses, word lists and tests (published with this study) before measuring the full set, and tested on the 2,078,952 papers the pilot never saw. Pre-ChatGPT papers whose abstracts were revised after November 2022 were left out of the baseline.

What we found. The twelve best-known 2024 tells rose from 4.5% of abstracts (2020 to 2022) to 11.9% in 2024, then fell to 3.0% in 2026, below where they started. Twelve other words, found in the pilot, rose from 1.9% to 12.4% over the same years, at least 3.5 times their baseline in all eight arXiv subject groups, overtaking the old set in the second half of 2025. The 50 style words that rose most in 2024 and in 2026 barely overlap (6 of them in common). Sentence length did not change.

What it means. The tells moved; the use did not go away. The standard excess-word method puts a lower bound of 68% on the share of 2026 abstracts processed by a language model when it uses the words that rose in 2026, and 20% when it uses the words that rose in 2024. Word-frequency estimates built on 2023 to 2024 vocabulary now undercount.

Key numbers

Sample

Every paper on arXiv, harvested on 4 October 2026 from arXiv's OAI-PMH service in the arXivRaw format, which carries the date of every version. 2,122,337 abstracts met the rules below. arXiv serves the latest version's abstract, which is why revision dates matter.

Method

Pilot, 4 October 2026: Firecrawl's Research Index returned 41,657 arXiv abstracts for 38 categories x 3 fixed topics x 2015 to 2026. The pilot produced the 2026 word set and the idea; it is not part of any test.

Pre-registration: hypotheses H1 to H5, the frozen word sets, the excess-word method, thresholds and exclusions were committed to the ScriptGrain repository before the census was analysed. Seven clarifications and one deviation followed, each before the census was measured and each listed with its reason.

Census: every arXiv record from arXiv's OAI-PMH service (3,195,084 records), filtered by the rules above.

Frozen-set tests: share of abstracts with at least one word of each set, 2020 to 2022 against 2024 and 2026 (January to September), with 95% confidence intervals from 2,000 bootstrap resamples of subject categories.

Excess words: for 2023 to 2026, each word's share against a straight-line extension of its 2021 to 2022 trend. Words were chosen on papers submitted in odd months and scored on papers submitted in even months, so no word is scored on the papers that selected it. Words already in 5% or more of abstracts were left out of the word sets (the bound becomes unstable).

Labelling: the 351 top candidates were labelled content or style by two independent model passes (agreement 343/351, Cohen's kappa 0.892) and a third pass for the 8 disagreements.

Instruments: average sentence length, punctuation and the AI-tell scan were measured on every abstract with the code ScriptGrain runs (measureTextFeatures, scanCliches).

Placebo (added after the confirmatory run): the lower-bound method was run five years earlier, 2016 to 2021, where no language models existed, to measure how much excess ordinary drift produces.

Variables measured

What each measure means

Share of abstracts
The percentage of abstracts that contain the word (or at least one word of a set) at least once, matched as a whole lowercase word after LaTeX maths and commands are removed.
2024 tells (the old set)
Twelve words fixed before the census was measured: innovative, utilized, advancements, pivotal, facilitates, firstly, tackle, showcasing, underscores, delve, delves, intricate. All were named as ChatGPT-era over-used words in the published studies cited below.
2026 tells (the new set)
Twelve words fixed before the census was measured: underexplored, grounded, aligns, synthesizes, replaces, leaving, recovers, expose, structurally, reproducible, treats, vanishes. They were found in the pilot sample, so the census test uses only papers the pilot did not include.
Excess word
A word whose share of abstracts in a year is above what its 2021 to 2022 trend predicts (the method of Kobak et al., 2025). The gap is the observed share minus the predicted share.
Style word
An excess word a writer could use on any subject ("remains", "enhance"), as opposed to a content word that names a topic ("llm", "agentic"). Every candidate was labelled by two independent model passes and a third for disagreements; all labels are published.
Lower bound
If a share s of abstracts were written with a model, and such an abstract always contained a word from the set, the set's share would be (1 minus s) times the predicted share plus s. Solving for s gives the smallest share consistent with the data. The true share is higher whenever the model sometimes avoids every word in the set.
Baseline (2020 to 2022)
Abstracts first submitted from January 2020 to November 2022, keeping only papers with no revision dated after 30 November 2022, so no baseline abstract can have been rewritten in the ChatGPT era.

Findings

What this study does not show

Two sets of tells, half-year by half-year

Share of abstracts using at least one word of each twelve-word set. Both sets were fixed before the census was measured.

Line chart: the 2024 tells rise from about 2% of abstracts in 2015 to 12.1% in early 2024 and fall to 2.0% by late 2026; the 2026 tells stay near 2% until 2024 and rise to 13.3%.
Source: ScriptGrain SGR-008, CC BY 4.0.

The frozen sets, tested

Papers outside the pilot. Baseline papers revised after 30 November 2022 are left out. Abstracts: 446,418 (2020 to 2022), 233,900 (2024), 262,617 (2026, January to September). Intervals: 95%, bootstrap over subject categories.

Word set2020 to 2022202420262024 minus baseline2026 minus 2024
2024 tells4.54%11.85%3.03%+7.31 (+5.4 to +8.8)-8.82 (-11.0 to -6.3)
2026 tells1.90%3.31%12.39%+1.41 (+0.9 to +1.8)+9.08 (+6.8 to +10.9)
Control: however, results, using, show, method73.48%74.88%72.64%+1.39 (+0.8 to +1.8)-2.24 (-3.7 to -0.4)

By subject group

Share of abstracts using a 2026 tell, by the arXiv group of the paper's first category, and the lower bound from both sets of 50 style words. 'Too uncertain' marks a bound whose 95% interval is wider than 40 points, which happens in the smaller groups.

Group2026 tells, 2020 to 20222026 tells, 2026Times baselineLower bound 2024Lower bound 2026
Computer science2.46%18.70%7.654%80%
Electrical engineering and systems1.68%9.56%5.746%too uncertain
Statistics2.22%11.53%5.246%too uncertain
Physics and astronomy1.65%6.18%3.729%63%
Mathematics1.32%4.64%3.514%49%
Quantitative biology2.19%15.13%6.9too uncertaintoo uncertain
Quantitative finance1.70%11.45%6.8too uncertaintoo uncertain
Economics1.38%9.77%7.1too uncertaintoo uncertain

Three 2024 tells

Line chart: delve, showcasing and innovative peak in 2024 and fall by 2026; delve from 0.35% to 0.02%.
Source: ScriptGrain SGR-008, CC BY 4.0.

Three 2026 tells

Line chart: grounded, recovers and underexplored rise from 2024 to 2026; grounded to 2.44% of abstracts.
Source: ScriptGrain SGR-008, CC BY 4.0.

The style words that rose most, 2024

Top 15 of the 50, scored on even-month papers: the share predicted by the word's 2021 to 2022 trend, the share observed, and the ratio.

WordPredictedObservedRatio
challenges4.82%9.71%2.0x
additionally2.58%7.53%2.9x
findings3.33%7.42%2.2x
enhance2.23%6.61%3.0x
insights2.15%5.37%2.5x
crucial3.05%6.50%2.1x
comprehensive2.51%6.29%2.5x
enhancing0.66%3.93%6.0x
capabilities2.05%5.78%2.8x
introduces2.12%4.86%2.3x
effectively3.32%6.63%2.0x
particularly2.48%5.61%2.3x
diverse2.72%6.32%2.3x
leveraging1.91%4.67%2.5x
enabling1.85%4.52%2.4x

The style words that rose most, 2026

WordPredictedObservedRatio
remains3.40%12.70%3.7x
yet3.82%11.41%3.0x
rather2.12%8.78%4.2x
establish2.73%8.31%3.0x
explicit2.57%8.07%3.1x
enabling2.13%7.97%3.7x
every2.38%6.65%2.8x
remain1.78%6.86%3.9x
fixed2.82%7.01%2.5x
yields1.67%6.18%3.7x
improves3.41%8.03%2.4x
practical3.52%7.26%2.1x
structured1.00%5.42%5.4x
findings3.67%7.28%2.0x
consistently1.04%5.38%5.2x

Lower bounds by year

Line chart: with 2026 words the lower bound rises from 12% in 2023 to 68% in 2026; with 2024 words it peaks at 40% in 2025 and falls to 20%.
Source: ScriptGrain SGR-008, CC BY 4.0.
Year2024 words2026 wordsBoth
202313% (8 to 17%)12% (9 to 14%)17% (11 to 22%)
202429% (20 to 37%)23% (18 to 27%)33% (25 to 40%)
202540% (31 to 47%)48% (41 to 53%)51% (43 to 58%)
202620% (13 to 26%)68% (62 to 73%)64% (56 to 70%)

Placebo: the same method before language models

The method shifted five years earlier: 2016 to 2017 as the trend, 2018 to 2021 as the 'after' years. The true answer is zero, so anything above zero is drift the method mistakes for model use.

WordsPlacebo 'bound' at 2021 (like 2026)Interval
The 2026 words0.5%-9 to 10%
The 2024 words10.2%4 to 16%
50 words re-selected on 2016 to 2021 data10.1%-2 to 19%

Sensitivity: all papers, revised ones included

Word set2020 to 202220242026
2024 tells4.52%11.85%3.03%
2026 tells1.94%3.31%12.39%

Other measures, 2020 to 2022 against 2026

Measure2020 to 20222026Difference (95% interval)
Words per abstract161.80174.99+13.20 (+10.02 to +16.21)
Average sentence length (words)24.4624.49+0.03 (-0.14 to +0.20)
Commas per sentence1.051.38+0.34 (+0.28 to +0.37)
Semicolons per 1,000 words0.571.05+0.48 (+0.37 to +0.58)
ScriptGrain AI-tell density per 1,000 words1.311.43+0.12 (+0.09 to +0.16)
Em dashes per 1,000 words0.020.07+0.05 (+0.04 to +0.06)

What 43 earlier studies estimated

We searched Firecrawl's Research Index with 20 fixed queries, screened 1,337 results to the studies that estimate how much scholarly text is written or edited with language models, and checked each paper exists and each number against its own text. 43 qualify. Methods differ (20 per-document detector, 13 excess vocabulary / word frequency, 5 distributional (population-level) model, 2 keyword floor, 2 other, 1 classifier trained on polished text), so the numbers are not one series. Detector-based studies often report a high rate before ChatGPT existed, so their figures are not shares of model-written text. The full table, with each quotation, is a download below.

StudyCorpusEstimate
Mingmeng Geng (2024)One million arXiv papers (abstracts)Fraction of LLM-style abstracts about 35% in computer science
Gwinyai Masukume (2024)PubMed and Scopus records for "delve", "realm" and "underscore"Co-use of "delve", "realm" and "underscore" up to 85-fold higher in 2023-2024 than in 2022 and earlier
Dmitry Kobak (2024)More than 15 million PubMed biomedical abstractsAt least 13.5% of 2024 PubMed abstracts processed with LLMs (lower bound; up to 40% in some subcorpora)
Sergio E Uribe (2024)299,695 dental research abstracts indexed in PubMedAbstracts with ChatGPT "signaling words": 47.1 per 10,000 before release vs 224.2 per 10,000 after
Weixin Liang (2024)1,121,912 preprints and published papers on arXiv, bioRxiv and Nature portfolio journalsLLM-modified content up to 22% of computer science papers; up to 9% in mathematics and Nature portfolio
Kayvan Kousha (2025)Abstracts in six databases (Scopus, Web of Science, PubMed, PMC, Dimensions, OpenAlex); 2.4 million PMC open-access fullFrequency of "delve" up 1,500%, "underscore" up 1,000%, "intricate" up 700% between 2022 and 2024; not a share of papers
Andrew Gray (2025)Full text of published papers across all research fields indexed in DimensionsLLM tools likely involved in more than 10% of all published papers in 2024
A. Arezki (2025)64,444 unique abstracts from 38 urology journals (PubMed)Proportion of AI-like text 1.8% in 2020 and 5.3% in 2024
Morgan D. Sanger (2026)149,452 abstracts published by the American Society of Civil Engineers (journals and proceedings), civil and environment13.9% (2024) and 20.4% (2025) of abstracts likely LLM-written (conservative); frequency-shift estimate 15.3% and 26.2%
Mike Thelwall (2026)1.25 million MDPI articles (full text); 80 LLM-associated terms; 73 journals with 500+ articles in 2021 examined separatTerm family "underscore" increased up to 29-fold; word-frequency change, not a share of papers
Przemysław Czuma (2026)69,632 first-version medRxiv preprints with an extractable Discussion section (full-text XML)Em-dash in Discussion sections rose from 4.23% of preprints before ChatGPT to 11.58% after (20.3% in 2025)
Lena Holzwarth (2026)Full texts of open-access biomedical papers from PubMed Central89% of papers show excess LLM-associated vocabulary by end of 2025; 68% of Discussion paragraphs vs 32% of Methods paragraphs
Aron Lee (2026)398,296 Korean-language abstracts (KCI), with 47,165 Vietnamese abstracts for comparisonLower bound on LLM-processed Korean abstracts 3.5%, 10.5% and 16.1% for 2024-2026 (single-word); 7.8%, 20.6%, 33.0% (split-half set)
Serat M. Saad (2026)Full text of 207,111 astro-ph papers; 392 papers that disclose model use calibrate the assisted rateAbout 54% of 2025 astro-ph papers carry a language-model trace (at or above 36% under varied assumptions); 0.81% disclose it
M. Z. Naser (2026)221,425 funded grant abstracts from NSF (n = 96,020), NIH (n = 80,496) and UKRI (n = 44,909)Excess marker-word rate in post-ChatGPT grant abstracts of 4.0-8.4% (NSF), 1.4-3.7% (NIH), 2.8-6.8% (UKRI) above the pre-ChatGPT baseline
Kyle Siler (2026)Full texts of 7.3 million journal articles from four publishers (Elsevier, Frontiers, MDPI, PLoS); 228 focal wordsAn estimated 57% of published articles in 2025 (12% in 2023) showed evidence of LLM influence
Hadar Better (2026)2,770 abstracts from six major dental journalsHigh-suspicion abstracts (AI-marker score of 20 or more) rose from 18.0% to 25.8%; AI marker words per abstract rose from 0.69 to 0.97
Sota Nakamura (2026)Articles in Surgery Today and Surgical Case ReportsRare AI-related word density rose after 2023 (slope change beta = 1.05 per year); no share estimated

Limitations

Competing interests

ScriptGrain sells writing-voice measurement and generation, and the analyst builds it. The hypotheses, word lists and tests were written down and committed before the census was measured; every change after that is listed in the pre-registration with its reason. Word labelling and the literature extraction used Claude models; their outputs are published so they can be checked.

Reproducing this

Data

Released under CC BY 4.0: free to reuse, including commercially, with credit to ScriptGrain and a link to this page.

Citations

How to cite

ScriptGrain (2026). The AI tells moved: 2.1 million arXiv abstracts, 2015 to 2026 (Study SGR-008, conducted 4 October 2026). Dataset licensed CC BY 4.0. https://scriptgrain.com/research/the-ai-tells-moved

About the analyst

Jack Stovell has worked in finance and data for more than twelve years, building management reporting, forecasting and profitability models for advertising agencies and tech scale-ups, from SQL and Power BI reporting to board-level analysis. Since 2016 he has run Adapt Progress Evolve, an applied AI studio, where he builds and operates AI systems and data products: ScriptGrain's measurement of writing voice across 45 attributes, UK Spend, which brings 16.7 million rows of UK council spending into one queryable dataset, and more than thirty AI agents running in production. He designs ScriptGrain's studies and is accountable for every number in them.