# AI words in scientific abstracts, 2015 to 2026: delve fell, new tells rose

> Every arXiv abstract from January 2015 to September 2026 (2,122,337 papers), measured against word lists fixed in advance. The 2024 AI tells peaked and fell below their pre-ChatGPT level; a different set of words rose in every field, so lists built on 2024 models now miss most of the signal.

Canonical: https://scriptgrain.com/research/the-ai-tells-moved

# The AI tells moved: 2.1 million arXiv abstracts, 2015 to 2026

**SGR-008**, conducted 2026-10-04. Analyst: Jack Stovell, founder, ScriptGrain.

**Question:** When words like "delve" became a joke, did language models stop shaping scientific writing, or did the tells change?

## Summary

**Background.** Since 2024, studies have shown that words such as "delve", "intricate" and "underscore" spiked in scientific writing after ChatGPT, and that "delve" fell once it became a public joke. Estimates of how many papers are written with language-model help keep rising. One study found other model-favoured words still rising as "delve" fell (Geng and Trotta, 2025). We found no study among the 444 on this topic in our literature search that compared the full set of words rising in 2026 with 2024's, or tested whether the word lists behind the prevalence estimates still work.

**What we did.** We took every arXiv paper first submitted between January 2015 and September 2026 from arXiv's own metadata service, with the date of every revision, and kept 2,122,337 English abstracts of 50 words or more. A pilot on a 41,657-abstract sample suggested the result; we then wrote down the hypotheses, word lists and tests (published with this study) before measuring the full set, and tested on the 2,078,952 papers the pilot never saw. Pre-ChatGPT papers whose abstracts were revised after November 2022 were left out of the baseline.

**What we found.** The twelve best-known 2024 tells rose from 4.5% of abstracts (2020 to 2022) to 11.9% in 2024, then fell to 3.0% in 2026, below where they started. Twelve other words, found in the pilot, rose from 1.9% to 12.4% over the same years, at least 3.5 times their baseline in all eight arXiv subject groups, overtaking the old set in the second half of 2025. The 50 style words that rose most in 2024 and in 2026 barely overlap (6 of them in common). Sentence length did not change.

**What it means.** The tells moved; the use did not go away. The standard excess-word method puts a lower bound of 68% on the share of 2026 abstracts processed by a language model when it uses the words that rose in 2026, and 20% when it uses the words that rose in 2024. Word-frequency estimates built on 2023 to 2024 vocabulary now undercount.

## Key numbers

- Abstracts on arXiv using one of the best-known 2024 AI tells (delve, showcasing, underscores, pivotal and eight others) rose from 4.5% in 2020 to 2022 to 11.9% in 2024, then fell to 3.0% in 2026 (ScriptGrain, 2026; 2,122,337 abstracts).
- A different set of words (underexplored, grounded, aligns, recovers and eight others) rose from 1.9% of arXiv abstracts in 2020 to 2022 to 12.4% in 2026, and overtook the 2024 tells in the second half of 2025.
- "Delve" appeared in 0.35% of arXiv abstracts in 2024 and 0.02% in 2026.
- The excess-word method gives a lower bound of 68% for the share of 2026 arXiv abstracts processed by a language model when it uses 2026 vocabulary, and 20% when it uses 2024 vocabulary (ScriptGrain, 2026).

## Sample

Every paper on arXiv, harvested on 4 October 2026 from arXiv's OAI-PMH service in the arXivRaw format, which carries the date of every version. 2,122,337 abstracts met the rules below. arXiv serves the latest version's abstract, which is why revision dates matter.

- First submitted from January 2015 to September 2026 (from the arXiv identifier); October 2026 was incomplete
- At least 50 words after LaTeX was removed, in English
- Primary subject is the first category the authors listed

## Method

Pilot, 4 October 2026: Firecrawl's Research Index returned 41,657 arXiv abstracts for 38 categories x 3 fixed topics x 2015 to 2026. The pilot produced the 2026 word set and the idea; it is not part of any test.

Pre-registration: hypotheses H1 to H5, the frozen word sets, the excess-word method, thresholds and exclusions were committed to the ScriptGrain repository before the census was analysed. Seven clarifications and one deviation followed, each before the census was measured and each listed with its reason.

Census: every arXiv record from arXiv's OAI-PMH service (3,195,084 records), filtered by the rules above.

Frozen-set tests: share of abstracts with at least one word of each set, 2020 to 2022 against 2024 and 2026 (January to September), with 95% confidence intervals from 2,000 bootstrap resamples of subject categories.

Excess words: for 2023 to 2026, each word's share against a straight-line extension of its 2021 to 2022 trend. Words were chosen on papers submitted in odd months and scored on papers submitted in even months, so no word is scored on the papers that selected it. Words already in 5% or more of abstracts were left out of the word sets (the bound becomes unstable).

Labelling: the 351 top candidates were labelled content or style by two independent model passes (agreement 343/351, Cohen's kappa 0.892) and a third pass for the 8 disagreements.

Instruments: average sentence length, punctuation and the AI-tell scan were measured on every abstract with the code ScriptGrain runs (measureTextFeatures, scanCliches).

Placebo (added after the confirmatory run): the lower-bound method was run five years earlier, 2016 to 2021, where no language models existed, to measure how much excess ordinary drift produces.

### Variables measured

- Share of abstracts using each word and word set
- Excess word share against trend
- Lower bound on model-processed abstracts
- Average sentence length
- AI-tell density per 1,000 words
- Em dashes per 1,000 words
- Words per abstract

## What each measure means

## Findings

- **The 2024 tells rose, then fell below their pre-ChatGPT level** (4.5% to 11.9% to 3.0% of abstracts): The twelve best-known 2024 tells were in 4.54% of abstracts from 2020 to 2022, 11.85% in 2024 and 3.03% in 2026 (2024 minus baseline +5.4 to +8.8 points, 2026 minus 2024 -11.0 to -6.3 points, 95% intervals). 2026 is 1.5 points below the baseline (interval -2.1 to -0.9). By the second half of 2026 the set was at 2.0%, about its 2016 level. "Delve" went from 0.04% of abstracts in 2022 to 0.35% in 2024 and 0.02% in 2026.
- **A new set of words rose in every field** (1.9% to 12.4% of abstracts): Twelve words found in the pilot (underexplored, grounded, aligns, synthesizes, replaces, leaving, recovers, expose, structurally, reproducible, treats, vanishes) were in 1.90% of abstracts from 2020 to 2022, 3.31% in 2024 and 12.39% in 2026 (2026 minus 2024 +6.8 to +10.9 points). Tested only on papers outside the pilot, the rise held in all eight subject groups: computer science 7.6 times, electrical engineering and systems 5.7 times, statistics 5.2 times, physics and astronomy 3.7 times, mathematics 3.5 times, quantitative biology 6.9 times, quantitative finance 6.8 times, economics 7.1 times. The two sets crossed in the second half of 2025 and the new set has led in every half-year since.
- **The words that rose most in 2024 and in 2026 are different words** (6 of the top 50 style words in common (Jaccard 0.06)): Ranking every style word by how far it rose above its own trend, the top 50 for 2024 and the top 50 for 2026 share six words: critical, distinct, enabling, findings, practical, remains. The 2024 list leads with challenges, additionally, findings, enhance, insights, crucial, comprehensive, enhancing; the 2026 list with remains, yet, rather, establish, explicit, enabling, every, remain. Both lists are below in full.
- **Lists built on 2024 vocabulary now undercount** (at least 68% (2026 words) against 20% (2024 words) of 2026 abstracts): The same lower-bound method, on the same 2026 abstracts, gives at least 68% (interval 62 to 73%) with the 50 style words that rose most in 2026 and 20% (interval 13 to 26%) with the 50 that rose most in 2024. With 2024 words the bound peaks in 2025 (40%) and falls; with 2026 words it rises every year: 2023 12%, 2024 23%, 2025 48%, 2026 68%. A placebo run on 2016 to 2021, before any language model, gave 0% for the 2026 words and 10% for the 2024 words, which were already drifting upward (Study SGR-007 traces where those habits came from); re-running the whole selection on pre-LLM years produced up to 10% (interval up to 19%). Read the bounds with that margin in mind.
- **Sentence length did not move** (24.5 to 24.5 words per sentence): Average sentence length was 24.46 words in 2020 to 2022 and 24.49 in 2026 (difference +0.03; 90% interval -0.11 to +0.17, inside the pre-registered one-word margin). The change is in vocabulary, not rhythm.
- **No sign that old abstracts were re-polished** (4.4% against 4.5% (2024 tells, revised against unrevised)): 11.2% of 2020 to 2022 papers have a version dated after November 2022, so their current abstract may have been edited since. They carry the 2024 tells at 4.39% against 4.54% for papers never revised, and the 2026 words at 2.27% against 1.90%. Leaving them in or out changes no result (sensitivity table below).

## What this study does not show

- That voice measurement, ScriptGrain's or anyone's, can tell whether a text was written by AI. Nothing in this study tests that.
- Which papers were written with a language model. Word shares describe a body of writing, never one paper.
- Why the tells changed. Newer models, authors editing out words that became jokes, and people adopting model vocabulary second-hand all fit the data.
- Anything about journals, theses or fields arXiv covers thinly (most of medicine, the humanities and social sciences).

## Two sets of tells, half-year by half-year

Share of abstracts using at least one word of each twelve-word set. Both sets were fixed before the census was measured.

## The frozen sets, tested

Papers outside the pilot. Baseline papers revised after 30 November 2022 are left out. Abstracts: 446,418 (2020 to 2022), 233,900 (2024), 262,617 (2026, January to September). Intervals: 95%, bootstrap over subject categories.

Word set · 2020 to 2022 · 2024 · 2026 · 2024 minus baseline · 2026 minus 2024
2024 tells · 4.54% · 11.85% · 3.03% · +7.31 (+5.4 to +8.8) · -8.82 (-11.0 to -6.3)
2026 tells · 1.90% · 3.31% · 12.39% · +1.41 (+0.9 to +1.8) · +9.08 (+6.8 to +10.9)
Control: however, results, using, show, method · 73.48% · 74.88% · 72.64% · +1.39 (+0.8 to +1.8) · -2.24 (-3.7 to -0.4)

## By subject group

Share of abstracts using a 2026 tell, by the arXiv group of the paper's first category, and the lower bound from both sets of 50 style words. 'Too uncertain' marks a bound whose 95% interval is wider than 40 points, which happens in the smaller groups.

Group · 2026 tells, 2020 to 2022 · 2026 tells, 2026 · Times baseline · Lower bound 2024 · Lower bound 2026
Computer science · 2.46% · 18.70% · 7.6 · 54% · 80%
Electrical engineering and systems · 1.68% · 9.56% · 5.7 · 46% · too uncertain
Statistics · 2.22% · 11.53% · 5.2 · 46% · too uncertain
Physics and astronomy · 1.65% · 6.18% · 3.7 · 29% · 63%
Mathematics · 1.32% · 4.64% · 3.5 · 14% · 49%
Quantitative biology · 2.19% · 15.13% · 6.9 · too uncertain · too uncertain
Quantitative finance · 1.70% · 11.45% · 6.8 · too uncertain · too uncertain
Economics · 1.38% · 9.77% · 7.1 · too uncertain · too uncertain

## Three 2024 tells

## Three 2026 tells

## The style words that rose most, 2024

Top 15 of the 50, scored on even-month papers: the share predicted by the word's 2021 to 2022 trend, the share observed, and the ratio.

Word · Predicted · Observed · Ratio
challenges · 4.82% · 9.71% · 2.0x
additionally · 2.58% · 7.53% · 2.9x
findings · 3.33% · 7.42% · 2.2x
enhance · 2.23% · 6.61% · 3.0x
insights · 2.15% · 5.37% · 2.5x
crucial · 3.05% · 6.50% · 2.1x
comprehensive · 2.51% · 6.29% · 2.5x
enhancing · 0.66% · 3.93% · 6.0x
capabilities · 2.05% · 5.78% · 2.8x
introduces · 2.12% · 4.86% · 2.3x
effectively · 3.32% · 6.63% · 2.0x
particularly · 2.48% · 5.61% · 2.3x
diverse · 2.72% · 6.32% · 2.3x
leveraging · 1.91% · 4.67% · 2.5x
enabling · 1.85% · 4.52% · 2.4x

## The style words that rose most, 2026

Word · Predicted · Observed · Ratio
remains · 3.40% · 12.70% · 3.7x
yet · 3.82% · 11.41% · 3.0x
rather · 2.12% · 8.78% · 4.2x
establish · 2.73% · 8.31% · 3.0x
explicit · 2.57% · 8.07% · 3.1x
enabling · 2.13% · 7.97% · 3.7x
every · 2.38% · 6.65% · 2.8x
remain · 1.78% · 6.86% · 3.9x
fixed · 2.82% · 7.01% · 2.5x
yields · 1.67% · 6.18% · 3.7x
improves · 3.41% · 8.03% · 2.4x
practical · 3.52% · 7.26% · 2.1x
structured · 1.00% · 5.42% · 5.4x
findings · 3.67% · 7.28% · 2.0x
consistently · 1.04% · 5.38% · 5.2x

## Lower bounds by year

Year · 2024 words · 2026 words · Both
2023 · 13% (8 to 17%) · 12% (9 to 14%) · 17% (11 to 22%)
2024 · 29% (20 to 37%) · 23% (18 to 27%) · 33% (25 to 40%)
2025 · 40% (31 to 47%) · 48% (41 to 53%) · 51% (43 to 58%)
2026 · 20% (13 to 26%) · 68% (62 to 73%) · 64% (56 to 70%)

## Placebo: the same method before language models

The method shifted five years earlier: 2016 to 2017 as the trend, 2018 to 2021 as the 'after' years. The true answer is zero, so anything above zero is drift the method mistakes for model use.

Words · Placebo 'bound' at 2021 (like 2026) · Interval
The 2026 words · 0.5% · -9 to 10%
The 2024 words · 10.2% · 4 to 16%
50 words re-selected on 2016 to 2021 data · 10.1% · -2 to 19%

## Sensitivity: all papers, revised ones included

Word set · 2020 to 2022 · 2024 · 2026
2024 tells · 4.52% · 11.85% · 3.03%
2026 tells · 1.94% · 3.31% · 12.39%

## Other measures, 2020 to 2022 against 2026

Measure · 2020 to 2022 · 2026 · Difference (95% interval)
Words per abstract · 161.80 · 174.99 · +13.20 (+10.02 to +16.21)
Average sentence length (words) · 24.46 · 24.49 · +0.03 (-0.14 to +0.20)
Commas per sentence · 1.05 · 1.38 · +0.34 (+0.28 to +0.37)
Semicolons per 1,000 words · 0.57 · 1.05 · +0.48 (+0.37 to +0.58)
ScriptGrain AI-tell density per 1,000 words · 1.31 · 1.43 · +0.12 (+0.09 to +0.16)
Em dashes per 1,000 words · 0.02 · 0.07 · +0.05 (+0.04 to +0.06)

## What 43 earlier studies estimated

We searched Firecrawl's Research Index with 20 fixed queries, screened 1,337 results to the studies that estimate how much scholarly text is written or edited with language models, and checked each paper exists and each number against its own text. 43 qualify. Methods differ (20 per-document detector, 13 excess vocabulary / word frequency, 5 distributional (population-level) model, 2 keyword floor, 2 other, 1 classifier trained on polished text), so the numbers are not one series. Detector-based studies often report a high rate before ChatGPT existed, so their figures are not shares of model-written text. The full table, with each quotation, is a download below.

Study · Corpus · Estimate
Mingmeng Geng (2024) · One million arXiv papers (abstracts) · Fraction of LLM-style abstracts about 35% in computer science
Gwinyai Masukume (2024) · PubMed and Scopus records for "delve", "realm" and "underscore" · Co-use of "delve", "realm" and "underscore" up to 85-fold higher in 2023-2024 than in 2022 and earlier
Dmitry Kobak (2024) · More than 15 million PubMed biomedical abstracts · At least 13.5% of 2024 PubMed abstracts processed with LLMs (lower bound; up to 40% in some subcorpora)
Sergio E Uribe (2024) · 299,695 dental research abstracts indexed in PubMed · Abstracts with ChatGPT "signaling words": 47.1 per 10,000 before release vs 224.2 per 10,000 after
Weixin Liang (2024) · 1,121,912 preprints and published papers on arXiv, bioRxiv and Nature portfolio journals · LLM-modified content up to 22% of computer science papers; up to 9% in mathematics and Nature portfolio
Kayvan Kousha (2025) · Abstracts in six databases (Scopus, Web of Science, PubMed, PMC, Dimensions, OpenAlex); 2.4 million PMC open-access full · Frequency of "delve" up 1,500%, "underscore" up 1,000%, "intricate" up 700% between 2022 and 2024; not a share of papers
Andrew Gray (2025) · Full text of published papers across all research fields indexed in Dimensions · LLM tools likely involved in more than 10% of all published papers in 2024
A. Arezki (2025) · 64,444 unique abstracts from 38 urology journals (PubMed) · Proportion of AI-like text 1.8% in 2020 and 5.3% in 2024
Morgan D. Sanger (2026) · 149,452 abstracts published by the American Society of Civil Engineers (journals and proceedings), civil and environment · 13.9% (2024) and 20.4% (2025) of abstracts likely LLM-written (conservative); frequency-shift estimate 15.3% and 26.2%
Mike Thelwall (2026) · 1.25 million MDPI articles (full text); 80 LLM-associated terms; 73 journals with 500+ articles in 2021 examined separat · Term family "underscore" increased up to 29-fold; word-frequency change, not a share of papers
Przemysław Czuma (2026) · 69,632 first-version medRxiv preprints with an extractable Discussion section (full-text XML) · Em-dash in Discussion sections rose from 4.23% of preprints before ChatGPT to 11.58% after (20.3% in 2025)
Lena Holzwarth (2026) · Full texts of open-access biomedical papers from PubMed Central · 89% of papers show excess LLM-associated vocabulary by end of 2025; 68% of Discussion paragraphs vs 32% of Methods paragraphs
Aron Lee (2026) · 398,296 Korean-language abstracts (KCI), with 47,165 Vietnamese abstracts for comparison · Lower bound on LLM-processed Korean abstracts 3.5%, 10.5% and 16.1% for 2024-2026 (single-word); 7.8%, 20.6%, 33.0% (split-half set)
Serat M. Saad (2026) · Full text of 207,111 astro-ph papers; 392 papers that disclose model use calibrate the assisted rate · About 54% of 2025 astro-ph papers carry a language-model trace (at or above 36% under varied assumptions); 0.81% disclose it
M. Z. Naser (2026) · 221,425 funded grant abstracts from NSF (n = 96,020), NIH (n = 80,496) and UKRI (n = 44,909) · Excess marker-word rate in post-ChatGPT grant abstracts of 4.0-8.4% (NSF), 1.4-3.7% (NIH), 2.8-6.8% (UKRI) above the pre-ChatGPT baseline
Kyle Siler (2026) · Full texts of 7.3 million journal articles from four publishers (Elsevier, Frontiers, MDPI, PLoS); 228 focal words · An estimated 57% of published articles in 2025 (12% in 2023) showed evidence of LLM influence
Hadar Better (2026) · 2,770 abstracts from six major dental journals · High-suspicion abstracts (AI-marker score of 20 or more) rose from 18.0% to 25.8%; AI marker words per abstract rose from 0.69 to 0.97
Sota Nakamura (2026) · Articles in Surgery Today and Surgical Case Reports · Rare AI-related word density rose after 2023 (slope change beta = 1.05 per year); no share estimated

## Limitations

- arXiv only. It is mostly physics, mathematics and computer science, and biomedicine is thin; the published PubMed studies cover that ground.
- Word shares describe a body of writing. They cannot say which papers used a model, and people also pick up model vocabulary by reading it.
- The topic mix changed: far more 2026 papers are about language models and agents. Style words were separated from topic words by labelling, and the new set rose in mathematics and physics too, but some topic effect may remain in the bounds.
- The lower bounds assume each word's 2021 to 2022 trend would have continued in a straight line. The placebo shows ordinary drift can add about 10 points over four years, more for the 2024 words, which were already rising.
- arXiv serves each paper's latest abstract. Baseline papers revised after November 2022 were removed; for 2024 and 2026 the abstract may postdate first submission (31% of 2024 papers have a version more than 90 days after the first).
- The 2026 word set came from a pilot sample chosen by semantic search, not at random. That is why the census test excludes the pilot papers.
- Content or style labels were assigned by Claude models (agreement 343 of 351 words). A different labeller would move a few words between the lists.
- 2026 covers January to September only.

## Competing interests

ScriptGrain sells writing-voice measurement and generation, and the analyst builds it. The hypotheses, word lists and tests were written down and committed before the census was measured; every change after that is listed in the pre-registration with its reason. Word labelling and the literature extraction used Claude models; their outputs are published so they can be checked.

## Reproducing this

- Pre-registration, every script and every result file: docs/research/sgr-008-preregistration.md and scripts/experiments/sgr-008/ in the ScriptGrain repository (census-harvest.py, census-build.py, census-analyse.py, census-instruments.mjs, census-instruments-analyse.py, census-bounds.py, census-placebo.py, write-entry.py).
- Source: arXiv's OAI-PMH service (oaipmh.arxiv.org), format arXivRaw. arXiv metadata is CC0, so anyone can harvest the same records and re-run the scripts.
- Instruments: supabase/functions/_shared/voice-match.ts (measureTextFeatures) and ai-cliches.ts (scanCliches), the code the product runs.
- Abstract text is not republished; the downloads carry counts and shares only.

## Data

- [Word set shares by half-year, 2015 to 2026 (CSV)](/research/sgr-008-half-year.csv) (csv)
- [Every candidate and frozen word, share of abstracts by year, with labels (CSV)](/research/sgr-008-words-by-year.csv) (csv)
- [Excess-word statistics, selection and scoring halves (CSV)](/research/sgr-008-excess-words.csv) (csv)
- [Literature evidence table, 43 studies with quotations (CSV)](/research/sgr-008-literature.csv) (csv)
- [Every result, including intervals, bounds and placebo (JSON)](/research/sgr-008-summary.json) (json)

Released under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/): free to reuse, including commercially, with credit to ScriptGrain and a link to this page.

## Citations

- [arXiv OAI-PMH service](https://info.arxiv.org/help/oa/index.html)
- [Kobak et al., Delving into LLM-assisted writing in biomedical publications through excess vocabulary (Science Advances, 2025; arXiv 2406.07016)](https://arxiv.org/abs/2406.07016)
- [Liang et al., Quantifying large language model usage in scientific papers (Nature Human Behaviour, 2025; arXiv 2404.01268)](https://arxiv.org/abs/2404.01268)
- [Geng and Trotta, Human-LLM coevolution: evidence from academic writing (arXiv 2502.09606)](https://arxiv.org/abs/2502.09606)
- [How much are LLMs changing the language of academic papers after ChatGPT? (arXiv 2509.09596)](https://arxiv.org/abs/2509.09596)
- [More than half of recent astronomy papers are written with language-model assistance (arXiv 2609.10664)](https://arxiv.org/abs/2609.10664)
- [Firecrawl Research Index](https://docs.firecrawl.dev/api-reference/endpoint/research-search-papers)
- [ScriptGrain SGR-007, Where did the AI-isms come from? (uses this study's census and word lists)](https://scriptgrain.com/research/where-did-ai-isms-come-from)

## How to cite

ScriptGrain (2026). The AI tells moved: 2.1 million arXiv abstracts, 2015 to 2026 (Study SGR-008, conducted 4 October 2026). Dataset licensed CC BY 4.0. https://scriptgrain.com/research/the-ai-tells-moved

## About the analyst

Jack Stovell has worked in finance and data for more than twelve years, building management reporting, forecasting and profitability models for advertising agencies and tech scale-ups, from SQL and Power BI reporting to board-level analysis. Since 2016 he has run Adapt Progress Evolve, an applied AI studio, where he builds and operates AI systems and data products: ScriptGrain's measurement of writing voice across 45 attributes, UK Spend, which brings 16.7 million rows of UK council spending into one queryable dataset, and more than thirty AI agents running in production. He designs ScriptGrain's studies and is accountable for every number in them.

- [jstov.uk](https://jstov.uk)
- [LinkedIn](https://www.linkedin.com/in/jackstovell)
- [GitHub](https://github.com/StovBuilds)

## About ScriptGrain (not a finding)

ScriptGrain measures writing voice: the habits that make a writer recognisable, such as sentence rhythm, punctuation and vocabulary. The AI-tell scan used here is the one behind the free [AI cliché checker](https://scriptgrain.com/tools/ai-cliche-checker). Neither is a test of who, or what, wrote a text.
