The words ChatGPT overuses, with the evidence
Search for 'words ChatGPT overuses' and you will find hundreds of lists. Almost none cite a source, and most are copies of copies: a 2023 observation about one model, pasted forward year after year as if nothing had changed. This page is the other kind of list. Every word on it carries its evidence: which study measured it, in what corpus, at what magnitude where one exists, and, critically, which model generation the evidence dates from.
That last part matters more than it sounds. Overuse profiles are not stable. 'Delve', the most famous tell of all, was heavily overused in 2023 and early 2024, became less frequent later in 2024, then dropped off sharply in 2025, according to the editors who maintain Wikipedia's catalogue of AI writing signs. A list frozen in 2023 is describing a model that no longer answers your prompts.
One framing note before the table: this is a measurement page, not a detection tool and not an evasion guide. Finding these words in a text does not make the text AI-written, and removing them does not make a text human. ScriptGrain measures style; the point of this page is to show what the evidence actually supports.
What counts as evidence here
Each entry below is tagged with one or more evidence types, in descending order of rigour. Corpus study: a peer-reviewed frequency analysis over a defined corpus. The anchor here is Juzek and Ward's COLING 2025 paper 'Why Does ChatGPT Delve So Much?', which analysed 26.7 million PubMed abstracts (5.2 billion tokens, 1975 to May 2024) and identified 21 focal words whose post-2022 frequency spikes are likely the result of LLM usage. A companion at larger scale is Kobak and colleagues' Science Advances study of over 15 million biomedical abstracts (2010 to 2024), which identified around 900 'excess vocabulary' words and estimated that at least 13.5% of 2024 abstracts were processed with LLMs, rising to roughly 40% in some subcorpora.
Speech corpus: measured spoken usage. The Max Planck Institute for Human Development analysed 740,249 hours of human speech (360,445 YouTube academic talks and 771,591 podcast episodes, over 7.35 billion transcribed words) to track ChatGPT-preferred words entering spontaneous conversation.
Community catalogue: the continuously edited word lists on Wikipedia's 'Signs of AI writing' project page. These are editor observations, not peer-reviewed measurements, but they have one property no published study matches: they are re-dated per model generation, which makes them the best public record of how tells drift. Where an entry below rests only on the catalogue, treat the magnitude as unquantified.
The list
Magnitudes are specific to the corpus they were measured in; a 6,697% rise in scientific abstracts does not mean the word rose 6,697% in blog posts. Percentage figures from Juzek and Ward compare 2020 with 2024 frequencies in PubMed abstracts. Model generations follow the eras used by Wikipedia's catalogue: GPT-3.5/GPT-4 era (2023 to mid-2024), GPT-4o era (mid-2024 to mid-2025), GPT-5 era (mid-2025 onward).
| Word or phrase | Evidence type | What the evidence says | Model generation of the evidence |
|---|---|---|---|
| delve (delves, delved, delving) | Corpus study + speech corpus + community catalogue | 'delves' rose 6,697% in PubMed abstracts from 2020 to 2024 (Juzek and Ward); also rose significantly in academic YouTube talks (p=0.010, Max Planck); Wikipedia editors record it dropping off sharply in 2025 | GPT-3.5/GPT-4 era; fading since late 2024 |
| intricate / intricacies | Corpus study + speech corpus + community catalogue | 'intricate' rose 611% in PubMed abstracts from 2020 to 2024; 'intricacies' also increased abruptly in spontaneous human speech | GPT-3.5/GPT-4 era |
| underscore / underscores | Corpus study + community catalogue | 'underscores' rose 904% in PubMed abstracts from 2020 to 2024; still catalogued through the GPT-4o era | GPT-3.5/GPT-4 and GPT-4o eras |
| showcase / showcasing | Corpus study + community catalogue | one of Juzek and Ward's 21 focal words; the only stem Wikipedia's catalogue lists across all three eras up to GPT-5 | GPT-3.5 through GPT-5 era (persistent) |
| boast / boasts | Corpus study + speech corpus + community catalogue | focal word in PubMed abstracts; 'boast' is among the top GPT words growing roughly 25% to 50% per year in spoken usage | GPT-3.5/GPT-4 era |
| realm | Corpus study | one of the 21 focal words; named by the FSU researchers alongside 'delve' and 'intricate' as a signature spike | GPT-3.5/GPT-4 era |
| garner / garnered | Corpus study + community catalogue | 'garnered' is a focal word in PubMed abstracts; 'garner' appears in the catalogue's 2023 to mid-2024 list | GPT-3.5/GPT-4 era |
| emphasizing | Corpus study + community catalogue | focal word in PubMed abstracts; catalogued in every Wikipedia era list including GPT-5 | GPT-3.5 through GPT-5 era (persistent) |
| comprehend / comprehending | Corpus study + speech corpus | 'comprehending' is a focal word; 'comprehend' is among the top GPT words growing roughly 25% to 50% per year in spoken usage | GPT-3.5/GPT-4 era |
| surpass / surpasses / surpassing | Corpus study | two of the 21 focal words ('surpasses', 'surpassing') in PubMed abstracts | GPT-3.5/GPT-4 era |
| groundbreaking | Corpus study | one of the 21 focal words in PubMed abstracts | GPT-3.5/GPT-4 era |
| advancements | Corpus study | one of the 21 focal words in PubMed abstracts | GPT-3.5/GPT-4 era |
| align / aligns / align with | Corpus study + community catalogue | 'aligns' is a focal word; the phrase 'align with' enters the catalogue in the mid-2024 to mid-2025 list | GPT-3.5/GPT-4 and GPT-4o eras |
| meticulous / meticulously | Speech corpus + community catalogue | among the top GPT words growing roughly 25% to 50% per year in spoken usage; catalogued for 2023 to mid-2024 | GPT-3.5/GPT-4 era |
| swift | Speech corpus | among the words showing a significant post-ChatGPT increase in spontaneous human speech | GPT-3.5/GPT-4 era |
| crucial | Community catalogue | catalogued in both the GPT-4 era and GPT-4o era lists; core of the puffery pattern 'plays a crucial role' | GPT-3.5/GPT-4 and GPT-4o eras |
| pivotal | Community catalogue | catalogued in both the GPT-4 era and GPT-4o era lists; twin of 'crucial' in inflated-significance phrasing | GPT-3.5/GPT-4 and GPT-4o eras |
| tapestry | Community catalogue | catalogued for 2023 to mid-2024 only; absent from later era lists, a worked example of a decayed tell | GPT-3.5/GPT-4 era (dated) |
| testament | Community catalogue | catalogued for 2023 to mid-2024, typically as 'stands as a testament'; absent from later lists | GPT-3.5/GPT-4 era (dated) |
| vibrant | Community catalogue | catalogued in both the GPT-4 era and GPT-4o era lists | GPT-3.5/GPT-4 and GPT-4o eras |
| landscape | Community catalogue | catalogued for 2023 to mid-2024, familiar from 'the ever-evolving landscape'; absent from later lists | GPT-3.5/GPT-4 era (dated) |
| interplay | Community catalogue | catalogued for 2023 to mid-2024; absent from later era lists | GPT-3.5/GPT-4 era (dated) |
| bolstered | Community catalogue | catalogued in both the GPT-4 era and GPT-4o era lists | GPT-3.5/GPT-4 and GPT-4o eras |
| enduring | Community catalogue | catalogued in both the GPT-4 era and GPT-4o era lists | GPT-3.5/GPT-4 and GPT-4o eras |
| foster / fostering | Community catalogue | 'fostering' enters the catalogue in the mid-2024 to mid-2025 list | GPT-4o era |
| enhance | Community catalogue | catalogued in the GPT-4o era list and carried into the GPT-5 era list | GPT-4o and GPT-5 eras (current) |
| highlighting | Community catalogue | catalogued in the GPT-4o era list and carried into the GPT-5 era list | GPT-4o and GPT-5 eras (current) |
| additionally | Community catalogue | catalogued for 2023 to mid-2024 as a connective overused at paragraph openings | GPT-3.5/GPT-4 era |
| valuable / valuable insights | Community catalogue | catalogued for 2023 to mid-2024; 'valuable insights' also flagged as an editorialising marker | GPT-3.5/GPT-4 era |
Why these words: the hypotheses
Juzek and Ward tested the obvious explanations and eliminated most of them. They report failing to find evidence that lexical overrepresentation stems from model architecture or algorithmic choices. They also report failing to find evidence that it comes from training or fine-tuning data: the focal words appear far less frequently in corpora such as arXiv and Wikipedia than they do in ChatGPT output, so the model is not simply mirroring what it read.
The hypothesis left standing, though not proven, is reinforcement learning from human feedback. In their model testing, Llama 2-Chat, the RLHF-tuned variant, showed considerably less surprise at text laden with focal words than its base model did, which is consistent with preference tuning teaching models that this vocabulary is what good answers sound like. The FSU team's framing is that human preference data is a source of the lexical preferences, the word choices, of these large language models. That is a hypothesis about the pipeline, not a demonstrated mechanism.
Their human experiment complicates the story further: participants appeared to react differently to 'delve' than to other focal words, so a single cause may not explain every word on the list. The authors are explicit that the lack of transparency surrounding model development remains an obstacle to settling the question. This page reports these as hypotheses because that is what they are.
Tells decay: a 2023 list is not a 2026 list
The strongest evidence that word lists rot comes from the one source that dates its entries. Wikipedia's catalogue lists roughly nineteen overused words for the 2023 to mid-2024 era, about twelve for mid-2024 to mid-2025, and only four (emphasizing, enhance, highlighting, showcasing) for mid-2025 onward. The list did not just shrink; it churned.
Concrete drift, per the catalogue: 'delve' was famously overused in 2023 and early 2024, became less frequent later in 2024, then dropped off sharply in 2025. 'Tapestry', 'testament', 'landscape' and 'interplay' are catalogued only for the GPT-4 era and vanish from later lists. Meanwhile 'enhance' and 'highlighting' only enter with the GPT-4o era. A checker still scanning for 'rich tapestry' in 2026 is auditing a retired model.
The corpus studies can track the same churn quantitatively: the excess-vocabulary dataset behind the Kobak study has been updated with monthly values through July 2025 precisely because yearly snapshots go stale. The practical rule: any list of AI words that does not say which model generation it describes should be assumed to describe GPT-3.5 or GPT-4, and to be partly obsolete.
The spillover: these words are entering human speech
The Max Planck Institute for Human Development analysed 740,249 hours of human discourse, 360,445 YouTube academic talks and 771,591 podcast episodes screened for unscripted speech, covering four years before ChatGPT's release through May 2024. Words preferentially generated by ChatGPT, including 'delve', 'comprehend', 'boast', 'swift' and 'meticulous', increased abruptly in spontaneous human speech after the release, with top GPT words growing at roughly 25% to 50% per year. 'Delve' rose significantly in academic YouTube talks (p=0.010) and in science and technology, business and education podcasts.
The team also ran a preregistered experiment with 496 participants: a brief interaction with a chatbot led participants to adopt its words as their own, and the effect persisted past a distractor task. Co-author Levin Brinkmann put it plainly: 'The patterns that are stored in AI technology seem to be transmitting back to the human mind.'
This is the quiet killer of word-level tells. A marker works because of a frequency gap between model output and human baseline; as humans absorb the vocabulary, the gap closes from both ends. Worse, it feeds back: the FSU researchers note that AI-influenced human writing re-enters the training data of future models, reinforcing the very patterns being measured. Every year a word spends on a viral 'AI words' list, it becomes a weaker marker, partly because humans read the list too.
What this list cannot tell you
Using these words does not make a text AI-written. Humans wrote 'delve' for centuries before November 2022, and academic English used every word in the table at some baseline long before LLMs existed. The corpus studies measure frequency shifts across millions of documents; they license conclusions about corpora, not verdicts on individual texts. Kobak and colleagues' estimate that at least 13.5% of 2024 biomedical abstracts were LLM-processed is a corpus-level lower bound, not a test you can run on one abstract.
The lists are also literal. Wikipedia's catalogue stresses that a word being overused by AI does not imply its synonyms are overused; the evidence attaches to specific tokens in specific corpora, and base rates differ across fields, registers and dialects.
Finally, the traffic runs both ways in practice: deleting listed words from a text does not remove the other statistical regularities these studies measure, and sprinkling them in does not make human text machine-made. This page exists to document what models actually do and how fast that changes, not to referee authorship disputes. If you want to see how any text, yours or a model's, actually distributes its vocabulary, that is a measurement job, and measurement is what ScriptGrain does.
Methodology
Compiled from peer-reviewed corpus studies (Juzek and Ward, COLING 2025; Kobak et al, Science Advances 2025), a large-scale speech corpus analysis (Yakura et al, Max Planck Institute for Human Development), and the editor-maintained catalogue at Wikipedia's 'Signs of AI writing' project page. Every magnitude on this page traces to one of the listed sources; nothing is reproduced from unsourced listicles.
Evidence standard: corpus studies outrank the speech corpus, which outranks the community catalogue; catalogue-only entries carry no magnitude and are labelled as such. Percentage increases from Juzek and Ward compare 2020 with 2024 word frequencies in PubMed abstracts and apply to that corpus only.
Last checked 2026-08-03. Update policy: the table is re-checked when a major model generation ships, because overuse profiles change between generations. Entries whose evidence no longer holds are re-dated and kept as documented decay, not silently deleted.
Sources
- Juzek and Ward, 'Why Does ChatGPT Delve So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models', COLING 2025
- Juzek and Ward, full text preprint (arXiv 2412.11385)
- Florida State University News, 'Why does ChatGPT delve so much?' (17 February 2025)
- Wikipedia, 'Signs of AI writing' (project page, continuously updated)
- Yakura et al, 'Empirical evidence of Large Language Models' influence on human spoken communication', Max Planck Institute for Human Development (arXiv 2409.01754)
- Kobak et al, 'Delving into LLM-assisted writing in biomedical publications through excess vocabulary', Science Advances (2025)
- berenslab/llm-excess-vocab: excess-words dataset and monthly updates through July 2025
- Scientific American, 'ChatGPT Is Changing the Words We Use in Conversation'
ScriptGrain · Why ScriptGrain · Sample voice profiles · The Journal · Support