The words ChatGPT overuses, with the evidence

Search for 'words ChatGPT overuses' and you will find hundreds of lists. Almost none cite a source, and most are copies of copies: a 2023 observation about one model, pasted forward year after year as if nothing had changed. This page is the other kind of list. Every word on it carries its evidence: which study measured it, in what corpus, at what magnitude where one exists, and, critically, which model generation the evidence dates from.

That last part matters more than it sounds. Overuse profiles are not stable. 'Delve', the most famous tell of all, was heavily overused in 2023 and early 2024, became less frequent later in 2024, then dropped off sharply in 2025, according to the editors who maintain Wikipedia's catalogue of AI writing signs. A list frozen in 2023 is describing a model that no longer answers your prompts.

One framing note before the table: this is a measurement page, not a detection tool and not an evasion guide. Finding these words in a text does not make the text AI-written, and removing them does not make a text human. ScriptGrain measures style; the point of this page is to show what the evidence actually supports.

What counts as evidence here

Each entry below is tagged with one or more evidence types, in descending order of rigour. Corpus study: a peer-reviewed frequency analysis over a defined corpus. The anchor here is Juzek and Ward's COLING 2025 paper 'Why Does ChatGPT Delve So Much?', which analysed 26.7 million PubMed abstracts (5.2 billion tokens, 1975 to May 2024) and identified 21 focal words whose post-2022 frequency spikes are likely the result of LLM usage. A companion at larger scale is Kobak and colleagues' Science Advances study of over 15 million biomedical abstracts (2010 to 2024), which identified around 900 'excess vocabulary' words and estimated that at least 13.5% of 2024 abstracts were processed with LLMs, rising to roughly 40% in some subcorpora.

Speech corpus: measured spoken usage. The Max Planck Institute for Human Development analysed 740,249 hours of human speech (360,445 YouTube academic talks and 771,591 podcast episodes, over 7.35 billion transcribed words) to track ChatGPT-preferred words entering spontaneous conversation.

Community catalogue: the continuously edited word lists on Wikipedia's 'Signs of AI writing' project page. These are editor observations, not peer-reviewed measurements, but they have one property no published study matches: they are re-dated per model generation, which makes them the best public record of how tells drift. Where an entry below rests only on the catalogue, treat the magnitude as unquantified.

The list

Magnitudes are specific to the corpus they were measured in; a 6,697% rise in scientific abstracts does not mean the word rose 6,697% in blog posts. Percentage figures from Juzek and Ward compare 2020 with 2024 frequencies in PubMed abstracts. Model generations follow the eras used by Wikipedia's catalogue: GPT-3.5/GPT-4 era (2023 to mid-2024), GPT-4o era (mid-2024 to mid-2025), GPT-5 era (mid-2025 onward).

Word or phraseEvidence typeWhat the evidence saysModel generation of the evidence
delve (delves, delved, delving)Corpus study + speech corpus + community catalogue'delves' rose 6,697% in PubMed abstracts from 2020 to 2024 (Juzek and Ward); also rose significantly in academic YouTube talks (p=0.010, Max Planck); Wikipedia editors record it dropping off sharply in 2025GPT-3.5/GPT-4 era; fading since late 2024
intricate / intricaciesCorpus study + speech corpus + community catalogue'intricate' rose 611% in PubMed abstracts from 2020 to 2024; 'intricacies' also increased abruptly in spontaneous human speechGPT-3.5/GPT-4 era
underscore / underscoresCorpus study + community catalogue'underscores' rose 904% in PubMed abstracts from 2020 to 2024; still catalogued through the GPT-4o eraGPT-3.5/GPT-4 and GPT-4o eras
showcase / showcasingCorpus study + community catalogueone of Juzek and Ward's 21 focal words; the only stem Wikipedia's catalogue lists across all three eras up to GPT-5GPT-3.5 through GPT-5 era (persistent)
boast / boastsCorpus study + speech corpus + community cataloguefocal word in PubMed abstracts; 'boast' is among the top GPT words growing roughly 25% to 50% per year in spoken usageGPT-3.5/GPT-4 era
realmCorpus studyone of the 21 focal words; named by the FSU researchers alongside 'delve' and 'intricate' as a signature spikeGPT-3.5/GPT-4 era
garner / garneredCorpus study + community catalogue'garnered' is a focal word in PubMed abstracts; 'garner' appears in the catalogue's 2023 to mid-2024 listGPT-3.5/GPT-4 era
emphasizingCorpus study + community cataloguefocal word in PubMed abstracts; catalogued in every Wikipedia era list including GPT-5GPT-3.5 through GPT-5 era (persistent)
comprehend / comprehendingCorpus study + speech corpus'comprehending' is a focal word; 'comprehend' is among the top GPT words growing roughly 25% to 50% per year in spoken usageGPT-3.5/GPT-4 era
surpass / surpasses / surpassingCorpus studytwo of the 21 focal words ('surpasses', 'surpassing') in PubMed abstractsGPT-3.5/GPT-4 era
groundbreakingCorpus studyone of the 21 focal words in PubMed abstractsGPT-3.5/GPT-4 era
advancementsCorpus studyone of the 21 focal words in PubMed abstractsGPT-3.5/GPT-4 era
align / aligns / align withCorpus study + community catalogue'aligns' is a focal word; the phrase 'align with' enters the catalogue in the mid-2024 to mid-2025 listGPT-3.5/GPT-4 and GPT-4o eras
meticulous / meticulouslySpeech corpus + community catalogueamong the top GPT words growing roughly 25% to 50% per year in spoken usage; catalogued for 2023 to mid-2024GPT-3.5/GPT-4 era
swiftSpeech corpusamong the words showing a significant post-ChatGPT increase in spontaneous human speechGPT-3.5/GPT-4 era
crucialCommunity cataloguecatalogued in both the GPT-4 era and GPT-4o era lists; core of the puffery pattern 'plays a crucial role'GPT-3.5/GPT-4 and GPT-4o eras
pivotalCommunity cataloguecatalogued in both the GPT-4 era and GPT-4o era lists; twin of 'crucial' in inflated-significance phrasingGPT-3.5/GPT-4 and GPT-4o eras
tapestryCommunity cataloguecatalogued for 2023 to mid-2024 only; absent from later era lists, a worked example of a decayed tellGPT-3.5/GPT-4 era (dated)
testamentCommunity cataloguecatalogued for 2023 to mid-2024, typically as 'stands as a testament'; absent from later listsGPT-3.5/GPT-4 era (dated)
vibrantCommunity cataloguecatalogued in both the GPT-4 era and GPT-4o era listsGPT-3.5/GPT-4 and GPT-4o eras
landscapeCommunity cataloguecatalogued for 2023 to mid-2024, familiar from 'the ever-evolving landscape'; absent from later listsGPT-3.5/GPT-4 era (dated)
interplayCommunity cataloguecatalogued for 2023 to mid-2024; absent from later era listsGPT-3.5/GPT-4 era (dated)
bolsteredCommunity cataloguecatalogued in both the GPT-4 era and GPT-4o era listsGPT-3.5/GPT-4 and GPT-4o eras
enduringCommunity cataloguecatalogued in both the GPT-4 era and GPT-4o era listsGPT-3.5/GPT-4 and GPT-4o eras
foster / fosteringCommunity catalogue'fostering' enters the catalogue in the mid-2024 to mid-2025 listGPT-4o era
enhanceCommunity cataloguecatalogued in the GPT-4o era list and carried into the GPT-5 era listGPT-4o and GPT-5 eras (current)
highlightingCommunity cataloguecatalogued in the GPT-4o era list and carried into the GPT-5 era listGPT-4o and GPT-5 eras (current)
additionallyCommunity cataloguecatalogued for 2023 to mid-2024 as a connective overused at paragraph openingsGPT-3.5/GPT-4 era
valuable / valuable insightsCommunity cataloguecatalogued for 2023 to mid-2024; 'valuable insights' also flagged as an editorialising markerGPT-3.5/GPT-4 era

Why these words: the hypotheses

Juzek and Ward tested the obvious explanations and eliminated most of them. They report failing to find evidence that lexical overrepresentation stems from model architecture or algorithmic choices. They also report failing to find evidence that it comes from training or fine-tuning data: the focal words appear far less frequently in corpora such as arXiv and Wikipedia than they do in ChatGPT output, so the model is not simply mirroring what it read.

The hypothesis left standing, though not proven, is reinforcement learning from human feedback. In their model testing, Llama 2-Chat, the RLHF-tuned variant, showed considerably less surprise at text laden with focal words than its base model did, which is consistent with preference tuning teaching models that this vocabulary is what good answers sound like. The FSU team's framing is that human preference data is a source of the lexical preferences, the word choices, of these large language models. That is a hypothesis about the pipeline, not a demonstrated mechanism.

Their human experiment complicates the story further: participants appeared to react differently to 'delve' than to other focal words, so a single cause may not explain every word on the list. The authors are explicit that the lack of transparency surrounding model development remains an obstacle to settling the question. This page reports these as hypotheses because that is what they are.

Tells decay: a 2023 list is not a 2026 list

The strongest evidence that word lists rot comes from the one source that dates its entries. Wikipedia's catalogue lists roughly nineteen overused words for the 2023 to mid-2024 era, about twelve for mid-2024 to mid-2025, and only four (emphasizing, enhance, highlighting, showcasing) for mid-2025 onward. The list did not just shrink; it churned.

Concrete drift, per the catalogue: 'delve' was famously overused in 2023 and early 2024, became less frequent later in 2024, then dropped off sharply in 2025. 'Tapestry', 'testament', 'landscape' and 'interplay' are catalogued only for the GPT-4 era and vanish from later lists. Meanwhile 'enhance' and 'highlighting' only enter with the GPT-4o era. A checker still scanning for 'rich tapestry' in 2026 is auditing a retired model.

The corpus studies can track the same churn quantitatively: the excess-vocabulary dataset behind the Kobak study has been updated with monthly values through July 2025 precisely because yearly snapshots go stale. The practical rule: any list of AI words that does not say which model generation it describes should be assumed to describe GPT-3.5 or GPT-4, and to be partly obsolete.

The spillover: these words are entering human speech

The Max Planck Institute for Human Development analysed 740,249 hours of human discourse, 360,445 YouTube academic talks and 771,591 podcast episodes screened for unscripted speech, covering four years before ChatGPT's release through May 2024. Words preferentially generated by ChatGPT, including 'delve', 'comprehend', 'boast', 'swift' and 'meticulous', increased abruptly in spontaneous human speech after the release, with top GPT words growing at roughly 25% to 50% per year. 'Delve' rose significantly in academic YouTube talks (p=0.010) and in science and technology, business and education podcasts.

The team also ran a preregistered experiment with 496 participants: a brief interaction with a chatbot led participants to adopt its words as their own, and the effect persisted past a distractor task. Co-author Levin Brinkmann put it plainly: 'The patterns that are stored in AI technology seem to be transmitting back to the human mind.'

This is the quiet killer of word-level tells. A marker works because of a frequency gap between model output and human baseline; as humans absorb the vocabulary, the gap closes from both ends. Worse, it feeds back: the FSU researchers note that AI-influenced human writing re-enters the training data of future models, reinforcing the very patterns being measured. Every year a word spends on a viral 'AI words' list, it becomes a weaker marker, partly because humans read the list too.

What this list cannot tell you

Using these words does not make a text AI-written. Humans wrote 'delve' for centuries before November 2022, and academic English used every word in the table at some baseline long before LLMs existed. The corpus studies measure frequency shifts across millions of documents; they license conclusions about corpora, not verdicts on individual texts. Kobak and colleagues' estimate that at least 13.5% of 2024 biomedical abstracts were LLM-processed is a corpus-level lower bound, not a test you can run on one abstract.

The lists are also literal. Wikipedia's catalogue stresses that a word being overused by AI does not imply its synonyms are overused; the evidence attaches to specific tokens in specific corpora, and base rates differ across fields, registers and dialects.

Finally, the traffic runs both ways in practice: deleting listed words from a text does not remove the other statistical regularities these studies measure, and sprinkling them in does not make human text machine-made. This page exists to document what models actually do and how fast that changes, not to referee authorship disputes. If you want to see how any text, yours or a model's, actually distributes its vocabulary, that is a measurement job, and measurement is what ScriptGrain does.

Methodology

Compiled from peer-reviewed corpus studies (Juzek and Ward, COLING 2025; Kobak et al, Science Advances 2025), a large-scale speech corpus analysis (Yakura et al, Max Planck Institute for Human Development), and the editor-maintained catalogue at Wikipedia's 'Signs of AI writing' project page. Every magnitude on this page traces to one of the listed sources; nothing is reproduced from unsourced listicles.

Evidence standard: corpus studies outrank the speech corpus, which outranks the community catalogue; catalogue-only entries carry no magnitude and are labelled as such. Percentage increases from Juzek and Ward compare 2020 with 2024 word frequencies in PubMed abstracts and apply to that corpus only.

Last checked 2026-08-03. Update policy: the table is re-checked when a major model generation ships, because overuse profiles change between generations. Entries whose evidence no longer holds are re-dated and kept as documented decay, not silently deleted.

Sources

ScriptGrain · Why ScriptGrain · Sample voice profiles · The Journal · Support