# Wikipedia's Signs of AI Writing: What the List Says and Where It Fails · ScriptGrain

> Wikipedia's Signs of AI writing is an editors' advice page that lists chatbot habits. It is useful for spotting patterns but proves nothing about a single text, and the page says so itself.

Canonical: https://scriptgrain.com/blog/wikipedia-signs-of-ai-writing

# Wikipedia's Signs of AI Writing: What the List Says and Where It Fails

*By Jack Stovell [X](https://x.com/StovBuilds) · [Instagram](https://www.instagram.com/stovbuilds) · [LinkedIn](https://www.linkedin.com/in/jackstovell/) · 2026-10-10 · Guides*

Wikipedia's "Signs of AI writing" is an advice page that volunteer editors wrote to help other editors spot undisclosed chatbot text. It lists habits such as title case headings, heavy boldface, the rule of three and promotional tone. It is useful evidence of patterns and weak evidence about any single text, and the page says so itself.

## What Wikipedia's Signs of AI writing page is

The page sits in Wikipedia's project space and belongs to WikiProject AI Cleanup, the group of editors who review suspected chatbot edits. A notice at the top says it "is not an encyclopedia article or a Wikipedia policy". It opens with a plain claim, "LLMs tend to have an identifiable writing style", and then calls itself a set of observations rather than rules, built from real examples found on Wikipedia.

So it is a field guide for editors. It gives no score and makes no ruling. Its job is to show a reviewer where to look when a submission reads oddly. It also warns editors against simply fixing the surface signs, because that makes the text harder to detect.

It is also a live document. Anyone can edit it, and it changes often. Everything below describes [the page](https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing) as it stood when we read it on 9 October 2026. Check it again on the day if you need to quote it.

## The signs, grouped by type

The page uses its own headings. The main groups are:

- **Content.** Undue emphasis on significance and legacy, superficial analysis, vague attribution to unnamed experts, outline-style conclusions about challenges and future prospects, and advert-like language. The page puts it bluntly: "LLMs have serious problems keeping a neutral tone."
- **Language and grammar.** A high density of "AI vocabulary", avoiding plain "is" and "are", the stock contrast shapes it calls negative parallelisms, and the rule of three.
- **Style.** Title case in headings, too much boldface, vertical lists where each item opens with a bold inline header and a colon, emoji used as formatting, odd tables, heading errors, and curly quotation marks.
- **Communication intended for the user.** Leftovers from a chat, such as collaborative asides, placeholder text, and disclaimers about how sources should be used.
- **Markup and citations.** Markdown pasted into wikitext, chatbot citation artefacts, broken links, invalid DOIs and ISBNs, and book citations with no page numbers.
- **Edit summaries and comments.** Signs specific to talk-page comments and edit summaries.

Two sections near the end matter as much as the list. "Ineffective indicators" names things that do not reliably point to AI: perfect grammar, "bland" or "robotic" prose, formal or academic prose, and transition words in isolation. "Historical indicators" holds signs that were common in older models and are much rarer now. The em dash has moved there, under a heading dated 2022 to September 2026, and the page notes that some AI companies have tried to make newer chatbots use fewer dashes.

For a wider catalogue beyond Wikipedia's, see our [markers of AI writing](https://scriptgrain.com/reference/markers-of-ai-writing).

## Which signs you can count and which need a judgement

Most write-ups present the list as one flat set of checks. It actually splits into two kinds.

Signs you can count:

1. **Punctuation rates.** Dashes, semicolons and curly quotes per 1,000 words.
2. **Sentence-length variance.** Whether sentence lengths stay in a narrow band or range from very short to long.
3. **Vocabulary.** How often a fixed list of favoured words appears, and how varied the wording is overall.
4. **Formatting.** Bold phrases per paragraph, bulleted lists with bold lead-ins, heading depth, tables, emoji.

Each gives you a number you can compare with your own earlier writing.

Signs that need a reader:

- **Editorialising.** Whether a sentence tells the reader what to think depends on the subject and on what a neutral version would say.
- **Promotional tone.** A product page is meant to sell. An encyclopaedia entry is not. The same sentence can be fine in one and a sign in the other.
- **Superficial analysis and vague attribution.** Spotting these means knowing what the sources actually support.

Software can find candidate sentences for these. Deciding whether a paragraph is quietly selling something still takes a person who knows what the text is for.

## Why no single sign proves anything

Take any sign and ask who else produces it. People use the rule of three because it is an old rhetorical device, and professional writers have always used em dashes. Copywriters write promotional prose for a living. On curly quotes the page is direct: "Curly quotes alone do not prove LLM use", because published works use them all the time. It adds that ChatGPT and DeepSeek typically produce curly quotes while Gemini and Claude typically do not, so even this sign depends on the model.

Wikipedia states the caveat itself. The page says "Not all text featuring these indicators is AI-generated" and "Many elements of AI writing can be found in editorials, blogs, or fan fiction." It calls the signs "only potential signs of a problem, not the problem itself". Of em dashes, it says the sign "is most useful when taken in combination with other indicators, not by itself." Editors on the [talk page](https://en.wikipedia.org/wiki/Wikipedia_talk:Signs_of_AI_writing) make the same point when they weigh up new signs, describing one candidate as "not really meaningful on its own".

It also warns readers about their own judgement: "Humans are notoriously bad at distinguishing human and LLM-generated text." And it tells editors not to rely only on detection tools, which perform better than chance but have "non-trivial error rates". Our [AI detector accuracy](https://scriptgrain.com/reference/ai-detector-accuracy) reference sets out what published tests found.

This is how false positives happen. A list of habits will always catch people who simply have those habits. The page notes, for example, that editors who are not native English speakers may avoid repeating words, which looks like one of the listed tells. For more on that risk, read [AI detectors and non-native English writers](https://scriptgrain.com/blog/ai-detectors-and-non-native-english-writers).

The second problem is that the tells move. The page says "what is typical for GPT-5 is not necessarily characteristic of GPT-4 or Gemini", which is why it keeps a historical section. Our study of 2.1 million arXiv abstracts, [the AI tells moved](https://scriptgrain.com/research/the-ai-tells-moved), found the same thing at scale. The words tied to 2024 models peaked and then fell, a different set rose, and lists built on 2024 output now miss most of the signal. Any list describes the period it was written in.

## How to check your own draft against the list

You can do most of this by hand, for free.

1. **Scan the formatting.** Count bold phrases, headings and bulleted lists with bold lead-ins. If a short piece has more structure than argument, cut some of the structure.
2. **Search for punctuation habits.** Use find in your editor to look for em dashes, semicolons and curly quotes. Note how many there are per page, and ask whether you would have typed that many yourself.
3. **Check the stock phrases.** Search for transitions such as "Additionally", "Notably" and "In conclusion". One is fine. A run of them is worth rewriting.
4. **Read the tone.** Mark every sentence that praises the subject or tells the reader why it matters. Keep the ones you could back with a source.
5. **Compare with something you wrote before.** This step is what gives the others meaning. A dash rate that looks odd in a stranger's text may be normal for you, and you only find out by checking both.

Reading aloud helps with rhythm. If every sentence takes the same breath, vary the lengths.

After a while, eyeballing gives you impressions when what you need is rates. ScriptGrain's free [AI slop checker](https://scriptgrain.com/tools/ai-slop-checker) counts listed phrases, sentence shapes and em dashes per 1,000 words and flags uniform sentence length, with no signup. A clean result means the passage did not trip those counts. It does not prove who wrote it.

## Questions

### Is Wikipedia's list an AI detector?

No. It is an advice page for human editors, and it gives no score. It tells editors not to rely solely on automated detectors either, and it treats its own signs as reasons to look more closely.

### Can I use the list to prove someone used AI?

No. The page says plainly that not all text with these indicators is AI-generated and that many of the patterns appear in human writing. A cluster of signs is a reason to ask questions or check sources. It is not grounds for an accusation.

### Why do em dashes and the rule of three show up so often?

Both are habits of polished, published prose, which chatbots learned from in bulk. The page says LLMs "overuse the rule of three" and were much more inclined to use em dashes than non-professional human writers in the same genre. It now lists the em dash as a historical indicator, because newer chatbots use fewer of them. There is more in [why ChatGPT uses em dashes](https://scriptgrain.com/blog/why-does-chatgpt-use-em-dashes).

### Do these signs change as models change?

Yes. The page keeps a separate section for indicators that older models produced and newer ones rarely do, and it notes that each model and version writes differently. Treat any list as a snapshot of the time it was written.

### What is the difference between spotting AI patterns and matching your own voice?

Spotting patterns asks whether text looks like generic chatbot output. Matching your voice asks whether it sounds like you. A draft can avoid every sign on Wikipedia's list and still sound like nobody in particular. ScriptGrain builds a voice profile of 45 attributes in 8 layers from your own writing, including dash habits, sentence-length variance, comma density and vocabulary, and scores drafts against it. Your writing is measured and never used to train models. The free plan gives you one profile and one full analysis, with no card needed.

## More from the Journal

- [What Is Stylometric Analysis and Why Does It Matter?](https://scriptgrain.com/blog/what-is-stylometric-analysis)
- [What AI Detector Does Turnitin Use? Its Own Model and Score, Explained](https://scriptgrain.com/blog/what-ai-detector-does-turnitin-use)
- [What AI Detector Do Universities Use? Tools, Policies and Limits](https://scriptgrain.com/blog/what-ai-detector-do-universities-use)

[More from the ScriptGrain Journal](https://scriptgrain.com/blog)
