# Why Does ChatGPT Love the Em Dash? · ScriptGrain

> How one punctuation mark became the most famous AI tell, why the habit probably formed, and why it is weaker evidence than everyone thinks.

Canonical: https://scriptgrain.com/blog/why-does-chatgpt-use-em-dashes

# Why Does ChatGPT Love the Em Dash?

*By Jack Stovell · 2026-08-31 · Research*

Here's the thing about that little horizontal line everyone's suddenly obsessed with: it's just punctuation. It joins clauses. It's been doing that job since long before anyone typed a prompt into a chatbot. And yet somehow, in the space of about two years, it became the single most cited piece of evidence that a paragraph was written by a machine. Type "why does AI use so many dashes" into a search bar and watch the autocomplete fill in before you finish. Everyone's noticed. Nobody's fully explained it. So let's try.

I'll say upfront: this article is deliberately, almost comically, dash-free. Not one appears anywhere in this piece, including in the title, which is why we're naming the punctuation mark rather than printing it. That's not a gimmick for its own sake (well, it's a bit of a gimmick). It's also a demonstration. If a site can talk about a mark at length without once reaching for it, that tells you the mark was never load-bearing to begin with. It's a habit, not a necessity.

What's probably going on

Nobody outside the labs knows the exact mechanics of why large language models over-index on this particular connector, and I'm not going to pretend otherwise. But there are two plausible, boring, non-conspiratorial explanations that get thrown around, and both are worth taking seriously as hypotheses rather than facts.

The first is training data register. These models learn from enormous piles of text, and a meaningful slice of professionally edited writing, essays, magazine features, certain strands of literary fiction, leans on that connector as a stylish way to insert a pause or an aside without starting a new sentence. If the model has absorbed a lot of that register, it's going to reproduce the habit at a rate that reflects the source material, not some deliberate design choice.

The second is preference tuning. Once a base model gets fine-tuned on human feedback, raters tend to reward prose that reads smoothly, that flows, that doesn't stop and start awkwardly. A dash is a brilliant tool for stitching two related thoughts into something that feels unbroken. If flow is what's getting upvoted during training, you'd expect the model to learn that this mark is a cheap, reliable way to manufacture that feeling. Neither of these is confirmed. Both are consistent with what we observe. That's a hypothesis, not a study, and I'm not going to dress it up as more than that.

Why it's a weak tell anyway

Here's where it gets genuinely funny. The mark that everyone treats as a smoking gun is also one of the most celebrated tools in the English stylist's kit. Writers have used it for centuries to add a jolt of pace, to interrupt themselves, to let a sentence swerve somewhere unexpected. Plenty of very human, very distinctive prose is dense with it. So the "tell" isn't really a tell at all; it's a punctuation mark that both machines and skilled humans happen to like, for overlapping but not identical reasons.

And tells decay. Fast. Whatever quirk is fashionable to point at this month gets patched, tuned, or trained around within a couple of model generations, because labs read the same commentary you do. The dash obsession might already be softening in newer releases; by the time this article is a year old, it could read as a dated observation about an early, more mannered phase of these tools. That's the trouble with treating any single surface habit as proof of anything. It's a snapshot, not a law.

Which is exactly why we don't build our own approach around spotting one flashy mark. At ScriptGrain, we measure 45 separate attributes of how a piece of writing actually behaves: sentence length variance, contraction rate, how someone opens paragraphs, how they hedge or don't, where their rhythm gets choppy and where it settles. One punctuation habit is a single data point in that picture, not the whole picture. We've got a fuller catalogue of these surface markers and what they do and don't prove on our reference page, if you want the longer, more careful version of this argument.

To be fair, none of this is an argument for hiding anything. We're not in the business of telling you how to dodge a detector, and we won't pretend that's a legitimate goal. The actual point is authenticity: understanding what your own writing measurably does, so that whatever you publish, assisted or not, still sounds like you rather than like the median output of a very large, very well-read averaging machine.

If you're curious what your own writing looks like under that kind of measurement, there's a free browser tool at scriptgrain.com/tools/writing-style-analysis, no account needed. Build a full voice profile and run one analysis free, no card required. Plans start at £12 a month after that.

At the end of the day, the mark isn't the villain. It's just punctuation that got famous for the wrong reasons.

[More from the ScriptGrain Journal](https://scriptgrain.com/blog)
