What does normal writing look like? Three centuries of human prose, measured
SGR-006, conducted 2026-10-04. Analyst: Jack Stovell, founder, ScriptGrain.
Question: For each habit of prose that can be counted, what is the normal range in human writing, by genre, and how has it moved from the 1700s to the web of 2021, before ChatGPT?
Summary
Background. Claims about writing (sentences are getting shorter, nobody uses semicolons any more, informal writing is new) are common and rarely measured across genres and centuries with one instrument. And any measure of a single piece of writing needs a reference point: long sentences compared with what?
What we did. We fixed three hypotheses, the genres and the sampling before measuring anything. We drew 30 public-domain books for each genre and era from Project Gutenberg (ten genres; the 1700s, the 1800s and the early 1900s, dated by when the author turned 40), three passages from each, and added 2,200 pages from Common Crawl's 2021 web and 1,000 arXiv abstracts from before ChatGPT. Every text was measured with the same deterministic code ScriptGrain uses, for sixteen habits.
What we found. Non-fiction sentences shortened steadily: 32.2 words on average in the 1700s, 28.6 in the 1800s, 24.5 in the early 1900s and 15.1 on the 2021 web. Semicolons fell from 12.1 to 1.5 per 1,000 words. Contractions, though, did not drift up in books at all (4.4 to 4.1 per 1,000 words); they more than tripled only on the web. And fiction was not always the short-sentence genre: in the 1700s, novels ran longer than history.
What it means. Some of the best-known changes in English prose are real and old (shorter sentences, fewer semicolons); others arrived suddenly with the web (contractions, 'you'). The percentile tables give every measure a reference range by genre and era, so a figure from any text can be read against what was normal for its kind.
Key numbers
- Average sentence length in English non-fiction fell from 32.2 words in the 1700s to 28.6 in the 1800s and 24.5 in the early 1900s (ScriptGrain, 2026; 450 non-fiction books).
- On the web in 2021, news, Wikipedia and how-to pages averaged 15.1 words per sentence; counting only full prose lines, news and how-to pages averaged 19.3, still well below the 25.0 of early-1900s non-fiction books.
- Semicolons in English non-fiction fell from 12.1 per 1,000 words in the 1700s to 5.9 in the early 1900s and 1.5 on the 2021 web.
- Contractions such as "don't" barely changed in English non-fiction books over two centuries (4.4 per 1,000 words in the 1700s, 4.1 in the early 1900s), then more than tripled on the 2021 web (13.9).
- In the 1700s, novels had longer sentences than history books (36.6 words against 30.9); by the early 1900s fiction was the shorter by 8.7 words.
- Dashes in English non-fiction peaked in the early 1900s at 6.6 per 1,000 words, against 1.3 on the 2021 web.
- The share of 'you' among I, we and you rose from 11% in 1700s non-fiction to 37% on the 2021 web.
Sample
Books: Project Gutenberg English texts in ten genres (fiction, essays, history, science, philosophy and religion, travel, biography, letters, children's books, drama; poetry excluded), 30 per genre and era where at least 30 exist, in a seeded random order, with three passages of about 1,000 words at 20%, 50% and 80% of each book. Three cells (1700s essays, children's books and drama) had too few books and were dropped, not filled. Web: 100 pages from each of 22 sources in nine kinds of writing, from Common Crawl's October 2021 crawl, in SGR-007's frozen order. Abstracts: 1,000 arXiv abstracts from 2015 to November 2022 with no later revision, from SGR-008's census.
- English; poetry excluded
- Books of at least 5,000 words; web pages of at least 200 words of article text
- Written or crawled before ChatGPT's launch (30 November 2022)
Method
Pre-registration: hypotheses, genres, eras and sampling were committed to the public repository before any text was measured (docs/research/sgr-006-preregistration.md).
Measurement: every text scored with the same deterministic code ScriptGrain uses, for sixteen habits, on all text. No language model reads or scores any text.
Re-run: the first books harvest was found to contain rows written by two runs at once, with some books mislabelled. It was discarded before analysis and the books were harvested again by a single run; the harvester now refuses to write into an existing file.
Check added after the main run (not pre-registered): web pages and the non-fiction book passages re-measured on full prose lines only (8 words or more, ending in sentence punctuation), to see whether the web's shorter sentences come from headings and lists.
Variables measured
- Words per sentence, sentence length variation, share of short sentences
- Commas per sentence; semicolons, exclamation marks, question marks and dashes per 1,000 words
- Contractions per 1,000 words; parentheticals
- Short-word and long-word shares
- I, we and you shares of pronouns
- AI-tell density
What each measure means
- Era of a book
- The catalogue gives authors' birth years, not books' dates, so a book is placed in the century in which its author turned 40. 'Early 1900s' ends around 1930, where public-domain texts stop.
- Non-fiction (for the hypotheses)
- Books: history, science, philosophy and religion, travel and biography, pooled. 2021 web: news (UK and US), Wikipedia and how-to pages, pooled.
- Words per sentence
- The product's sentence measure on all of a text. On web pages it also counts headings, captions, list items and reference entries as short sentences, which books do not have; see Checks.
- Dashes
- Em dashes, and the double hyphens plain-text editions use for them, per 1,000 words. En dashes are not counted, and British news often uses spaced en dashes, so web figures for dashes run low.
- Percentiles
- For each genre and era, the 10th, 25th, 50th (median), 75th and 90th percentile of each habit across its texts. The median is the typical text; the middle half lies between the 25th and 75th.
- 95% interval
- From 10,000 bootstrap resamples of texts. A hypothesis holds only if every one of its comparisons excludes zero in the predicted direction.
Findings
- H2 (supported): non-fiction sentences got shorter, century after century (32.2 → 28.6 → 24.5 → 15.1 words per sentence): Early-1900s non-fiction books averaged 7.7 fewer words per sentence than 1700s books (95% interval 6.4 to 9.0 fewer), and 2021 web non-fiction 9.4 fewer again. Part of the web figure comes from headings and lists counted as sentences; on full prose lines only, news and how-to pages averaged 19.3 words against 25.0 for early-1900s books, a smaller gap in the same direction.
- H3 (not supported): contractions did not creep up in books; they jumped on the web (4.4 → 4.1 in books, 13.9 on the 2021 web, per 1,000 words): We predicted a rise from the 1700s to the early 1900s. Non-fiction books were flat (difference −0.3, interval −1.0 to +0.4). The jump came after: 2021 web non-fiction used 9.7 more per 1,000 words than early-1900s books, and still 15.4 per 1,000 on full prose lines only.
- H1 (not supported): fiction was not always the short-sentence genre (1700s fiction 36.6 words per sentence; 1800s 21.3; early 1900s 17.6): In the 1800s and early 1900s, fiction's sentences were shorter than history, science and philosophy, as predicted. In the 1700s they were not: novels ran 5.8 words longer than history, and the comparisons with science and philosophy and religion were inconclusive. Fiction's sentences shortened faster than any non-fiction genre's, by about 19 words between the 1700s and the early 1900s.
- Exploratory: semicolons faded and dashes rose, then fell (semicolons 12.1 → 1.5; dashes peaked at 6.6 in the early 1900s): In non-fiction, semicolons fell in every era, from 12.1 per 1,000 words in the 1700s to 5.9 in the early 1900s and 1.5 on the 2021 web. Dashes rose from 2.7 to 6.6 and were rare on the 2021 web (1.3): books of a century ago used them about five times as often as web pages did just before ChatGPT. Not pre-registered; read as description.
What this study does not show
- What any individual writer of an era wrote like. The baseline describes genres and eras; single books vary widely, as the percentile ranges show.
- Why prose changed. Printing, schooling, readership and the web all plausibly played a part; the study measures the change, not its causes.
- That voice measurement, ScriptGrain's or anyone's, can tell whether a text was written by AI. Nothing in this study tests that.
- Anything after 2022. The modern samples stop before ChatGPT on purpose, so the baseline is human writing.
Sentences through three centuries
Punctuation through three centuries
The baseline: normal sentence length by genre and era
Pre-registered hypotheses and results
| Hypothesis | Result | Test |
|---|---|---|
| H1: in every era, fiction's sentences are shorter than history, science and philosophy and religion | Not supported | 1700s vs history: +5.8 (+1.8 to +9.8); 1700s vs science: +2.2 (−2.0 to +6.4); 1700s vs philosophy and religion: +2.5 (−1.9 to +6.8); 1800s vs history: −8.3 (−10.7 to −5.8); 1800s vs science: −7.6 (−10.0 to −5.1); 1800s vs philosophy and religion: −6.8 (−9.5 to −4.1); early 1900s vs history: −8.7 (−10.8 to −6.6); early 1900s vs science: −4.0 (−5.7 to −2.2); early 1900s vs philosophy and religion: −7.9 (−9.7 to −6.1) |
| H2: non-fiction sentences shorter in the early 1900s than the 1700s, and shorter again on the 2021 web | Supported | 1700s to early 1900s: −7.7 (−9.0 to −6.4); early 1900s to 2021 web: −9.4 (−10.1 to −8.7) |
| H3: contractions higher in the early 1900s than the 1700s, and higher again on the 2021 web | Not supported | 1700s to early 1900s: −0.3 (−1.0 to +0.4); early 1900s to 2021 web: +9.7 (+9.0 to +10.5) |
The baseline: typical values by genre and era
Medians (the typical text) for six habits. The 10th, 25th, 75th and 90th percentiles for all sixteen habits are in the summary file.
| Genre | Era | Texts | Words per sentence | Contractions per 1,000 | Semicolons per 1,000 | Dashes per 1,000 | Long-word share | 'You' share of pronouns |
|---|---|---|---|---|---|---|---|---|
| Biography | 1700s | 90 | 27.3 | 3.6 | 10.5 | 0.6 | 0.13 | 0.03 |
| Drama | 1700s | 90 | 16.2 | 24.0 | 12.4 | 10.3 | 0.09 | 0.24 |
| Fiction | 1700s | 90 | 31.2 | 3.9 | 13.7 | 1.2 | 0.12 | 0.14 |
| History | 1700s | 90 | 27.9 | 3.2 | 9.6 | 0.7 | 0.13 | 0.04 |
| Letters | 1700s | 90 | 26.6 | 6.7 | 11.4 | 2.8 | 0.12 | 0.23 |
| Philosophy and religion | 1700s | 90 | 32.2 | 1.5 | 11.8 | 0.0 | 0.13 | 0.04 |
| Science | 1700s | 90 | 33.3 | 1.8 | 12.9 | 0.0 | 0.14 | 0.00 |
| Travel | 1700s | 90 | 31.8 | 2.8 | 10.8 | 0.7 | 0.13 | 0.00 |
| Biography | 1800s | 90 | 25.7 | 3.7 | 7.6 | 3.1 | 0.12 | 0.06 |
| Children's books | 1800s | 90 | 20.4 | 12.3 | 6.7 | 3.4 | 0.08 | 0.32 |
| Drama | 1800s | 90 | 15.1 | 14.0 | 10.3 | 5.9 | 0.10 | 0.26 |
| Essays | 1800s | 90 | 28.5 | 3.6 | 7.8 | 3.6 | 0.15 | 0.13 |
| Fiction | 1800s | 90 | 19.2 | 9.8 | 7.2 | 4.9 | 0.11 | 0.32 |
| History | 1800s | 90 | 29.3 | 3.7 | 6.2 | 2.9 | 0.14 | 0.00 |
| Letters | 1800s | 90 | 24.8 | 5.7 | 5.8 | 3.8 | 0.12 | 0.22 |
| Philosophy and religion | 1800s | 90 | 26.3 | 3.6 | 6.9 | 2.1 | 0.13 | 0.00 |
| Science | 1800s | 90 | 27.8 | 1.0 | 6.8 | 2.7 | 0.16 | 0.00 |
| Travel | 1800s | 90 | 27.9 | 3.0 | 7.6 | 2.9 | 0.13 | 0.00 |
| Biography | early 1900s | 90 | 23.8 | 4.7 | 3.0 | 1.9 | 0.13 | 0.04 |
| Children's books | early 1900s | 90 | 15.5 | 20.2 | 1.0 | 2.0 | 0.08 | 0.32 |
| Drama | early 1900s | 90 | 10.7 | 22.0 | 2.4 | 7.0 | 0.08 | 0.38 |
| Essays | early 1900s | 90 | 22.2 | 5.4 | 5.5 | 3.9 | 0.11 | 0.07 |
| Fiction | early 1900s | 90 | 16.2 | 15.6 | 4.0 | 4.9 | 0.10 | 0.31 |
| History | early 1900s | 90 | 25.7 | 1.9 | 3.4 | 1.0 | 0.15 | 0.00 |
| Letters | early 1900s | 90 | 24.3 | 5.8 | 3.0 | 1.9 | 0.14 | 0.20 |
| Philosophy and religion | early 1900s | 90 | 25.6 | 3.1 | 5.9 | 1.0 | 0.15 | 0.07 |
| Science | early 1900s | 90 | 21.7 | 1.0 | 2.9 | 1.7 | 0.15 | 0.00 |
| Travel | early 1900s | 90 | 24.1 | 3.5 | 3.6 | 2.6 | 0.13 | 0.02 |
| Essay-mill essays | 2021 web | 300 | 15.6 | 3.5 | 0.9 | 0.0 | 0.24 | 0.53 |
| How-to content | 2021 web | 200 | 12.2 | 9.3 | 0.0 | 0.0 | 0.15 | 0.83 |
| Marketing and SEO blogs | 2021 web | 300 | 15.2 | 19.7 | 0.0 | 0.5 | 0.16 | 0.86 |
| Medium posts | 2021 web | 100 | 14.4 | 10.9 | 0.0 | 2.4 | 0.16 | 0.35 |
| News (UK and US) | 2021 web | 500 | 17.2 | 16.0 | 0.0 | 0.0 | 0.16 | 0.11 |
| Press releases | 2021 web | 300 | 19.5 | 7.7 | 0.0 | 2.1 | 0.27 | 0.00 |
| Q&A forum | 2021 web | 100 | 11.1 | 12.7 | 1.0 | 0.0 | 0.15 | 0.39 |
| Self-help and productivity | 2021 web | 300 | 14.2 | 20.1 | 1.4 | 1.4 | 0.13 | 0.49 |
| Wikipedia | 2021 web | 100 | 10.3 | 3.5 | 1.9 | 0.0 | 0.22 | 0.07 |
| Scientific abstracts (arXiv) | 2015-2022 | 1,000 | 24.1 | 0.0 | 0.0 | 0.0 | 0.28 | 0.00 |
Checks
Full prose lines only (added after the main run): web pages and book passages re-measured on lines of eight or more words ending in sentence punctuation. News: 17.2 words per sentence on all text, 19.7 on prose lines (394 pages). How-to pages: 13.7 and 17.7 (98). Early-1900s non-fiction books: 24.6 and 25.0. Wikipedia pages could not be matched back to the crawl index for this check, so its web figures should be read as including headings and lists.
The first books harvest was discarded after two runs were found to have written into the same file, mislabelling some books' genre and era. The re-run is the data published here; the harvester now refuses to append to an existing file.
Limitations
- Books are dated by when their author turned 40, not by publication; a book written young or old can sit in a neighbouring era.
- Project Gutenberg texts come from particular editions, and punctuation can be an editor's or printer's; semicolons and dashes are the most exposed.
- The web's sentence measure counts headings, captions, list items and references as short sentences, which pulls its figures down, most for Wikipedia (see Checks).
- Genres are assigned from catalogue subjects, so a book can sit in more than one (travel letters, for example) and some labels are broad.
- English only, and mostly British and American writing.
Competing interests
ScriptGrain sells writing-voice measurement and generation, and the analyst builds it. The hypotheses, genres and sampling were fixed and published before anything was measured, nothing was re-run or selected (except one disclosed re-run after a file fault, below), and the scripts are public, so the numbers can be checked without taking our word for them.
Reproducing this
- Pre-registration, sources, harvester, analysis and checks: docs/research/sgr-006-preregistration.md and scripts/experiments/sgr-006/ in the ScriptGrain repository.
- Every text is identified in the data (Project Gutenberg number and passage, web page URL and capture time, arXiv id), so any can be re-read and measured.
Data
- Every text measured, sixteen habits (CSV) (csv)
- Baseline: percentiles by genre and era, hypotheses, trends (JSON) (json)
- Prose-line check (JSON) (json)
Released under CC BY 4.0: free to reuse, including commercially, with credit to ScriptGrain and a link to this page.
Citations
How to cite
ScriptGrain (2026). What does normal writing look like? Three centuries of human prose, measured (Study SGR-006, conducted 4 October 2026). Dataset licensed CC BY 4.0. https://scriptgrain.com/research/what-does-normal-writing-look-like
About the analyst
Jack Stovell has worked in finance and data for more than twelve years, building management reporting, forecasting and profitability models for advertising agencies and tech scale-ups, from SQL and Power BI reporting to board-level analysis. Since 2016 he has run Adapt Progress Evolve, an applied AI studio, where he builds and operates AI systems and data products: ScriptGrain's measurement of writing voice across 45 attributes, UK Spend, which brings 16.7 million rows of UK council spending into one queryable dataset, and more than thirty AI agents running in production. He designs ScriptGrain's studies and is accountable for every number in them.