What does normal writing look like? Three centuries of human prose, measured

SGR-006, conducted 2026-10-04. Analyst: Jack Stovell, founder, ScriptGrain.

Question: For each habit of prose that can be counted, what is the normal range in human writing, by genre, and how has it moved from the 1700s to the web of 2021, before ChatGPT?

Summary

Background. Claims about writing (sentences are getting shorter, nobody uses semicolons any more, informal writing is new) are common and rarely measured across genres and centuries with one instrument. And any measure of a single piece of writing needs a reference point: long sentences compared with what?

What we did. We fixed three hypotheses, the genres and the sampling before measuring anything. We drew 30 public-domain books for each genre and era from Project Gutenberg (ten genres; the 1700s, the 1800s and the early 1900s, dated by when the author turned 40), three passages from each, and added 2,200 pages from Common Crawl's 2021 web and 1,000 arXiv abstracts from before ChatGPT. Every text was measured with the same deterministic code ScriptGrain uses, for sixteen habits.

What we found. Non-fiction sentences shortened steadily: 32.2 words on average in the 1700s, 28.6 in the 1800s, 24.5 in the early 1900s and 15.1 on the 2021 web. Semicolons fell from 12.1 to 1.5 per 1,000 words. Contractions, though, did not drift up in books at all (4.4 to 4.1 per 1,000 words); they more than tripled only on the web. And fiction was not always the short-sentence genre: in the 1700s, novels ran longer than history.

What it means. Some of the best-known changes in English prose are real and old (shorter sentences, fewer semicolons); others arrived suddenly with the web (contractions, 'you'). The percentile tables give every measure a reference range by genre and era, so a figure from any text can be read against what was normal for its kind.

Key numbers

Sample

Books: Project Gutenberg English texts in ten genres (fiction, essays, history, science, philosophy and religion, travel, biography, letters, children's books, drama; poetry excluded), 30 per genre and era where at least 30 exist, in a seeded random order, with three passages of about 1,000 words at 20%, 50% and 80% of each book. Three cells (1700s essays, children's books and drama) had too few books and were dropped, not filled. Web: 100 pages from each of 22 sources in nine kinds of writing, from Common Crawl's October 2021 crawl, in SGR-007's frozen order. Abstracts: 1,000 arXiv abstracts from 2015 to November 2022 with no later revision, from SGR-008's census.

Method

Pre-registration: hypotheses, genres, eras and sampling were committed to the public repository before any text was measured (docs/research/sgr-006-preregistration.md).

Measurement: every text scored with the same deterministic code ScriptGrain uses, for sixteen habits, on all text. No language model reads or scores any text.

Re-run: the first books harvest was found to contain rows written by two runs at once, with some books mislabelled. It was discarded before analysis and the books were harvested again by a single run; the harvester now refuses to write into an existing file.

Check added after the main run (not pre-registered): web pages and the non-fiction book passages re-measured on full prose lines only (8 words or more, ending in sentence punctuation), to see whether the web's shorter sentences come from headings and lists.

Variables measured

What each measure means

Era of a book
The catalogue gives authors' birth years, not books' dates, so a book is placed in the century in which its author turned 40. 'Early 1900s' ends around 1930, where public-domain texts stop.
Non-fiction (for the hypotheses)
Books: history, science, philosophy and religion, travel and biography, pooled. 2021 web: news (UK and US), Wikipedia and how-to pages, pooled.
Words per sentence
The product's sentence measure on all of a text. On web pages it also counts headings, captions, list items and reference entries as short sentences, which books do not have; see Checks.
Dashes
Em dashes, and the double hyphens plain-text editions use for them, per 1,000 words. En dashes are not counted, and British news often uses spaced en dashes, so web figures for dashes run low.
Percentiles
For each genre and era, the 10th, 25th, 50th (median), 75th and 90th percentile of each habit across its texts. The median is the typical text; the middle half lies between the 25th and 75th.
95% interval
From 10,000 bootstrap resamples of texts. A hypothesis holds only if every one of its comparisons excludes zero in the predicted direction.

Findings

What this study does not show

Sentences through three centuries

Line chart: non-fiction sentence length falls from 32.2 words in the 1700s to 15.1 on the 2021 web; fiction falls faster, from 36.6 to 17.6.
Source: ScriptGrain SGR-006, CC BY 4.0. The 2021 web figure includes headings and lists counted as sentences; see Checks.

Punctuation through three centuries

Three small line charts: semicolons fall steadily, dashes rise to the early 1900s then fall on the web, contractions stay flat in books then rise sharply on the web.
Source: ScriptGrain SGR-006, CC BY 4.0.

The baseline: normal sentence length by genre and era

Range chart of sentence length for every genre and era: 1700s books highest, drama and children's books lowest among books, 2021 web pages lower still, arXiv abstracts near early-1900s books.
Source: ScriptGrain SGR-006, CC BY 4.0. Full percentiles for all sixteen habits are in the data.

Pre-registered hypotheses and results

HypothesisResultTest
H1: in every era, fiction's sentences are shorter than history, science and philosophy and religionNot supported1700s vs history: +5.8 (+1.8 to +9.8); 1700s vs science: +2.2 (−2.0 to +6.4); 1700s vs philosophy and religion: +2.5 (−1.9 to +6.8); 1800s vs history: −8.3 (−10.7 to −5.8); 1800s vs science: −7.6 (−10.0 to −5.1); 1800s vs philosophy and religion: −6.8 (−9.5 to −4.1); early 1900s vs history: −8.7 (−10.8 to −6.6); early 1900s vs science: −4.0 (−5.7 to −2.2); early 1900s vs philosophy and religion: −7.9 (−9.7 to −6.1)
H2: non-fiction sentences shorter in the early 1900s than the 1700s, and shorter again on the 2021 webSupported1700s to early 1900s: −7.7 (−9.0 to −6.4); early 1900s to 2021 web: −9.4 (−10.1 to −8.7)
H3: contractions higher in the early 1900s than the 1700s, and higher again on the 2021 webNot supported1700s to early 1900s: −0.3 (−1.0 to +0.4); early 1900s to 2021 web: +9.7 (+9.0 to +10.5)

The baseline: typical values by genre and era

Medians (the typical text) for six habits. The 10th, 25th, 75th and 90th percentiles for all sixteen habits are in the summary file.

GenreEraTextsWords per sentenceContractions per 1,000Semicolons per 1,000Dashes per 1,000Long-word share'You' share of pronouns
Biography1700s9027.33.610.50.60.130.03
Drama1700s9016.224.012.410.30.090.24
Fiction1700s9031.23.913.71.20.120.14
History1700s9027.93.29.60.70.130.04
Letters1700s9026.66.711.42.80.120.23
Philosophy and religion1700s9032.21.511.80.00.130.04
Science1700s9033.31.812.90.00.140.00
Travel1700s9031.82.810.80.70.130.00
Biography1800s9025.73.77.63.10.120.06
Children's books1800s9020.412.36.73.40.080.32
Drama1800s9015.114.010.35.90.100.26
Essays1800s9028.53.67.83.60.150.13
Fiction1800s9019.29.87.24.90.110.32
History1800s9029.33.76.22.90.140.00
Letters1800s9024.85.75.83.80.120.22
Philosophy and religion1800s9026.33.66.92.10.130.00
Science1800s9027.81.06.82.70.160.00
Travel1800s9027.93.07.62.90.130.00
Biographyearly 1900s9023.84.73.01.90.130.04
Children's booksearly 1900s9015.520.21.02.00.080.32
Dramaearly 1900s9010.722.02.47.00.080.38
Essaysearly 1900s9022.25.45.53.90.110.07
Fictionearly 1900s9016.215.64.04.90.100.31
Historyearly 1900s9025.71.93.41.00.150.00
Lettersearly 1900s9024.35.83.01.90.140.20
Philosophy and religionearly 1900s9025.63.15.91.00.150.07
Scienceearly 1900s9021.71.02.91.70.150.00
Travelearly 1900s9024.13.53.62.60.130.02
Essay-mill essays2021 web30015.63.50.90.00.240.53
How-to content2021 web20012.29.30.00.00.150.83
Marketing and SEO blogs2021 web30015.219.70.00.50.160.86
Medium posts2021 web10014.410.90.02.40.160.35
News (UK and US)2021 web50017.216.00.00.00.160.11
Press releases2021 web30019.57.70.02.10.270.00
Q&A forum2021 web10011.112.71.00.00.150.39
Self-help and productivity2021 web30014.220.11.41.40.130.49
Wikipedia2021 web10010.33.51.90.00.220.07
Scientific abstracts (arXiv)2015-20221,00024.10.00.00.00.280.00

Checks

Full prose lines only (added after the main run): web pages and book passages re-measured on lines of eight or more words ending in sentence punctuation. News: 17.2 words per sentence on all text, 19.7 on prose lines (394 pages). How-to pages: 13.7 and 17.7 (98). Early-1900s non-fiction books: 24.6 and 25.0. Wikipedia pages could not be matched back to the crawl index for this check, so its web figures should be read as including headings and lists.

The first books harvest was discarded after two runs were found to have written into the same file, mislabelling some books' genre and era. The re-run is the data published here; the harvester now refuses to append to an existing file.

Limitations

Competing interests

ScriptGrain sells writing-voice measurement and generation, and the analyst builds it. The hypotheses, genres and sampling were fixed and published before anything was measured, nothing was re-run or selected (except one disclosed re-run after a file fault, below), and the scripts are public, so the numbers can be checked without taking our word for them.

Reproducing this

Data

Released under CC BY 4.0: free to reuse, including commercially, with credit to ScriptGrain and a link to this page.

Citations

How to cite

ScriptGrain (2026). What does normal writing look like? Three centuries of human prose, measured (Study SGR-006, conducted 4 October 2026). Dataset licensed CC BY 4.0. https://scriptgrain.com/research/what-does-normal-writing-look-like

About the analyst

Jack Stovell has worked in finance and data for more than twelve years, building management reporting, forecasting and profitability models for advertising agencies and tech scale-ups, from SQL and Power BI reporting to board-level analysis. Since 2016 he has run Adapt Progress Evolve, an applied AI studio, where he builds and operates AI systems and data products: ScriptGrain's measurement of writing voice across 45 attributes, UK Spend, which brings 16.7 million rows of UK council spending into one queryable dataset, and more than thirty AI agents running in production. He designs ScriptGrain's studies and is accountable for every number in them.