# What does normal writing look like? 5,720 texts, 1700s to 2021, measured · ScriptGrain Research

> A baseline for human writing: 5,720 texts and 6.0 million words, from 840 public-domain books in ten genres across three centuries to 2,200 pages of the 2021 web and 1,000 scientific abstracts, measured with the product's own instruments. Percentile tables for sixteen habits, three pre-registered hypotheses, and a free dataset under CC BY 4.0.

Canonical: https://scriptgrain.com/research/what-does-normal-writing-look-like

# What does normal writing look like? Three centuries of human prose, measured

**SGR-006**, conducted 2026-10-04. Analyst: Jack Stovell, founder, ScriptGrain.

**Question:** For each habit of prose that can be counted, what is the normal range in human writing, by genre, and how has it moved from the 1700s to the web of 2021, before ChatGPT?

## Summary

**Background.** Claims about writing (sentences are getting shorter, nobody uses semicolons any more, informal writing is new) are common and rarely measured across genres and centuries with one instrument. And any measure of a single piece of writing needs a reference point: long sentences compared with what?

**What we did.** We fixed three hypotheses, the genres and the sampling before measuring anything. We drew 30 public-domain books for each genre and era from Project Gutenberg (ten genres; the 1700s, the 1800s and the early 1900s, dated by when the author turned 40), three passages from each, and added 2,200 pages from Common Crawl's 2021 web and 1,000 arXiv abstracts from before ChatGPT. Every text was measured with the same deterministic code ScriptGrain uses, for sixteen habits.

**What we found.** Non-fiction sentences shortened steadily: 32.2 words on average in the 1700s, 28.6 in the 1800s, 24.5 in the early 1900s and 15.1 on the 2021 web. Semicolons fell from 12.1 to 1.5 per 1,000 words. Contractions, though, did not drift up in books at all (4.4 to 4.1 per 1,000 words); they more than tripled only on the web. And fiction was not always the short-sentence genre: in the 1700s, novels ran longer than history.

**What it means.** Some of the best-known changes in English prose are real and old (shorter sentences, fewer semicolons); others arrived suddenly with the web (contractions, 'you'). The percentile tables give every measure a reference range by genre and era, so a figure from any text can be read against what was normal for its kind.

## Key numbers

- Average sentence length in English non-fiction fell from 32.2 words in the 1700s to 28.6 in the 1800s and 24.5 in the early 1900s (ScriptGrain, 2026; 450 non-fiction books).
- On the web in 2021, news, Wikipedia and how-to pages averaged 15.1 words per sentence; counting only full prose lines, news and how-to pages averaged 19.3, still well below the 25.0 of early-1900s non-fiction books.
- Semicolons in English non-fiction fell from 12.1 per 1,000 words in the 1700s to 5.9 in the early 1900s and 1.5 on the 2021 web.
- Contractions such as "don't" barely changed in English non-fiction books over two centuries (4.4 per 1,000 words in the 1700s, 4.1 in the early 1900s), then more than tripled on the 2021 web (13.9).
- In the 1700s, novels had longer sentences than history books (36.6 words against 30.9); by the early 1900s fiction was the shorter by 8.7 words.
- Dashes in English non-fiction peaked in the early 1900s at 6.6 per 1,000 words, against 1.3 on the 2021 web.
- The share of 'you' among I, we and you rose from 11% in 1700s non-fiction to 37% on the 2021 web.

## Sample

Books: Project Gutenberg English texts in ten genres (fiction, essays, history, science, philosophy and religion, travel, biography, letters, children's books, drama; poetry excluded), 30 per genre and era where at least 30 exist, in a seeded random order, with three passages of about 1,000 words at 20%, 50% and 80% of each book. Three cells (1700s essays, children's books and drama) had too few books and were dropped, not filled. Web: 100 pages from each of 22 sources in nine kinds of writing, from Common Crawl's October 2021 crawl, in SGR-007's frozen order. Abstracts: 1,000 arXiv abstracts from 2015 to November 2022 with no later revision, from SGR-008's census.

- English; poetry excluded
- Books of at least 5,000 words; web pages of at least 200 words of article text
- Written or crawled before ChatGPT's launch (30 November 2022)

## Method

Pre-registration: hypotheses, genres, eras and sampling were committed to the public repository before any text was measured (docs/research/sgr-006-preregistration.md).

Measurement: every text scored with the same deterministic code ScriptGrain uses, for sixteen habits, on all text. No language model reads or scores any text.

Re-run: the first books harvest was found to contain rows written by two runs at once, with some books mislabelled. It was discarded before analysis and the books were harvested again by a single run; the harvester now refuses to write into an existing file.

Check added after the main run (not pre-registered): web pages and the non-fiction book passages re-measured on full prose lines only (8 words or more, ending in sentence punctuation), to see whether the web's shorter sentences come from headings and lists.

### Variables measured

- Words per sentence, sentence length variation, share of short sentences
- Commas per sentence; semicolons, exclamation marks, question marks and dashes per 1,000 words
- Contractions per 1,000 words; parentheticals
- Short-word and long-word shares
- I, we and you shares of pronouns
- AI-tell density

## What each measure means

## Findings

- **H2 (supported): non-fiction sentences got shorter, century after century** (32.2 → 28.6 → 24.5 → 15.1 words per sentence): Early-1900s non-fiction books averaged 7.7 fewer words per sentence than 1700s books (95% interval 6.4 to 9.0 fewer), and 2021 web non-fiction 9.4 fewer again. Part of the web figure comes from headings and lists counted as sentences; on full prose lines only, news and how-to pages averaged 19.3 words against 25.0 for early-1900s books, a smaller gap in the same direction.
- **H3 (not supported): contractions did not creep up in books; they jumped on the web** (4.4 → 4.1 in books, 13.9 on the 2021 web, per 1,000 words): We predicted a rise from the 1700s to the early 1900s. Non-fiction books were flat (difference −0.3, interval −1.0 to +0.4). The jump came after: 2021 web non-fiction used 9.7 more per 1,000 words than early-1900s books, and still 15.4 per 1,000 on full prose lines only.
- **H1 (not supported): fiction was not always the short-sentence genre** (1700s fiction 36.6 words per sentence; 1800s 21.3; early 1900s 17.6): In the 1800s and early 1900s, fiction's sentences were shorter than history, science and philosophy, as predicted. In the 1700s they were not: novels ran 5.8 words longer than history, and the comparisons with science and philosophy and religion were inconclusive. Fiction's sentences shortened faster than any non-fiction genre's, by about 19 words between the 1700s and the early 1900s.
- **Exploratory: semicolons faded and dashes rose, then fell** (semicolons 12.1 → 1.5; dashes peaked at 6.6 in the early 1900s): In non-fiction, semicolons fell in every era, from 12.1 per 1,000 words in the 1700s to 5.9 in the early 1900s and 1.5 on the 2021 web. Dashes rose from 2.7 to 6.6 and were rare on the 2021 web (1.3): books of a century ago used them about five times as often as web pages did just before ChatGPT. Not pre-registered; read as description.

## What this study does not show

- What any individual writer of an era wrote like. The baseline describes genres and eras; single books vary widely, as the percentile ranges show.
- Why prose changed. Printing, schooling, readership and the web all plausibly played a part; the study measures the change, not its causes.
- That voice measurement, ScriptGrain's or anyone's, can tell whether a text was written by AI. Nothing in this study tests that.
- Anything after 2022. The modern samples stop before ChatGPT on purpose, so the baseline is human writing.

## Sentences through three centuries

## Punctuation through three centuries

## The baseline: normal sentence length by genre and era

## Pre-registered hypotheses and results

Hypothesis · Result · Test
H1: in every era, fiction's sentences are shorter than history, science and philosophy and religion · Not supported · 1700s vs history: +5.8 (+1.8 to +9.8); 1700s vs science: +2.2 (−2.0 to +6.4); 1700s vs philosophy and religion: +2.5 (−1.9 to +6.8); 1800s vs history: −8.3 (−10.7 to −5.8); 1800s vs science: −7.6 (−10.0 to −5.1); 1800s vs philosophy and religion: −6.8 (−9.5 to −4.1); early 1900s vs history: −8.7 (−10.8 to −6.6); early 1900s vs science: −4.0 (−5.7 to −2.2); early 1900s vs philosophy and religion: −7.9 (−9.7 to −6.1)
H2: non-fiction sentences shorter in the early 1900s than the 1700s, and shorter again on the 2021 web · Supported · 1700s to early 1900s: −7.7 (−9.0 to −6.4); early 1900s to 2021 web: −9.4 (−10.1 to −8.7)
H3: contractions higher in the early 1900s than the 1700s, and higher again on the 2021 web · Not supported · 1700s to early 1900s: −0.3 (−1.0 to +0.4); early 1900s to 2021 web: +9.7 (+9.0 to +10.5)

## The baseline: typical values by genre and era

Medians (the typical text) for six habits. The 10th, 25th, 75th and 90th percentiles for all sixteen habits are in the summary file.

Genre · Era · Texts · Words per sentence · Contractions per 1,000 · Semicolons per 1,000 · Dashes per 1,000 · Long-word share · 'You' share of pronouns
Biography · 1700s · 90 · 27.3 · 3.6 · 10.5 · 0.6 · 0.13 · 0.03
Drama · 1700s · 90 · 16.2 · 24.0 · 12.4 · 10.3 · 0.09 · 0.24
Fiction · 1700s · 90 · 31.2 · 3.9 · 13.7 · 1.2 · 0.12 · 0.14
History · 1700s · 90 · 27.9 · 3.2 · 9.6 · 0.7 · 0.13 · 0.04
Letters · 1700s · 90 · 26.6 · 6.7 · 11.4 · 2.8 · 0.12 · 0.23
Philosophy and religion · 1700s · 90 · 32.2 · 1.5 · 11.8 · 0.0 · 0.13 · 0.04
Science · 1700s · 90 · 33.3 · 1.8 · 12.9 · 0.0 · 0.14 · 0.00
Travel · 1700s · 90 · 31.8 · 2.8 · 10.8 · 0.7 · 0.13 · 0.00
Biography · 1800s · 90 · 25.7 · 3.7 · 7.6 · 3.1 · 0.12 · 0.06
Children's books · 1800s · 90 · 20.4 · 12.3 · 6.7 · 3.4 · 0.08 · 0.32
Drama · 1800s · 90 · 15.1 · 14.0 · 10.3 · 5.9 · 0.10 · 0.26
Essays · 1800s · 90 · 28.5 · 3.6 · 7.8 · 3.6 · 0.15 · 0.13
Fiction · 1800s · 90 · 19.2 · 9.8 · 7.2 · 4.9 · 0.11 · 0.32
History · 1800s · 90 · 29.3 · 3.7 · 6.2 · 2.9 · 0.14 · 0.00
Letters · 1800s · 90 · 24.8 · 5.7 · 5.8 · 3.8 · 0.12 · 0.22
Philosophy and religion · 1800s · 90 · 26.3 · 3.6 · 6.9 · 2.1 · 0.13 · 0.00
Science · 1800s · 90 · 27.8 · 1.0 · 6.8 · 2.7 · 0.16 · 0.00
Travel · 1800s · 90 · 27.9 · 3.0 · 7.6 · 2.9 · 0.13 · 0.00
Biography · early 1900s · 90 · 23.8 · 4.7 · 3.0 · 1.9 · 0.13 · 0.04
Children's books · early 1900s · 90 · 15.5 · 20.2 · 1.0 · 2.0 · 0.08 · 0.32
Drama · early 1900s · 90 · 10.7 · 22.0 · 2.4 · 7.0 · 0.08 · 0.38
Essays · early 1900s · 90 · 22.2 · 5.4 · 5.5 · 3.9 · 0.11 · 0.07
Fiction · early 1900s · 90 · 16.2 · 15.6 · 4.0 · 4.9 · 0.10 · 0.31
History · early 1900s · 90 · 25.7 · 1.9 · 3.4 · 1.0 · 0.15 · 0.00
Letters · early 1900s · 90 · 24.3 · 5.8 · 3.0 · 1.9 · 0.14 · 0.20
Philosophy and religion · early 1900s · 90 · 25.6 · 3.1 · 5.9 · 1.0 · 0.15 · 0.07
Science · early 1900s · 90 · 21.7 · 1.0 · 2.9 · 1.7 · 0.15 · 0.00
Travel · early 1900s · 90 · 24.1 · 3.5 · 3.6 · 2.6 · 0.13 · 0.02
Essay-mill essays · 2021 web · 300 · 15.6 · 3.5 · 0.9 · 0.0 · 0.24 · 0.53
How-to content · 2021 web · 200 · 12.2 · 9.3 · 0.0 · 0.0 · 0.15 · 0.83
Marketing and SEO blogs · 2021 web · 300 · 15.2 · 19.7 · 0.0 · 0.5 · 0.16 · 0.86
Medium posts · 2021 web · 100 · 14.4 · 10.9 · 0.0 · 2.4 · 0.16 · 0.35
News (UK and US) · 2021 web · 500 · 17.2 · 16.0 · 0.0 · 0.0 · 0.16 · 0.11
Press releases · 2021 web · 300 · 19.5 · 7.7 · 0.0 · 2.1 · 0.27 · 0.00
Q&A forum · 2021 web · 100 · 11.1 · 12.7 · 1.0 · 0.0 · 0.15 · 0.39
Self-help and productivity · 2021 web · 300 · 14.2 · 20.1 · 1.4 · 1.4 · 0.13 · 0.49
Wikipedia · 2021 web · 100 · 10.3 · 3.5 · 1.9 · 0.0 · 0.22 · 0.07
Scientific abstracts (arXiv) · 2015-2022 · 1,000 · 24.1 · 0.0 · 0.0 · 0.0 · 0.28 · 0.00

## Checks

Full prose lines only (added after the main run): web pages and book passages re-measured on lines of eight or more words ending in sentence punctuation. News: 17.2 words per sentence on all text, 19.7 on prose lines (394 pages). How-to pages: 13.7 and 17.7 (98). Early-1900s non-fiction books: 24.6 and 25.0. Wikipedia pages could not be matched back to the crawl index for this check, so its web figures should be read as including headings and lists.

The first books harvest was discarded after two runs were found to have written into the same file, mislabelling some books' genre and era. The re-run is the data published here; the harvester now refuses to append to an existing file.

## Limitations

- Books are dated by when their author turned 40, not by publication; a book written young or old can sit in a neighbouring era.
- Project Gutenberg texts come from particular editions, and punctuation can be an editor's or printer's; semicolons and dashes are the most exposed.
- The web's sentence measure counts headings, captions, list items and references as short sentences, which pulls its figures down, most for Wikipedia (see Checks).
- Genres are assigned from catalogue subjects, so a book can sit in more than one (travel letters, for example) and some labels are broad.
- English only, and mostly British and American writing.

## Competing interests

ScriptGrain sells writing-voice measurement and generation, and the analyst builds it. The hypotheses, genres and sampling were fixed and published before anything was measured, nothing was re-run or selected (except one disclosed re-run after a file fault, below), and the scripts are public, so the numbers can be checked without taking our word for them.

## Reproducing this

- Pre-registration, sources, harvester, analysis and checks: docs/research/sgr-006-preregistration.md and scripts/experiments/sgr-006/ in the ScriptGrain repository.
- Every text is identified in the data (Project Gutenberg number and passage, web page URL and capture time, arXiv id), so any can be re-read and measured.

## Data

- [Every text measured, sixteen habits (CSV)](/research/sgr-006-texts.csv) (csv)
- [Baseline: percentiles by genre and era, hypotheses, trends (JSON)](/research/sgr-006-summary.json) (json)
- [Prose-line check (JSON)](/research/sgr-006-prose-check.json) (json)

Released under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/): free to reuse, including commercially, with credit to ScriptGrain and a link to this page.

## Citations

- [Project Gutenberg](https://www.gutenberg.org/)
- [Common Crawl, CC-MAIN-2021-43](https://commoncrawl.org/)
- [arXiv](https://arxiv.org/)

## How to cite

ScriptGrain (2026). What does normal writing look like? Three centuries of human prose, measured (Study SGR-006, conducted 4 October 2026). Dataset licensed CC BY 4.0. https://scriptgrain.com/research/what-does-normal-writing-look-like

## About the analyst

Jack Stovell has worked in finance and data for more than twelve years, building management reporting, forecasting and profitability models for advertising agencies and tech scale-ups, from SQL and Power BI reporting to board-level analysis. Since 2016 he has run Adapt Progress Evolve, an applied AI studio, where he builds and operates AI systems and data products: ScriptGrain's measurement of writing voice across 45 attributes, UK Spend, which brings 16.7 million rows of UK council spending into one queryable dataset, and more than thirty AI agents running in production. He designs ScriptGrain's studies and is accountable for every number in them.

- [jstov.uk](https://jstov.uk)
- [LinkedIn](https://www.linkedin.com/in/jackstovell)
- [GitHub](https://github.com/StovBuilds)

## About ScriptGrain (not a finding)

ScriptGrain's free tools place a piece of writing against reference ranges for habits like these. This study's tables are a published baseline anyone can use; they are not built into the product by this study.
