Authorship attribution
Authorship attribution is the task of working out who wrote a text by comparing its measurable style habits, above all its function words, with writing by candidate authors. It is the oldest application of stylometry, from the Federalist Papers to forensic cases, and it works best when the texts compared are the same kind of writing.
The numbers
- Disputed Federalist Papers Mosteller and Wallace attributed with a Bayesian analysis in 1964: 12 (ScriptGrain: what stylometry research actually shows)
- Best overall score in PAN 2021 authorship verification, within one kind of writing (English fanfiction): 0.9545 (Kestemont et al. (2021), Overview of the Cross-Domain Authorship Verification Task at PAN 2021)
- Best submitted overall score in PAN 2022, when the two texts were different kinds of writing; a naive baseline scored 0.600: 0.587 (Stamatatos et al. (2022), Overview of the Authorship Verification Task at PAN 2022)
- Sample size below which attribution on English novels was 'simply disastrous' (Eder, DH2010): 3,000 words (Eder, M. (2010). Does Size Matter? Authorship Attribution, Small Samples, Big Problem. DH2010)
What authorship attribution is
Authorship attribution is the task of working out which of several candidate authors most likely wrote a disputed text. That's different from authorship verification, which asks a narrower question: do these two texts share an author, full stop, without a lineup of suspects. Attribution needs candidates. Verification just needs a pair.
The classic case is the Federalist Papers. Mosteller and Wallace ran a Bayesian analysis in 1964 across all 146 essays by Jay, Hamilton and Madison, and used it to settle the 12 papers both Hamilton and Madison had claimed. Before that, the first serious quantitative attempt at anything like this is usually dated to Mendenhall's 1887 study of Shakespeare's plays. So the field is old. Older than most people assume when they hear "stylometry" and think it's a machine-learning trend from the last decade.
The underlying idea connects to what's sometimes called a stylometric fingerprint: the pattern of measurable habits that make one person's writing distinguishable from another's, at least statistically.
How it is done
Here's the mechanics. You take known writing samples from each candidate, measure a set of features, then measure the same features in the disputed text, then compare. The features that carry the most reliable signal aren't the vivid ones. They're the boring ones: function words. Stamatatos's 2009 survey put it plainly: function words "are used in a largely unconscious manner by the authors and they are topic-independent." That's the whole appeal. You can't fake unconscious habits easily, and they don't shift just because the topic did.
Published research has used sets of function words of different sizes: 303, 365, 480 and 675 words. No single "correct" list exists. Researchers pick a set, measure frequencies across candidates and the disputed text, and see which candidate's rates the disputed text resembles most closely. Simple in concept. Fiddly in execution, and entirely dependent on having enough text to measure in the first place.
Where it works and where it fails
Context decides a great deal here, and the PAN shared tasks make that obvious. PAN 2021 tackled authorship verification within a single variety of writing, English fanfiction, open set. The best system scored an AUC of 0.9869, with an overall score of 0.9545. Excellent numbers.
PAN 2022 changed one thing: the two texts being compared came from different discourse types, essays, emails, text messages, business memos. The best submitted system scored 0.587 overall. A naive character n-gram baseline, nothing clever, scored 0.600 and beat every submission. The only change was the kind of writing, and the results fell to near chance.
Eder's DH2010 work on English novels found something similar from another angle: "samples shorter than 5000 words provide a poor guessing," and "below the size of 3000 words, the obtained results are simply disastrous." Length matters, whatever the method.
Both findings point the same way for forensic use: the comparison means most with plenty of text and the same kind of writing on both sides; the five famous forensic cases are worth reading if you want the messy real-world version. For the fuller technical picture, the stylometry research page covers the field beyond this one task.
What it cannot tell you
Attribution gives you a probability among the candidates you fed it. Not proof. Never proof. It's weak on short texts, weaker still when the disputed text and the known samples come from different genres, because the features you're measuring start reflecting the genre more than the person.
ScriptGrain uses the same underlying measurements for a different job entirely. It builds a voice profile from a writer's own writing, with that writer's involvement, and scores new drafts against their own established profile. It does not identify anonymous authors, and it was never built to. Its AI-detect feature is a separate human-versus-AI read, and that too is not proof of who wrote something, just a signal worth weighing.
Worked example
two candidate authors, call them A and B, and one disputed letter. Both candidates have several thousand words of known writing available. The letter runs to about 2,000 words, comparable in length and, crucially, comparable in kind: it's a personal letter, and so are the samples from A and B.
say the comparison looks at four illustrative function-word rates, the frequency per thousand words of "of," "that," "which" and "but." Candidate A's known writing sits at roughly 9, 6, 2 and 4. Candidate B sits at 14, 3, 5 and 7. The disputed letter comes in at 9, 5, 2, and 4, tracking A far more closely across the board.
this points toward A, because the comparison held genre constant, personal letter against personal letter, though at about 2,000 words the letter is shorter than Eder found reliable for novels, so it is a lead, never a conclusion. Swap the disputed text for a business memo written by the same hypothetical author and the numbers would drift even with the author unchanged, because memos and letters lean on different function words for different reasons. That's the catch with attribution: it measures style, but style bends to context, and pretending otherwise is how confident-sounding conclusions turn out to be wrong.
Sources
- ScriptGrain: what stylometry research actually shows
- Stamatatos, E. (2009). A Survey of Modern Authorship Attribution Methods. JASIST 60(3), 538-556 (doi)
- Stamatatos et al. (2022), Overview of the Authorship Verification Task at PAN 2022