JOVANA
Explore Library Glossary Getting Started Three Levels Fields How it works Mission
Join the mission
Back to the library
Statistics 1908

The Probable Error of a Mean

William Sealy Gosset ("Student")

How far can you trust an average drawn from only a handful of measurements?

Choose your version
In depth · the introduction

How sure can you be of an average when you've only measured a few things?

Small samples deserve wider margins

When you measure something many thousands of times, you learn both its average and how spread out it is, and you can trust your error bars. But real experiments are usually small — four plots of barley, five patients, a handful of trials — and then you face a quieter problem: you have to guess the spread from the very same tiny set of numbers. That guess is shaky, so honest error bars must be wider than the textbook normal curve would suggest. Gosset worked out exactly how much wider, giving a new bell-shaped curve with heavier tails for small samples.

The brewer who wrote as 'Student'

William Sealy Gosset was not a professor but a chemist and head experimental brewer at the Guinness brewery in Dublin, trying to grow better barley and brew more consistent beer with only a few experimental batches to learn from. Guinness forbade its staff from publishing, afraid of leaking trade secrets, so when Gosset solved the small-sample problem in 1908 he published under the pen name 'Student.' To test his new curve he ran a startlingly hands-on experiment: he took 3,000 prisoners' body measurements, wrote each on cardboard, shuffled the deck, and dealt out 750 little samples of four — an early simulation done entirely by hand. The young Ronald Fisher later proved Gosset's result rigorously, fixed a subtle counting error, named the curve 'Student's t' in his friend's honour, and the two exchanged letters for decades.

Why it mattered

Almost every experiment that has ever mattered was run on a small sample — you rarely get thousands of patients or thousands of harvests. Before Gosset, statisticians had no honest way to draw conclusions from a few measurements; after him, they did. The t-test he made possible became the standard way to ask 'is this difference real or just luck?' in medicine, agriculture, psychology and, today, the A/B tests that decide which version of a web page you see. It let science be both small and rigorous at once.

Judging from a few reviews

Imagine rating a restaurant. From 300 reviews you can be fairly confident in the average score. From just 3 reviews, the average might be lucky or unlucky — so a sensible person hedges, allowing a much wider range of what the 'true' quality might be. Student's t-distribution is the exact mathematics of that hedging: the fewer measurements you have, the wider the curve spreads, and the more cautious your conclusion must be. With many measurements it tightens back into the familiar bell curve.

A chart with two bell-shaped curves: a dashed normal curve and a solid t-curve. A slider changes the sample size; smaller samples make the solid curve lower and wider, and a marker for the 95% cutoff slides outward.

Where it sits

Gosset's curve is a cornerstone of practical statistics, sitting alongside the Library's other foundations of reasoning under uncertainty. Bayes (1763) showed how a single observation should revise a belief; Kolmogorov (1933) gave probability its rigorous footing; Ronald Fisher (1935) built the whole science of experimental design on top of significance and the t-test. Every time a study reports a result 'with 95% confidence,' or a company decides one web design beat another, Gosset's small-sample mathematics is quietly at work.

The original document
Original source text
Student [W. S. Gosset] · The Probable Error of a Mean · Biometrika 6(1): 1–25 · March 1908
What the paper set out to do
The paper opens by noting a gap everyone had stepped around: the standard way of judging how close a sample mean lies to the true mean assumes the sample is large enough to treat its measured spread as if it were the real one. Gosset asks the unhandled question — what to do when the sample is small, so that the spread itself is only a shaky estimate.
Gosset's route to the distribution
Assuming the underlying measurements are normally distributed, he works out the distribution of the sample standard deviation, argues that for normal data the sample mean and the sample spread vary independently, and from these obtains the distribution of the standardised quantity z = (mean − true mean) / s. Lacking a fully rigorous proof, he fixes the form of the curve by matching its moments and tabulates it.
Checking it by experiment
To test his curve, Gosset took W. R. Macdonell's measurements of 3,000 criminals — height and left-middle-finger length — wrote each on a piece of cardboard, shuffled them, and dealt out 750 samples of four. Comparing the spread of those small samples against his formula was an early hand-run simulation, a forerunner of the Monte Carlo method.
[ … ]
Student [W. S. Gosset] · the Galton Laboratory & the Guinness brewery · 1908