population vs sample
/ POP-yuu-LAY-shun vurs SAM-pul /
Suppose you want to know the average height of every adult in a country. Measuring all of them is impossible — there are millions, and new people are born and die every day. So instead you measure a few thousand, carefully chosen, and use that smaller group to make a statement about everyone. The whole group you really care about is the population; the smaller group you actually observe is the sample. Almost all of statistics lives in the gap between these two.
More precisely, the population is the complete collection of every item the question is about — every policyholder, every claim that could occur, every possible roll of a die — together with the true numbers that describe it. Those true numbers, like the population mean or the population variance, are called parameters, and they are usually fixed but unknown. A sample is a subset that we observe, and numbers computed from the sample, like the sample mean, are called statistics. We use the statistics as estimates of the parameters. A key idea: if I draw a different sample, I get a different sample mean, even though the population mean never changes. The sample is random; the population is not.
For actuaries this distinction is the foundation of everything. The 'population' might be all the lives a mortality table is meant to describe, or every car-insurance loss that the portfolio could ever generate; the 'sample' is the finite stack of data the company has actually collected. The whole art is reaching honest conclusions about the unseen population from a limited, possibly biased sample. The deepest danger is a sample that does not represent the population — for example, estimating future claims only from the policyholders who stayed (survivors), which quietly distorts the picture.
An insurer cannot know the true average claim of every driver it might ever cover (the population). It has last year's 50,000 actual claims (a sample), averages them to get $1,820, and treats that as its best guess of the population's true mean claim.
The sample mean ($1,820) is a known statistic; the population mean it estimates is unknown.
A bigger sample reduces random error but does nothing to fix a biased one. A million phone-survey responses are still useless for the whole country if only people who answer phones are in it.