Statistics, Data & Modeling

hypothesis testing, p-values, and significance

/ hy-POTH-eh-sis; PEE-value /

Suppose a new underwriting rule is supposed to lower claims. After a year, claims are a bit lower — but is that a real effect, or just the luck of the draw, the kind of dip that happens by chance even if the rule did nothing? Hypothesis testing is the formal courtroom for that question. It starts by assuming the boring explanation is true (the rule does nothing), then asks whether the data are too surprising to square with that assumption.

The boring explanation is the null hypothesis; the interesting one is the alternative. We compute a test statistic measuring how far the data sit from what the null predicts, then a p-value: the probability of seeing data at least this extreme if the null were true. A small p-value (commonly below 0.05) means the data would be surprising under the null, so we 'reject the null' and call the result statistically significant. For instance, a p-value of 0.01 says results this striking would arise only 1% of the time by pure chance under the null. Crucially, a large p-value does not prove the null — it merely fails to provide evidence against it.

Actuaries use these tests to ask whether a rating variable truly affects loss, whether observed mortality differs from a standard table, or whether a model term earns its place. But the framework is widely abused, so honest use demands care. The 0.05 cutoff is an arbitrary convention, not a law of nature. Statistical significance is not practical importance — with enough data a trivial, useless difference becomes 'significant.' And a p-value is not the probability the null is true, nor the probability your finding is a fluke; it is a statement about data assuming the null. Running many tests and reporting only the significant ones (p-hacking) manufactures false discoveries.

An actuary tests whether a region's mortality exceeds the standard table. The test gives a p-value of 0.002, far below 0.05, so the excess is statistically significant — the region genuinely seems to have heavier mortality, not just an unlucky year.

A small p-value flags a result too extreme to comfortably blame on chance.

Statistical significance is not practical importance, and a p-value is not the probability the null is true. With huge data, a meaningless difference can be 'highly significant.'

Also called
significance testp-value假设检定顯著性檢驗