Credibility Theory

empirical Bayes credibility

/ em-PEER-ih-kul BAYZ /

Buhlmann credibility needs two hidden numbers to work — the expected process variance (EPV) and the variance of the hypothetical means (VHM). But in real life nobody hands you these; the 'population of risks' and its true variances are invisible. Empirical Bayes is the clever trick of estimating those structural parameters from the very same data you are about to credibility-weight. You use the whole portfolio to learn how risks scatter, then feed that back to rate each individual risk.

The recipe, in words: collect the experience of many risks; estimate the expected process variance by averaging each risk's own internal variance (the within-risk wobble); estimate the variance of the hypothetical means by taking the total spread of risk averages and subtracting off the part explained by that internal noise (because observed differences between risks are partly real and partly luck). Then k = EPV / VHM and Z = n/(n + k) follow. The phrase 'empirical Bayes' captures the spirit: it is a Bayesian-flavoured shrinkage of each risk toward the collective, but the prior is estimated from the data rather than assumed.

Empirical Bayes is how Buhlmann and Buhlmann-Straub are actually implemented in pricing offices. Its honesty is also its limit: because EPV and VHM are themselves estimates from finite data, the resulting k and Z carry estimation error, and the VHM estimate can even come out negative and be floored at zero (forcing Z = 0). With few risks or short histories the structural parameters are shaky, so the whole apparatus should be used with judgement, not as a black box.

A motor book has 200 policyholders, each observed several years. Averaging every driver's own year-to-year variance estimates EPV = 0.38; the spread of drivers' averages minus that noise estimates VHM = 0.047, giving k about 8.1.

The portfolio teaches its own structural parameters; those then set each risk's Z.

The estimated VHM can be negative because it is a subtraction of two noisy quantities; convention sets it to zero, which makes Z = 0. That is a sign the data cannot reliably tell the risks apart, not a computational glitch to ignore.

Also called
empirical Bayes estimationnonparametric Buhlmann estimation经验贝叶斯估计經驗貝氏估計