Sanov's theorem
/ SAH-nof /
Sanov's theorem is the large deviation principle for the empirical distribution itself, not merely for the empirical mean. Where Cramer asks how unlikely it is for the average of n samples to be atypical, Sanov asks the far richer question: how unlikely is it for the entire empirical histogram of n samples to look like some chosen distribution nu other than the true law mu? This is called the level-2 LDP (Cramer is level 1), and its rate function is one of the most important objects in probability and information theory.
Let X_1, ..., X_n be iid samples from a law mu, and let L_n = (1/n) Sum delta_{X_i} be the empirical measure, a random probability distribution. Sanov's theorem states that L_n satisfies an LDP on the space of probability measures (with the weak topology, or stronger tau-topology in the refined version) with speed n and good rate function equal to the relative entropy I(nu) = H(nu | mu) = integral of log(dnu/dmu) dnu. In words: the probability that the empirical measure of n iid mu-samples resembles nu decays like e^(-n H(nu|mu)). Relative entropy vanishes uniquely at nu = mu, recovering the Glivenko-Cantelli convergence of L_n to mu, and grows as nu departs from mu.
Sanov is the bridge between probability and information theory and the source of the maximum entropy / Gibbs conditioning principle. If you condition n iid samples on their empirical mean being atypical, the conditional empirical measure converges to the I-projection: the distribution closest to mu in relative entropy subject to the imposed constraint, which is exactly the exponentially tilted (Gibbs) measure. Cramer's theorem can be derived from Sanov by the contraction principle, pushing the empirical measure forward through the mean functional. The hypothesis that needs care is the topology: Sanov holds in the weak topology generally, and in the finer tau-topology under additional integrability, but the relative entropy rate function is the same.
Roll a fair die n times; the true law mu is uniform on {1,...,6}. The chance the empirical frequencies come out close to nu = (1/2, 1/10, 1/10, 1/10, 1/10, 1/10) decays like e^(-n H(nu|mu)), where H(nu|mu) = Sum nu_i log(nu_i / (1/6)) = Sum nu_i log(6 nu_i). The cheapest non-uniform histogram dominates exactly as Sanov predicts.
Sanov: the empirical distribution deviates at rate relative entropy H(nu|mu).
Relative entropy H(nu|mu) is finite only when nu is absolutely continuous with respect to mu; if nu puts mass where mu has none, H = +infinity, correctly signalling that such an empirical measure is impossible (not just exponentially rare) for mu-samples.