Cramer's theorem
/ KRAH-mair /
Cramer's theorem is the founding result of large deviations: it gives the exact exponential rate at which the empirical mean of independent identically distributed random variables deviates from its expectation. It is the large-deviation analogue of the law of large numbers, answering the question the LLN raises but does not address — not just that the mean concentrates, but how astronomically unlikely a fixed-size deviation is, and what governs that.
Let X_1, X_2, ... be iid real random variables and S_n = (X_1 + ... + X_n)/n the empirical mean. Cramer's theorem states that the laws of S_n satisfy an LDP with speed n and rate function I given by the Legendre-Fenchel transform of the logarithmic moment generating function: I(x) = sup over theta of (theta x - log E[e^(theta X_1)]). The rate function is convex, non-negative, and vanishes exactly at the mean E[X_1]. The intuition behind the rate function is the exponential tilt: to make the empirical mean sit near an atypical value x, the cheapest mechanism is to replace the true law by the exponentially tilted (Cramer-transformed) law dmu_theta proportional to e^(theta X) dmu, with theta chosen so that x becomes the new mean; the cost of this change of measure, measured by relative entropy, is exactly I(x).
Cramer's theorem holds in R under mild conditions and extends to R^d (Cramer in R^d) and to separable Banach spaces (Cramer-Donsker-Varadhan) when the log-mgf is finite in a neighbourhood of the origin, a condition ensuring all exponential moments and a good rate function. If the moment generating function is finite only at the origin (heavy tails), Cramer's theorem in its classical exponential form fails and deviations follow a polynomial 'big jump' principle instead — the deviation is realised not by many small contributions but by a single large summand. So the light-tail hypothesis cannot be dropped silently.
For iid Exponential(1) variables, log E[e^(theta X)] = -log(1 - theta) for theta < 1, so I(x) = sup_theta (theta x + log(1 - theta)) = x - 1 - log x for x > 0. It vanishes at x = 1 (the mean) and tells you, for example, that the chance the average of n exponentials exceeds 2 decays like e^(-n(1 - log 2)) = e^(-0.307 n).
Cramer: the rate is the Legendre transform of the log-mgf.
The light-tail (finite exponential moments near 0) hypothesis is essential. For heavy-tailed summands the rare deviation is dominated by one big jump, the decay is polynomial not exponential, and Cramer's exponential rate function does not describe it.