Large Deviations Theory

the exponential change of measure

The exponential change of measure — also called exponential tilting, the Cramer transform, or the Esscher transform — is the single most important technique in large deviations, the engine behind both the lower bound of the LDP and the practical computation of rare-event probabilities. The idea is disarmingly simple: a rare event is hard to study because it almost never happens under the true law, so instead change the probability law to a tilted one under which that rare event becomes typical, do an easy law-of-large-numbers calculation there, and pay for the change of measure with an explicit Radon-Nikodym factor.

Given a law mu and a parameter theta, the tilted (exponentially changed) measure is dmu_theta proportional to e^(theta X) dmu, with normaliser the moment generating function: dmu_theta/dmu = e^(theta X) / M(theta), M(theta) = E[e^(theta X)]. Under mu_theta the mean of X shifts to the derivative (log M)'(theta), so by choosing theta to solve (log M)'(theta) = x you make the desired deviation x the new typical value. The relative entropy of the tilted law with respect to mu is exactly the Cramer rate function: H(mu_theta | mu) = I(x), so the rate function literally is the cost of the optimal change of measure. For the LDP lower bound, you write the probability of the rare set as an expectation under mu_theta of the inverse likelihood ratio, the likelihood ratio contributes the e^(-n I(x)) factor, and an ordinary LLN under mu_theta supplies the rest.

Beyond proofs, exponential tilting is the foundation of importance sampling for rare-event simulation: naively estimating P(rare) by Monte Carlo needs exponentially many samples, but sampling under the optimally tilted law and reweighting gives an unbiased estimator with bounded relative error. The same construction is Girsanov's theorem in continuous time (the stochastic exponential is the path-space tilt), is the Gibbs/Boltzmann measure in statistical mechanics (the tilt by energy at inverse temperature beta), and is the Esscher transform in actuarial and financial pricing. The caveat: the optimal tilt exists only where the moment generating function is finite and steep; for heavy tails no finite theta makes the deviation typical, tilting fails, and the rare event is instead realised by a single big jump.

To estimate P(mean of n Exp(1) variables > 2) by simulation, naive Monte Carlo wastes nearly all samples. Tilt to Exp(rate 1/2) (the theta = 1/2 tilt whose mean is 2), sample there, and reweight by the likelihood ratio. The estimator concentrates with bounded relative error, and the leading exponential rate e^(-n(2 - 1 - log 2)) matches Cramer exactly.

Tilt so the rare event becomes typical; pay with the likelihood ratio.

The tilt mu_theta must be absolutely continuous with respect to mu (and vice versa), and the optimal theta exists only where the mgf is finite. For heavy-tailed laws no exponential tilt works; the correct importance-sampling change of measure is then a power-law / big-jump construction, not an exponential one.

Also called
exponential tiltingCramer transformEsscher transformtilted measure傾斜測度