the contraction principle
The contraction principle is the transfer rule of large deviations: it says how a large deviation principle pushes forward through a continuous map. If you know the LDP for some primitive random object and you care about a continuous function of it, you do not need to start over — the image automatically satisfies an LDP, and you get its rate function by a clean minimisation. This is what makes large deviations modular: prove one foundational LDP, then derive many others by contraction.
Suppose X_n satisfies an LDP on a space X with good rate function I, and T: X -> Y is continuous. Then the image random variables Y_n = T(X_n) satisfy an LDP on Y with the good rate function J(y) = inf over x with T(x) = y of I(x) (the infimum of I over the preimage of y, with the convention inf over the empty set = +infinity). The intuition is exactly right and worth stating in words: the cheapest way for the image to deviate to y is to take the cheapest preimage x that maps to y, so the cost of y is the minimal cost among all causes x of y. Goodness of I is what guarantees the infimum is attained and that J is again a good rate function.
The contraction principle is everywhere. Cramer's theorem follows from Sanov's by contracting the empirical measure through the mean functional nu -> integral x dnu, turning relative entropy into the Legendre transform. Freidlin-Wentzell rate functions arise by contracting Schilder's Brownian LDP through the Ito map that sends a driving path to the SDE solution. A crucial honesty point: the image of a convex rate function under a nonlinear contraction is typically non-convex, which is precisely how multiple equally-likely deviation mechanisms, metastable wells and phase coexistence enter the theory — they are visible as multiple local minima of the contracted J that a Legendre transform could never produce.
Let X_n satisfy an LDP on R with good rate I, and take T(x) = x^2. Then Y_n = X_n^2 satisfies an LDP with J(y) = min(I(sqrt(y)), I(-sqrt(y))) for y >= 0 and J(y) = +infinity for y < 0. The two preimages compete and the cheaper one wins — a non-monotone, possibly non-convex J even when I is convex.
Contraction: cost of the image = cheapest cost among its preimages.
The map must be continuous and the source rate function good (compact level sets) for the standard contraction principle; for merely measurable or discontinuous maps the formula can fail. An inverse contraction (pulling an LDP back through a map) needs an exponential approximation argument, not just continuity.