Schilder's theorem
/ SHIL-der /
Schilder's theorem is the large deviation principle for Brownian motion in the small-noise limit, and it is the gateway from the large deviations of sums to the large deviations of paths. It asks: if you scale Brownian motion down by a small factor sqrt(epsilon), how unlikely is it that the resulting tiny-noise path wanders close to some prescribed deterministic trajectory? The answer is an LDP on path space whose rate function is the kinetic energy of the path — and this is the seed from which all small-noise (Freidlin-Wentzell) theory grows.
Let B be standard Brownian motion on [0,1] and consider the scaled process sqrt(epsilon) B as epsilon -> 0. Schilder's theorem states that the laws of sqrt(epsilon) B satisfy an LDP on the space of continuous paths C[0,1] (with the uniform topology), with speed 1/epsilon and good rate function the Cameron-Martin energy: I(phi) = (1/2) integral from 0 to 1 of (phi'(t))^2 dt if phi is absolutely continuous with phi(0) = 0 and square-integrable derivative, and I(phi) = +infinity otherwise. So P(sqrt(epsilon) B stays near phi) is roughly e^(-(1/2 epsilon) integral (phi')^2). The finite-energy paths are exactly the Cameron-Martin space, the reproducing kernel Hilbert space of Brownian motion, which is no accident: the rate function is the squared Cameron-Martin norm.
Schilder's theorem is the canonical 'level-3' / process-level LDP and the prototype for proving path-space LDPs by exponential tightness plus finite-dimensional Cramer bounds. Its real power comes through the contraction principle: pushing Schilder's LDP through the (continuous) Ito solution map of a small-noise SDE yields the Freidlin-Wentzell action functional, which controls exit times, escape paths and metastable transitions. The infinite-dimensional setting matters — exponential tightness on path space (using a modulus-of-continuity estimate for Brownian motion) is the non-trivial ingredient, since the unscaled paths are nowhere differentiable and only the rescaled finite-energy limits enter the rate function.
The cheapest path for sqrt(epsilon) B to reach level 1 at time 1 is the straight line phi(t) = t, with energy I = (1/2) integral_0^1 1 dt = 1/2. So P(sqrt(epsilon) B_1 near 1) is roughly e^(-1/(2 epsilon)), matching the Gaussian tail of B_1 ~ N(0,1) scaled by epsilon: P(sqrt(epsilon) B_1 > 1) ~ e^(-1/(2 epsilon)).
Schilder: rare Brownian paths cost their kinetic energy (1/2) integral (phi')^2.
The minimiser must be a finite-energy (Cameron-Martin, absolutely continuous L^2-derivative) path; Brownian sample paths themselves are nowhere differentiable and have infinite rate, so the 'most likely rare path' is smooth even though typical Brownian paths are not. The energy charges only the regular Cameron-Martin directions.