Advanced Martingale Theory

the martingale central limit theorem

The classical CLT needs independent summands. But many sums in probability and statistics are NOT independent — they are martingale differences: each term is conditionally mean-zero given the past, though it may depend on the past in other ways. The martingale central limit theorem (MCLT) is the far-reaching generalization that recovers asymptotic normality for such sums, with the role of the variance taken over by the (predictable) quadratic variation. It is the engine behind central limit theorems for Markov chains, stochastic approximation, maximum-likelihood and estimating-equation estimators, and the convergence of discrete martingales to Brownian motion.

A clean version (discrete-time, triangular array): let (X_{n,k}) be, for each n, a sequence of martingale differences with respect to a filtration (E[X_{n,k} given F_{n,k-1}] = 0). Suppose the predictable quadratic variation stabilizes, sum_k E[X_{n,k}^2 given F_{n,k-1}] -> sigma^2 in probability (a constant or a random limit), and a conditional Lindeberg condition holds, sum_k E[X_{n,k}^2 1_{|X_{n,k}| > epsilon} given F_{n,k-1}] -> 0 in probability for every epsilon > 0 (no single term dominates). Then the partial sums S_n = sum_k X_{n,k} converge in distribution to N(0, sigma^2) (a mixture of normals if sigma^2 is random). The continuous-time / functional form is sharper and more structural: a sequence of cadlag local martingales M^n with M^n_0 = 0 whose predictable quadratic variations converge, <M^n>_t -> c(t) for a deterministic continuous increasing c, and whose jumps are asymptotically negligible (a conditional-Lindeberg or maximal-jump condition), converges in the Skorokhod path space to a continuous Gaussian martingale with that variance function — for c(t) = t, the limit is Brownian motion. This is exactly how Donsker-type invariance principles and the Stroock-Varadhan martingale-problem approach to diffusions are organized.

Why it matters and the honest hypotheses: the MCLT replaces the independence of classical Lindeberg-Feller by adaptedness plus the two structural conditions — the quadratic variation converging (which sets the limiting variance) and a Lindeberg-type negligibility of jumps (which forces the limit to be Gaussian rather than, say, Poissonian). Both conditions are essential and cannot be dropped: without quadratic-variation convergence the variance is undetermined; without the Lindeberg/negligible-jumps condition the limit can be non-Gaussian (a compensated jump process converges to a Levy process, not a normal). When the limiting <M> is random, the limit is a variance mixture of normals (stable convergence is then the right mode), not a fixed normal — a subtlety that matters in statistics, where the observed (random) Fisher information appears as the conditional variance.

Rescale a simple symmetric random walk: S^n_t = (1/sqrt(n)) sum_{k <= nt} xi_k where the xi_k are independent fair-coin steps (a martingale-difference array). Its predictable quadratic variation is <S^n>_t = floor(nt)/n -> t, and the jumps 1/sqrt(n) are negligible, so the functional MCLT gives S^n -> Brownian motion in the Skorokhod topology — this is Donsker's invariance principle as an instance of the martingale CLT.

Donsker's theorem as a martingale CLT: quadratic variation -> t and vanishing jumps force the scaling limit to be Brownian motion.

The two hypotheses are not optional: dropping quadratic-variation convergence leaves the variance undetermined, and dropping the Lindeberg/negligible-jumps condition can make the limit a non-Gaussian Levy process; a random limiting bracket yields a normal variance-mixture, not a fixed normal.

Also called
martingale CLTMCLTfunctional martingale CLT鞅 CLT鞅泛函中央極限定理