the strong law as an ergodic theorem
This entry is the bridge between Volume I and ergodic theory: it shows that the strong law of large numbers for iid sequences is the single simplest instance of Birkhoff's theorem. The point is not a new proof for its own sake, but the recognition that the SLLN, far from being special to independence, is the surface of a much wider law of averages for stationary dependent data.
Here is the dictionary. Let (X_1, X_2, ...) be iid with E|X_1| < infinity. Realise them on the sequence space Omega = R^N with the product law P = (law of X_1)^(otimes N), and let T be the shift (T omega)_n = omega_(n+1) and f(omega) = omega_1 the first-coordinate function. Then X_k = f(T^(k-1) omega), so the sample mean (1/n) sum X_k is exactly the Birkhoff time average (1/n) sum f(T^k omega). The shift is measure-preserving (the product measure is shift-invariant — that is just stationarity) and ergodic (the Kolmogorov 0-1 law makes the invariant sigma-algebra trivial). Birkhoff therefore gives (1/n) sum X_k -> E[ f given I ] = E[X_1] almost surely. That is the strong law.
The payoff is generality. Replace "iid" by "stationary and ergodic" and the same argument gives the ergodic theorem for stationary processes: for any stationary ergodic sequence with E|X_1| < infinity, the sample average converges a.s. to E[X_1] even though the X_k are dependent. This is the workhorse behind consistency of estimators from time series, Monte Carlo with Markov chains (an ergodic chain in equilibrium obeys an SLLN), and information theory (the Shannon-McMillan-Breiman theorem). The honest caution: ergodicity is essential. A stationary but non-ergodic sequence still has converging averages, but the limit is the random E[ X_1 given I ] — for instance, draw a coin bias once and then flip forever: the average converges to that random bias, not to its overall mean.
A stationary ergodic Markov chain (X_n) in its stationary distribution pi obeys (1/n) sum_(k<n) g(X_k) -> sum_x pi(x) g(x) almost surely for any g with finite pi-mean. This is the SLLN for dependent data and the justification for time-averaging a single long MCMC run instead of many independent samples.
The same ergodic-theorem skeleton that yields the iid strong law yields the law of large numbers for stationary, dependent data.
Independence is a luxury, not a requirement: the strong law needs only stationarity + ergodicity + integrability. Drop ergodicity and the average still converges, but to a random limit, not to the mean.