Gaussian Processes & Gaussian Measures

the Borell-TIS inequality

/ bo-RELL /

How much does the supremum of a Gaussian process fluctuate? The Borell-TIS inequality (Borell, and independently Tsirelson-Ibragimov-Sudakov) gives a stunningly clean answer: the supremum concentrates around its mean (or median) with a Gaussian tail whose only relevant scale is the LARGEST single standard deviation in the process, not the number of variables. This is the cornerstone concentration result for Gaussian suprema and a workhorse of empirical-process theory and extreme-value bounds.

Let (X_t)_(t in T) be a centered Gaussian process that is almost surely bounded, write M = sup_(t in T) X_t, and let sigma^2 = sup_(t in T) E[X_t^2] be the largest pointwise variance. The inequality states P( | M - E[M] | >= u ) <= 2 exp( -u^2 / (2 sigma^2) ) for every u > 0 (and with E[M] replaceable by a median). In words: the supremum of a Gaussian process is itself a sub-Gaussian random variable with variance proxy sigma^2 — it concentrates exactly as tightly as a single Gaussian of the worst-case variance, no matter how rich or high-dimensional the index set T is. The proof route is the Gaussian isoperimetric / log-concavity machinery (or the Gaussian Poincare-type concentration for Lipschitz functions): M is a 1-Lipschitz function of the underlying Gaussian vector with the Euclidean-to-sigma scaling, and Lipschitz functions of Gaussians concentrate.

Why it matters: Borell-TIS reduces the hard problem 'how big is sup X_t' to two SEPARATE problems — the fluctuation (controlled, for free, by sigma alone) and the LOCATION E[M] (the genuinely hard part, attacked by Dudley's entropy bound or Sudakov-Fernique). It is the reason that for Gaussian suprema you only ever sweat the mean. Two honest caveats: the inequality controls the deviation of M from its OWN mean, not the value of that mean; and it requires the process to be a.s. bounded (otherwise M = +infinity and there is nothing to concentrate).

For n independent standard normals Z_1, ..., Z_n, each variance sigma^2 = 1, Borell-TIS gives P( | max_i Z_i - E[max_i Z_i] | >= u ) <= 2 exp(-u^2 / 2), with NO dependence on n. The mean grows like E[max Z_i] approximately sqrt(2 log n), but the fluctuation around it stays order 1 — fluctuations do not grow with dimension, only the location does.

The maximum of many Gaussians fluctuates at scale sigma regardless of n; only its mean drifts (like sqrt(2 log n)).

Borell-TIS bounds only the FLUCTUATION of the supremum around its mean (using sigma, the largest variance); finding the mean itself is the separate hard problem solved by entropy/comparison methods.

Also called
Borell-Tsirelson-Ibragimov-Sudakov inequalityGaussian concentration of the supremum高斯上確界集中不等式