the Gaussian log-Sobolev inequality
/ so-bo-LYOV /
The Gaussian log-Sobolev inequality (Gross) is the master functional inequality of the Gaussian world. It bounds an entropy-like measure of how non-constant a function is by the average size of its gradient — a single clean inequality that simultaneously encodes concentration of measure, the exponential convergence of the Ornstein-Uhlenbeck semigroup, and (equivalently) hypercontractivity. It is the dimension-free upgrade of the Poincare inequality and the reason Gaussian measures are so analytically friendly.
Let gamma be the standard Gaussian measure on R^n and define the entropy of f^2 by Ent_gamma(f^2) = integral f^2 log f^2 dgamma - (integral f^2 dgamma) log(integral f^2 dgamma). Gross's inequality states Ent_gamma(f^2) <= 2 integral | grad f |^2 dgamma for all smooth f, with the constant 2 sharp and, crucially, INDEPENDENT of the dimension n. (Compare the Gaussian Poincare inequality, Var_gamma(f) <= integral |grad f|^2 dgamma, which controls variance; LSI controls the stronger entropy functional and implies Poincare by linearizing near a constant.) The dimension-free constant is what makes it powerful in infinite dimensions and in product / high-dimensional settings, where it tensorizes: the LSI for gamma on R^1 automatically gives the LSI for gamma^n on R^n with the SAME constant. Gross's theorem is that this inequality is EQUIVALENT to Nelson's hypercontractivity of the OU semigroup, the two being Legendre-type transforms of one another along the time parameter.
Why it matters: from LSI one derives, by the Herbst argument, sub-Gaussian concentration for Lipschitz functions, P( f - E f >= u ) <= exp(-u^2 / 2) for 1-Lipschitz f — recovering the Gaussian isoperimetric concentration with an explicit constant; one derives exponential decay of entropy along the OU flow (hence fast mixing); and one gets a robust, tensorizable tool that survives in infinite dimensions where isoperimetry is delicate. The honest caveats: the gradient must be the EUCLIDEAN gradient matched to the Gaussian's covariance, and the dimension-free constant is special to log-concave measures — for general measures one only gets an LSI with a constant depending on the measure (via, e.g., the Bakry-Emery curvature criterion), if at all.
Apply LSI to f = exp(lambda g / 2) for a 1-Lipschitz g (the Herbst argument): the entropy becomes a differential inequality for the log-moment-generating function, which integrates to E[exp(lambda(g - E g))] <= exp(lambda^2 / 2). By Markov this is exactly the sub-Gaussian tail P(g - E g >= u) <= exp(-u^2 / 2) — Gaussian concentration produced from the gradient bound alone.
The log-Sobolev inequality plus the Herbst argument yields Gaussian concentration P(g - E g >= u) <= exp(-u^2/2) for Lipschitz g.
The constant 2 is sharp and DIMENSION-FREE (it tensorizes), which is special to log-concave measures; general measures need not satisfy any log-Sobolev inequality.