Moments. A full-dimensional log-concave law has exponentially decaying
coordinate tails. To see this directly, Prékopa applied to the density times
1 { x i ≥ t } \mathbf1_{\{x_i\ge t\}} 1 { x i ≥ t } shows that the survival function F i ( t ) F_i(t) F i ( t ) is
log-concave. For an unbounded upper tail, choose b > c b>c b > c with
0 < F i ( b ) < F i ( c ) 0<F_i(b)<F_i(c) 0 < F i ( b ) < F i ( c ) . Concavity of log F i \log F_i log F i gives for t ≥ b t\ge b t ≥ b
F i ( t ) ≤ F i ( c ) exp ( t − c b − c log F i ( b ) F i ( c ) ) . F_i(t)\le F_i(c)
\exp\left(\frac{t-c}{b-c}\log\frac{F_i(b)}{F_i(c)}\right). F i ( t ) ≤ F i ( c ) exp ( b − c t − c log F i ( c ) F i ( b ) ) . For a bounded upper tail a bound is immediate. Apply the same argument to
− X i -X_i − X i and use the finite-coordinate union bound for ∣ X ∣ > t |X|>t ∣ X ∣ > t .
Integrating k t k − 1 P ( ∣ X ∣ > t ) kt^{k-1}\mathbb P(|X|>t) k t k − 1 P ( ∣ X ∣ > t ) proves E ∣ X ∣ k < ∞ \mathbb E|X|^k<\infty E ∣ X ∣ k < ∞
for every real k > 0 k>0 k > 0 . The same holds on a proper affine support by using
coordinates there.
Convolution curvature. More generally let a full-dimensional law have
curvature at least a I aI a I with a ≥ 0 a\ge0 a ≥ 0 , meaning its density is proportional to
exp ( − a ∣ y ∣ 2 / 2 − W ( y ) ) \exp(-a|y|^2/2-W(y)) exp ( − a ∣ y ∣ 2 /2 − W ( y )) for an extended-real convex function W W W .
The law of X + ε G X+\sqrt\varepsilon G X + ε G has strictly positive density
ρ ε ( x ) = ( 2 π ε ) − n / 2 ∫ exp ( − ∣ x − y ∣ 2 / ( 2 ε ) ) d μ ( y ) . \rho_\varepsilon(x)=(2\pi\varepsilon)^{-n/2}
\int\exp(-|x-y|^2/(2\varepsilon))\,d\mu(y). ρ ε ( x ) = ( 2 π ε ) − n /2 ∫ exp ( − ∣ x − y ∣ 2 / ( 2 ε )) d μ ( y ) . Every derivative of the Gaussian kernel is bounded, uniformly in its
argument, so differentiation under the probability integral shows
ρ ε \rho_\varepsilon ρ ε is smooth. The completion of squares
a 2 ∣ y ∣ 2 + ∣ x − y ∣ 2 2 ε = a ∣ x ∣ 2 2 ( 1 + a ε ) + 1 + a ε 2 ε ∣ y − x 1 + a ε ∣ 2 \frac a2|y|^2+\frac{|x-y|^2}{2\varepsilon}
=\frac{a|x|^2}{2(1+a\varepsilon)}
+\frac{1+a\varepsilon}{2\varepsilon}
\left|y-\frac{x}{1+a\varepsilon}\right|^2 2 a ∣ y ∣ 2 + 2 ε ∣ x − y ∣ 2 = 2 ( 1 + a ε ) a ∣ x ∣ 2 + 2 ε 1 + a ε ∣ ∣ y − 1 + a ε x ∣ ∣ 2 expresses ρ ε ( x ) \rho_\varepsilon(x) ρ ε ( x ) as
exp [ − a ∣ x ∣ 2 / ( 2 ( 1 + a ε ) ) ] \exp[-a|x|^2/(2(1+a\varepsilon))] exp [ − a ∣ x ∣ 2 / ( 2 ( 1 + a ε ))] times a log-concave function of x x x :
the residual integrand is jointly log-concave in ( x , y ) (x,y) ( x , y ) and Prékopa applies.
Thus V ε = − log ρ ε V_\varepsilon=-\log\rho_\varepsilon V ε = − log ρ ε has
D 2 V ε ⪰ a ( 1 + a ε ) − 1 I D^2V_\varepsilon\succeq a(1+a\varepsilon)^{-1}I D 2 V ε ⪰ a ( 1 + a ε ) − 1 I .
For the upper bound let π ε , x \pi_{\varepsilon,x} π ε , x be the probability law with
density proportional to exp [ − ∣ x − y ∣ 2 / ( 2 ε ) ] \exp[-|x-y|^2/(2\varepsilon)] exp [ − ∣ x − y ∣ 2 / ( 2 ε )] relative to μ \mu μ .
Differentiation of its mean gives
D 2 V ε ( x ) = ε − 1 I − ε − 2 Cov ( π ε , x ) ⪯ ε − 1 I . D^2V_\varepsilon(x)=\varepsilon^{-1}I
-\varepsilon^{-2}\operatorname{Cov}(\pi_{\varepsilon,x})
\preceq\varepsilon^{-1}I. D 2 V ε ( x ) = ε − 1 I − ε − 2 Cov ( π ε , x ) ⪯ ε − 1 I . All derivatives are legitimate: a polynomial in y y y times this Gaussian
kernel is bounded in y y y , locally uniformly in x x x .
After dividing the random vector by 1 + ε \sqrt{1+\varepsilon} 1 + ε , covariance
becomes ( Cov μ + ε I ) / ( 1 + ε ) ⪯ I (\operatorname{Cov}\mu+\varepsilon I)/(1+\varepsilon)\preceq I ( Cov μ + ε I ) / ( 1 + ε ) ⪯ I
and the Hessian bounds become
a ( 1 + ε ) 1 + a ε I ⪯ D 2 V μ ε ⪯ 1 + ε ε I . \frac{a(1+\varepsilon)}{1+a\varepsilon}I
\preceq D^2V_{\mu_\varepsilon}
\preceq\frac{1+\varepsilon}{\varepsilon}I. 1 + a ε a ( 1 + ε ) I ⪯ D 2 V μ ε ⪯ ε 1 + ε I . For 0 < a ≤ 1 0<a\le1 0 < a ≤ 1 the lower bound is at least a I aI a I , proving regularity.
In the common coupling, X + ε G → X X+\sqrt\varepsilon G\to X X + ε G → X almost surely and
for 0 < ε ≤ 1 0<\varepsilon\le1 0 < ε ≤ 1 ,
∣ X + ε G ∣ k ≤ 2 max { k − 1 , 0 } ( ∣ X ∣ k + ∣ G ∣ k ) . |X+\sqrt\varepsilon G|^k\le
2^{\max\{k-1,0\}}(|X|^k+|G|^k). ∣ X + ε G ∣ k ≤ 2 m a x { k − 1 , 0 } ( ∣ X ∣ k + ∣ G ∣ k ) . Dominated convergence proves weak convergence, convergence of every
polynomial moment, and convergence of all absolute moments. The extra scale
tends to one and preserves these conclusions.
Isotropic approximants. For isotropic X X X put
ε j = δ j = 1 / j \varepsilon_j=\delta_j=1/j ε j = δ j = 1/ j , X j = X + ε j G X_j=X+\sqrt{\varepsilon_j}G X j = X + ε j G , and define
d ν j ( x ) = Z j − 1 e − δ j ∣ x ∣ 2 / 2 ρ ε j ( x ) d x , Z j = E e − δ j ∣ X j ∣ 2 / 2 . d\nu_j(x)=Z_j^{-1}e^{-\delta_j|x|^2/2}\rho_{\varepsilon_j}(x)\,dx,
\qquad Z_j=\mathbb E e^{-\delta_j|X_j|^2/2}. d ν j ( x ) = Z j − 1 e − δ j ∣ x ∣ 2 /2 ρ ε j ( x ) d x , Z j = E e − δ j ∣ X j ∣ 2 /2 . Its smooth potential W j W_j W j has
δ j I ⪯ D 2 W j ⪯ ( δ j + ε j − 1 ) I \delta_jI\preceq D^2W_j\preceq(\delta_j+\varepsilon_j^{-1})I δ j I ⪯ D 2 W j ⪯ ( δ j + ε j − 1 ) I .
The same coupling gives Z j → 1 Z_j\to1 Z j → 1 and weak and all fixed-moment convergence
of ν j \nu_j ν j to μ \mu μ by dominated convergence. In particular its mean m j m_j m j
and covariance S j S_j S j satisfy m j → 0 m_j\to0 m j → 0 , S j → I S_j\to I S j → I . Every S j S_j S j is
positive definite because the density of ν j \nu_j ν j is strictly positive.
Let μ j \mu_j μ j be its image under x ↦ S j − 1 / 2 ( x − m j ) x\mapsto S_j^{-1/2}(x-m_j) x ↦ S j − 1/2 ( x − m j ) . This is
isotropic, and its potential is W j ( m j + S j 1 / 2 y ) W_j(m_j+S_j^{1/2}y) W j ( m j + S j 1/2 y ) up to a constant.
Its Hessian therefore lies between
δ j λ min ( S j ) I \delta_j\lambda_{\min}(S_j)I δ j λ m i n ( S j ) I and
( δ j + ε j − 1 ) λ max ( S j ) I (\delta_j+\varepsilon_j^{-1})\lambda_{\max}(S_j)I ( δ j + ε j − 1 ) λ m a x ( S j ) I .
These bounds are strictly positive and finite, so μ j \mu_j μ j is regular.
Since S j − 1 / 2 → I S_j^{-1/2}\to I S j − 1/2 → I and m j → 0 m_j\to0 m j → 0 , the coupling proves weak convergence.
For each fixed k > 0 k>0 k > 0 , the quantities
∣ S j − 1 / 2 ( X j − m j ) ∣ k |S_j^{-1/2}(X_j-m_j)|^k ∣ S j − 1/2 ( X j − m j ) ∣ k are bounded by an integrable constant multiple
of 1 + ∣ X ∣ k + ∣ G ∣ k 1+|X|^k+|G|^k 1 + ∣ X ∣ k + ∣ G ∣ k , uniformly in j j j . Thus every fixed moment converges too.
Appell limits. In a fixed dimension, each coefficient of the degree-d d d
Appell tensor is a polynomial in moments of orders at most d d d : invert the
formal moment generating series, whose constant coefficient is one, recursively
by degree. Hence moment convergence implies coefficientwise convergence of
these polynomials and all their spatial derivatives. Integrating any product
of two such polynomials uses only finitely many further moments, all
convergent here. In an orthonormal basis of Sym d R n \operatorname{Sym}^d\mathbb R^n Sym d R n ,
the Gram matrix of the normalized Appell polynomials therefore converges
entrywise. Its largest eigenvalue is c d ( μ j ) 2 c_d(\mu_j)^2 c d ( μ j ) 2 , so continuity of the
largest eigenvalue proves convergence of the coefficient norms. The same
finite expansion proves convergence of squared derivative norms for any
fixed tensor and derivative order, including mixed Gram entries.
Weak limits and test functions. For u ∈ C c ∞ u\in C_c^\infty u ∈ C c ∞ , the functions
u , u 2 , ∣ ∇ u ∣ 2 u,u^2,|\nabla u|^2 u , u 2 , ∣∇ u ∣ 2 are bounded and continuous. Weak convergence therefore
passes Var ν j u ≤ K ∫ ∣ ∇ u ∣ 2 d ν j \operatorname{Var}_{\nu_j}u\le K\int|\nabla u|^2d\nu_j Var ν j u ≤ K ∫ ∣∇ u ∣ 2 d ν j to ν \nu ν .
A full-dimensional log-concave law is absolutely continuous, so it remains
to justify the extension from smooth compact tests for an absolutely
continuous probability measure.
First let u u u be compactly supported Lipschitz. Its standard mollifications
u ε u_\varepsilon u ε converge uniformly to u u u , their gradients are uniformly
bounded by its Lipschitz constant, and those gradients converge Lebesgue
almost everywhere to ∇ u \nabla u ∇ u by differentiation of convolutions.
Absolute continuity and dominated convergence pass the inequality to u u u .
Next let h h h be bounded and locally Lipschitz with finite energy. Choose
0 ≤ χ ≤ 1 0\le\chi\le1 0 ≤ χ ≤ 1 smooth, compactly supported, and equal to one on the unit
ball, and set h R ( x ) = χ ( x / R ) h ( x ) h_R(x)=\chi(x/R)h(x) h R ( x ) = χ ( x / R ) h ( x ) . These are compactly supported Lipschitz;
h R → h h_R\to h h R → h in L 2 ( ν ) L^2(\nu) L 2 ( ν ) and
∇ h R − ∇ h = ( χ ( x / R ) − 1 ) ∇ h + R − 1 h ∇ χ ( x / R ) ⟶ 0 in L 2 ( ν ) . \nabla h_R-\nabla h=(\chi(x/R)-1)\nabla h
+R^{-1}h\nabla\chi(x/R)\longrightarrow0\quad\text{in }L^2(\nu). ∇ h R − ∇ h = ( χ ( x / R ) − 1 ) ∇ h + R − 1 h ∇ χ ( x / R ) ⟶ 0 in L 2 ( ν ) . The first term converges by dominated convergence and the second has norm
at most R − 1 ∥ h ∥ ∞ ∥ ∇ χ ∥ ∞ R^{-1}\|h\|_\infty\|\nabla\chi\|_\infty R − 1 ∥ h ∥ ∞ ∥∇ χ ∥ ∞ .
Thus the inequality holds for such h h h .
Finally let f f f be locally Lipschitz with energy
E = ∫ ∣ ∇ f ∣ 2 d ν < ∞ E=\int|\nabla f|^2d\nu<\infty E = ∫ ∣∇ f ∣ 2 d ν < ∞ , without assuming integrability of f f f .
Its truncations f N = max ( − N , min ( f , N ) ) f_N=\max(-N,\min(f,N)) f N = max ( − N , min ( f , N )) satisfy
Var ν f N ≤ K E \operatorname{Var}_\nu f_N\le KE Var ν f N ≤ K E by the chain rule. Choose b < ∞ b<\infty b < ∞
such that A = { ∣ f ∣ ≤ b } A=\{|f|\le b\} A = { ∣ f ∣ ≤ b } has probability p > 0 p>0 p > 0 . For N ≥ b N\ge b N ≥ b write
m N = ∫ f N d ν m_N=\int f_Nd\nu m N = ∫ f N d ν . Since f N = f f_N=f f N = f on A A A ,
p ( ∣ m N ∣ − b ) + 2 ≤ ∫ A ∣ f N − m N ∣ 2 d ν ≤ K E . p(|m_N|-b)_+^2\le\int_A|f_N-m_N|^2d\nu\le KE. p ( ∣ m N ∣ − b ) + 2 ≤ ∫ A ∣ f N − m N ∣ 2 d ν ≤ K E . Consequently
∫ f N 2 d ν ≤ K E + ( b + K E / p ) 2 . \int f_N^2d\nu\le KE+(b+\sqrt{KE/p})^2. ∫ f N 2 d ν ≤ K E + ( b + K E / p ) 2 . Fatou proves f ∈ L 2 ( ν ) f\in L^2(\nu) f ∈ L 2 ( ν ) . Then f N → f f_N\to f f N → f in L 2 ( ν ) L^2(\nu) L 2 ( ν ) , because
∣ f N − f ∣ ≤ ∣ f ∣ |f_N-f|\le|f| ∣ f N − f ∣ ≤ ∣ f ∣ , and passing the variance to the limit gives
Var ν f ≤ K E \operatorname{Var}_\nu f\le KE Var ν f ≤ K E as required.