We first establish a posterior inequality. Let Ψ : ( 0 , ∞ ) → [ 1 , ∞ ) \Psi:(0,\infty)\to[1,\infty) Ψ : ( 0 , ∞ ) → [ 1 , ∞ )
be continuous, and assume the regular coefficient bounds
c k ≤ Ψ ( a − 1 ) ( k − 1 ) / 2 c_k\le\Psi(a^{-1})^{(k-1)/2} c k ≤ Ψ ( a − 1 ) ( k − 1 ) /2 . Consider a strongly log-concave law with
covariance A ≻ 0 A\succ0 A ≻ 0 and density proportional to
exp ( − x T Λ x / 2 − V ( x ) ) \exp(-x^T\Lambda x/2-V(x)) exp ( − x T Λ x /2 − V ( x )) , Λ ≻ 0 \Lambda\succ0 Λ ≻ 0 , V V V extended-valued convex.
If M ⪰ A M\succeq A M ⪰ A and M ⪰ δ Λ − 1 M\succeq\delta\Lambda^{-1} M ⪰ δ Λ − 1 , then the law of
M − 1 / 2 ( X − E X ) M^{-1/2}(X-\mathbb EX) M − 1/2 ( X − E X ) has covariance at most I I I and curvature at least
δ I \delta I δ I . The latter follows by inversion and congruence, without
commuting matrices. Its full Appell expansion gives, for a polynomial f f f
of degree d d d ,
Var f ≤ ∑ k = 1 d J k − 1 ∥ E D k f ∥ M ⊗ k , J = Ψ ( δ − 1 ) . (1) \sqrt{\operatorname{Var}f}\le
\sum_{k=1}^d J^{k-1}\|\mathbb ED^kf\|_{M^{\otimes k}},
\qquad J=\sqrt{\Psi(\delta^{-1})}. \tag{1} Var f ≤ k = 1 ∑ d J k − 1 ∥ E D k f ∥ M ⊗ k , J = Ψ ( δ − 1 ) . ( 1 ) Here the factor 1 / k ! 1/k! 1/ k ! is incorporated in c k c_k c k . For a finite vector or
tensor output use the same inequality in the output Hilbert direct sum and
Minkowski’s inequality.
For a nonsmooth normalized posterior, convolve with N ( 0 , t I ) N(0,tI) N ( 0 , t I ) and divide
by 1 + t \sqrt{1+t} 1 + t . Its covariance is at most I I I and its Hessian lies between
δ ( 1 + t ) / ( 1 + δ t ) I \delta(1+t)/(1+\delta t)I δ ( 1 + t ) / ( 1 + δ t ) I and ( 1 + t ) t − 1 I (1+t)t^{-1}I ( 1 + t ) t − 1 I . The Hessian formula is
t − 1 I − t − 2 Cov ( X ∣ X + t G = y ) t^{-1}I-t^{-2}\operatorname{Cov}(X\mid X+\sqrt tG=y) t − 1 I − t − 2 Cov ( X ∣ X + t G = y ) before the final
rescaling; Brascamp–Lieb bounds the conditional covariance by
( δ + t − 1 ) − 1 I (\delta+t^{-1})^{-1}I ( δ + t − 1 ) − 1 I . These are the same regularization and nonsmooth
Brascamp–Lieb inputs used in the certified analytic and polynomial dossiers.
Fixed-order moments and Appell coefficients converge as t ↓ 0 t\downarrow0 t ↓ 0 .
Continuity of Ψ \Psi Ψ therefore proves (1) for the original posterior.
Start covariance-adapted localization from a compactly supported isotropic
law, using the construction in Theorem 7.1 . Denote its
covariance by A t A_t A t , precision by Λ t = ∫ 0 t A u − 1 d u \Lambda_t=\int_0^tA_u^{-1}du Λ t = ∫ 0 t A u − 1 d u , covariance
noise coefficients by U i , t U_{i,t} U i , t , and whitened noise matrices by
S i , t = A t − 1 / 2 U i , t A t − 1 / 2 S_{i,t}=A_t^{-1/2}U_{i,t}A_t^{-1/2} S i , t = A t − 1/2 U i , t A t − 1/2 . Its identities give
d A t = ∑ i U i , t d β i , t − A t d t , ∑ i S i , t 2 ⪯ 8 I . dA_t=\sum_iU_{i,t}\,d\beta_{i,t}-A_tdt,
\qquad \sum_iS_{i,t}^2\preceq8I. d A t = i ∑ U i , t d β i , t − A t d t , i ∑ S i , t 2 ⪯ 8 I . Fix τ , s > 0 \tau,s>0 τ , s > 0 , put ρ = s / τ \rho=s/\tau ρ = s / τ , and introduce
M t = A t + ρ ∫ 0 t A u d u ⪰ A t . M_t=A_t+\rho\int_0^tA_u\,du\succeq A_t. M t = A t + ρ ∫ 0 t A u d u ⪰ A t . For h j = E t D j f h_j=\mathbb E_tD^jf h j = E t D j f set
N j = E l o c ∥ h j ∥ M t ⊗ j 2 , L j = E l o c E t ∥ D j f − h j ∥ M t ⊗ j 2 . N_j=\mathbb E_{\rm loc}\|h_j\|_{M_t^{\otimes j}}^2,
\qquad L_j=\mathbb E_{\rm loc}\mathbb E_t
\|D^jf-h_j\|_{M_t^{\otimes j}}^2. N j = E loc ∥ h j ∥ M t ⊗ j 2 , L j = E loc E t ∥ D j f − h j ∥ M t ⊗ j 2 . We claim
N j ′ ≤ ( 4 j 2 + j ρ ) N j + 9 j 2 L j . (2) N_j'\le(4j^2+j\rho)N_j+9j^2L_j. \tag{2} N j ′ ≤ ( 4 j 2 + j ρ ) N j + 9 j 2 L j . ( 2 ) Indeed, the drift of M t M_t M t is ( ρ − 1 ) A t ⪯ ρ M t (\rho-1)A_t\preceq\rho M_t ( ρ − 1 ) A t ⪯ ρ M t .
With C = M t − 1 / 2 A t 1 / 2 C=M_t^{-1/2}A_t^{1/2} C = M t − 1/2 A t 1/2 , both C C T CC^T C C T and C T C C^TC C T C are contractions.
Thus its whitened noise matrices S ~ i = C S i C T \widetilde S_i=CS_iC^T S i = C S i C T obey
∑ i S ~ i 2 = C ( ∑ i S i C T C S i ) C T ⪯ 8 I . \sum_i\widetilde S_i^2
=C\Big(\sum_iS_iC^TCS_i\Big)C^T\preceq8I. i ∑ S i 2 = C ( i ∑ S i C T C S i ) C T ⪯ 8 I . For K = M t ⊗ j K=M_t^{\otimes j} K = M t ⊗ j the product rule bounds the drift by
[ j ρ + 8 ( j 2 ) ] K [j\rho+8\binom j2]K [ j ρ + 8 ( 2 j ) ] K . To see the bound for each pair of slots, their
commuting symmetric matrices satisfy
2 S ~ i ( h ) S ~ i ( l ) ⪯ ( S ~ i ( h ) ) 2 + ( S ~ i ( l ) ) 2 2\widetilde S_i^{(h)}\widetilde S_i^{(l)}
\preceq(\widetilde S_i^{(h)})^2+(\widetilde S_i^{(l)})^2 2 S i ( h ) S i ( l ) ⪯ ( S i ( h ) ) 2 + ( S i ( l ) ) 2 ;
sum over i i i . After congruence the noise of K K K is
V i = ∑ h = 1 j S ~ i ( h ) V_i=\sum_{h=1}^j\widetilde S_i^{(h)} V i = ∑ h = 1 j S i ( h ) , whence
∑ i V i 2 ⪯ 8 j 2 I \sum_iV_i^2\preceq8j^2I ∑ i V i 2 ⪯ 8 j 2 I .
The noise of h j h_j h j is
c i = E t [ ( D j f − h j ) ξ i ] c_i=\mathbb E_t[(D^jf-h_j)\xi_i] c i = E t [( D j f − h j ) ξ i ] , where ξ = A t − 1 / 2 ( X − m t ) \xi=A_t^{-1/2}(X-m_t) ξ = A t − 1/2 ( X − m t ) .
Weighted Bessel gives
Q j : = ∑ i ∥ c i ∥ K 2 ≤ E t ∥ D j f − h j ∥ K 2 Q_j:=\sum_i\|c_i\|_K^2\le\mathbb E_t\|D^jf-h_j\|_K^2 Q j := ∑ i ∥ c i ∥ K 2 ≤ E t ∥ D j f − h j ∥ K 2 .
The cross variation is bounded by
2 ∑ i ⟨ K 1 / 2 h j , V i K 1 / 2 c i ⟩ ≤ ∥ h j ∥ K 2 + 8 j 2 Q j . 2\sum_i\langle K^{1/2}h_j,V_iK^{1/2}c_i\rangle
\le\|h_j\|_K^2+8j^2Q_j. 2 i ∑ ⟨ K 1/2 h j , V i K 1/2 c i ⟩ ≤ ∥ h j ∥ K 2 + 8 j 2 Q j . Adding the quadratic variation of h j h_j h j proves (2), because
8 ( j 2 ) + 1 ≤ 4 j 2 8\binom j2+1\le4j^2 8 ( 2 j ) + 1 ≤ 4 j 2 and 1 + 8 j 2 ≤ 9 j 2 1+8j^2\le9j^2 1 + 8 j 2 ≤ 9 j 2 .
One may first stop the localization parameters on compact sets. Spatial
compactness bounds the covariance, polynomial tests and accumulated metric
on each deterministic time interval. The displayed Bessel and noise bounds
then remove these stops in the integrated inequalities.
Suppose inductively that all smaller degrees in the dimension range satisfy
c k ≤ A k − 1 c_k\le A^{k-1} c k ≤ A k − 1 , with A ≥ 1 A\ge1 A ≥ 1 . Applying the Appell expansion to each
D j f D^jf D j f in a whitened posterior, and increasing each newly introduced
covariance slot from A t A_t A t to M t M_t M t , yields
L j ≤ ∑ k = 1 d − j A k − 1 N j + k . (3) \sqrt{L_j}\le\sum_{k=1}^{d-j}A^{k-1}\sqrt{N_{j+k}}. \tag{3} L j ≤ k = 1 ∑ d − j A k − 1 N j + k . ( 3 ) All old output slots keep the weight M t M_t M t . Thus (3) uses no comparison
between metrics at different times.
Take f = P d [ T ] f=P_d[T] f = P d [ T ] , T ≠ 0 T\ne0 T = 0 , put Q = ( d ! ) 2 ∥ T ∥ H S 2 Q=(d!)^2\|T\|_{\rm HS}^2 Q = ( d ! ) 2 ∥ T ∥ HS 2 ,
w l = A 2 max ( l − 1 , 0 ) w_l=A^{2\max(l-1,0)} w l = A 2 m a x ( l − 1 , 0 ) and
E ( t ) = max 1 ≤ j ≤ d N j ( t ) / ( Q w d − j ) E(t)=\max_{1\le j\le d}N_j(t)/(Qw_{d-j}) E ( t ) = max 1 ≤ j ≤ d N j ( t ) / ( Q w d − j ) .
For 1 ≤ k ≤ l 1\le k\le l 1 ≤ k ≤ l one has
A k − 1 w l − k ≤ w l A^{k-1}\sqrt{w_{l-k}}\le\sqrt{w_l} A k − 1 w l − k ≤ w l . Hence
L j ≤ d 2 Q w d − j E L_j\le d^2Qw_{d-j}E L j ≤ d 2 Q w d − j E . Appell centering gives
N j ( 0 ) = 0 N_j(0)=0 N j ( 0 ) = 0 for j < d j<d j < d and N d ( 0 ) = Q N_d(0)=Q N d ( 0 ) = Q . Integrating (2), applying the
maximum, and using Gronwall therefore gives
E ( t ) ≤ e ( 13 d 4 + d ρ ) t . E(t)\le e^{(13d^4+d\rho)t}. E ( t ) ≤ e ( 13 d 4 + d ρ ) t . Keeping the zero initial conditions in variation of constants gives the
additional estimates
N j ( t ) ≤ 9 d 4 t e ( 13 d 4 + d ρ ) t Q w d − j ( j < d ) , N d ( t ) ≤ Q e ( 4 d 2 + d ρ ) t . (4) N_j(t)\le9d^4t e^{(13d^4+d\rho)t}Qw_{d-j}\quad(j<d),
\qquad N_d(t)\le Qe^{(4d^2+d\rho)t}. \tag{4} N j ( t ) ≤ 9 d 4 t e ( 13 d 4 + d ρ ) t Q w d − j ( j < d ) , N d ( t ) ≤ Q e ( 4 d 2 + d ρ ) t . ( 4 ) For 0 < ε < 1 0<\varepsilon<1 0 < ε < 1 take
τ = ε / d 6 , s = ε / d , δ = s τ = ε 2 / d 7 . \tau=\varepsilon/d^6,\qquad s=\varepsilon/d,
\qquad\delta=s\tau=\varepsilon^2/d^7. τ = ε / d 6 , s = ε / d , δ = s τ = ε 2 / d 7 . For every vector v v v , expansion of a square gives
∫ 0 τ v T A t v d t − τ 2 v T Λ τ − 1 v = ∫ 0 τ ∣ A t 1 / 2 v − τ A t − 1 / 2 Λ τ − 1 v ∣ 2 d t ≥ 0. \int_0^\tau v^TA_tv\,dt-\tau^2v^T\Lambda_\tau^{-1}v
=\int_0^\tau|A_t^{1/2}v-\tau A_t^{-1/2}\Lambda_\tau^{-1}v|^2dt\ge0. ∫ 0 τ v T A t v d t − τ 2 v T Λ τ − 1 v = ∫ 0 τ ∣ A t 1/2 v − τ A t − 1/2 Λ τ − 1 v ∣ 2 d t ≥ 0. Thus M τ ⪰ δ Λ τ − 1 M_\tau\succeq\delta\Lambda_\tau^{-1} M τ ⪰ δ Λ τ − 1 , and (1) applies at time
τ \tau τ with J = Ψ ( d 7 / ε 2 ) J=\sqrt{\Psi(d^7/\varepsilon^2)} J = Ψ ( d 7 / ε 2 ) .
The exponents in (4) are bounded respectively by 14 ε 14\varepsilon 14 ε and
5 ε 5\varepsilon 5 ε . When J ≤ A J\le A J ≤ A , the top term in (1) contributes at most
Q e 5 ε / 2 J d − 1 \sqrt Qe^{5\varepsilon/2}J^{d-1} Q e 5 ε /2 J d − 1 . Each of the at most d − 1 d-1 d − 1 lower
terms contributes at most
Q 3 d 2 τ e 7 ε A d − 2 \sqrt Q3d^2\sqrt\tau e^{7\varepsilon}A^{d-2} Q 3 d 2 τ e 7 ε A d − 2 , including k = d − 1 k=d-1 k = d − 1
since w 1 = 1 w_1=1 w 1 = 1 . Minkowski over localization randomness proves
E l o c Var τ f ≤ Q ( e 5 ε / 2 J d − 1 + 3 d 3 τ e 7 ε A d − 2 ) . \sqrt{\mathbb E_{\rm loc}\operatorname{Var}_\tau f}
\le\sqrt Q\big(e^{5\varepsilon/2}J^{d-1}
+3d^3\sqrt\tau e^{7\varepsilon}A^{d-2}\big). E loc Var τ f ≤ Q ( e 5 ε /2 J d − 1 + 3 d 3 τ e 7 ε A d − 2 ) . Bessel and the posterior martingale identity imply
( E l o c Var t f ) ′ = − E l o c ∣ Cov t ( f , ξ t ) ∣ 2 ≥ − E l o c Var t f (\mathbb E_{\rm loc}\operatorname{Var}_t f)'
=-\mathbb E_{\rm loc}|\operatorname{Cov}_t(f,\xi_t)|^2
\ge-\mathbb E_{\rm loc}\operatorname{Var}_t f ( E loc Var t f ) ′ = − E loc ∣ Cov t ( f , ξ t ) ∣ 2 ≥ − E loc Var t f .
Since τ ≤ ε \tau\le\varepsilon τ ≤ ε , the finite-degree bound is consequently
c d ≤ e 3 ε J d − 1 + 3 ε e 8 ε A d − 2 . (5) c_d\le e^{3\varepsilon}J^{d-1}
+3\sqrt\varepsilon e^{8\varepsilon}A^{d-2}. \tag{5} c d ≤ e 3 ε J d − 1 + 3 ε e 8 ε A d − 2 . ( 5 ) Apply this with Ψ ( x ) = Γ 2 ℓ r ( x ) 2 \Psi(x)=\Gamma^2\ell_r(x)^2 Ψ ( x ) = Γ 2 ℓ r ( x ) 2 and, simultaneously over all
laws, induct on d d d . Degree one is covariance normalization. For d ≥ 2 d\ge2 d ≥ 2
set α = r − 2 \alpha=r^{-2} α = r − 2 , ε = 2 − 20 α 2 \varepsilon=2^{-20}\alpha^2 ε = 2 − 20 α 2 , and
A = ( 1 + α ) Γ ℓ r ( d ) A=(1+\alpha)\Gamma\ell_r(d) A = ( 1 + α ) Γ ℓ r ( d ) . Monotonicity makes the smaller-degree
induction sufficient for (3). The elementary inequalities
x g ′ ( x ) g ( x ) ≤ 1 2 , g ( d 7 / ε 2 ) ≤ [ 7 + 2 log ( 1 / ε ) ] g ( d ) \frac{xg'(x)}{g(x)}\le\frac12,
\qquad g(d^7/\varepsilon^2)\le[7+2\log(1/\varepsilon)]g(d) g ( x ) x g ′ ( x ) ≤ 2 1 , g ( d 7 / ε 2 ) ≤ [ 7 + 2 log ( 1/ ε )] g ( d ) imply
log ℓ r ( d 7 / ε 2 ) ℓ r ( d ) ≤ 2 − ( r − 1 ) log [ 7 + 2 log ( 1 / ε ) ] ≤ α / 8 (6) \log\frac{\ell_r(d^7/\varepsilon^2)}{\ell_r(d)}
\le2^{-(r-1)}\log[7+2\log(1/\varepsilon)]\le\alpha/8 \tag{6} log ℓ r ( d ) ℓ r ( d 7 / ε 2 ) ≤ 2 − ( r − 1 ) log [ 7 + 2 log ( 1/ ε )] ≤ α /8 ( 6 ) above one universal r ∗ r_* r ∗ . For the first inequality, with t = e + x t=e+x t = e + x ,
t log t − 2 t + 2 e ≥ 0 t\log t-2t+2e\ge0 t log t − 2 t + 2 e ≥ 0 follows from its nonnegative derivative on [ e , ∞ ) [e,\infty) [ e , ∞ ) ;
integrate the logarithmic derivative and compose. The second follows from
e + d 7 / ε 2 ≤ ( e + d ) 7 / ε 2 e+d^7/\varepsilon^2\le(e+d)^7/\varepsilon^2 e + d 7 / ε 2 ≤ ( e + d ) 7 / ε 2 .
The last inequality in (6) holds uniformly in d d d because
r 2 2 − r log [ 7 + 2 log ( 2 20 r 4 ) ] → 0 r^22^{-r}\log[7+2\log(2^{20}r^4)]\to0 r 2 2 − r log [ 7 + 2 log ( 2 20 r 4 )] → 0 .
As log ( 1 + α / 4 ) ≥ α / 8 \log(1+\alpha/4)\ge\alpha/8 log ( 1 + α /4 ) ≥ α /8 and
log ( 1 + α ) − log ( 1 + α / 4 ) ≥ 3 α / 8 \log(1+\alpha)-\log(1+\alpha/4)\ge3\alpha/8 log ( 1 + α ) − log ( 1 + α /4 ) ≥ 3 α /8 , we have
J / A ≤ e − 3 α / 8 J/A\le e^{-3\alpha/8} J / A ≤ e − 3 α /8 . Divide (5) by A d − 1 A^{d-1} A d − 1 .
Its first term is at most e 3 ε − 3 α / 8 ≤ e − α / 4 ≤ 1 − α / 8 e^{3\varepsilon-3\alpha/8}
\le e^{-\alpha/4}\le1-\alpha/8 e 3 ε − 3 α /8 ≤ e − α /4 ≤ 1 − α /8 .
Its second is at most 3 ε e 8 ε / A ≤ α / 128 3\sqrt\varepsilon e^{8\varepsilon}/A
\le\alpha/128 3 ε e 8 ε / A ≤ α /128 ; here A ≥ 1 A\ge1 A ≥ 1 and the stated choice of ε \varepsilon ε
suffices. Their sum is less than one. This closes the degree induction
starting at degree two, with no depth-dependent small-degree threshold.
Finally condition an arbitrary isotropic law on increasing centered balls,
then center and whiten. Log-concavity gives convergence of moments of every
fixed order; the recursively determined Appell coefficients converge as
well. At each fixed degree the proved inequality therefore passes to the
limit. For covariance at most I I I , whitening and contraction in every tensor
slot transfer the same bounds; a singular covariance is treated on its
supporting subspace. Those subspaces remain in the specified dimension
range. This completes the proof.