Song–Zhang v1: the iteration of curvature profiles
Part of the first version of Song–Zhang, Chapter Song–Zhang, first version: polynomial estimates and curvature ; the reading order is on the full proofs page.
Overview. This reconstructs Section 6 of Song & Zhang, 2026 .
First a curvature bound is extended to affine normalizations of possibly
nonsmooth localization posteriors. Next the derivative hierarchy gives new
polynomial coefficients. Retaining its zero initial data makes the additional
coefficient loss tend to one at large depth. Finally the comparison
Theorem 7.2 closes a finite-depth induction, with every
initialization and admissibility threshold paid explicitly.
Put g ( x ) = log ( e + x ) g(x)=\log(e+x) g ( x ) = log ( e + x ) , ℓ 0 ( x ) = x \ell_0(x)=x ℓ 0 ( x ) = x and ℓ r = g ∘ r \ell_r=g^{\circ r} ℓ r = g ∘ r .
There exist a universal C 0 > 0 C_0>0 C 0 > 0 and universal constants Γ r ≤ C 0 4 r \Gamma_r\le C_0 4^r Γ r ≤ C 0 4 r
such that, in every dimension, every centered probability measure
ν = e − W d x \nu=e^{-W}dx ν = e − W d x with smooth W W W , covariance at most I I I , and
a I ⪯ D 2 W ⪯ b I aI\preceq D^2W\preceq bI a I ⪯ D 2 W ⪯ b I for some 0 < a ≤ b < ∞ 0<a\le b<\infty 0 < a ≤ b < ∞ , satisfies
C P ( ν ) ≤ Γ r 2 ℓ r ( a − 1 ) 2 ( r ≥ 1 ) . C_P(\nu)\le\Gamma_r^2\ell_r(a^{-1})^2\qquad(r\ge1). C P ( ν ) ≤ Γ r 2 ℓ r ( a − 1 ) 2 ( r ≥ 1 ) . This is Theorem 7.3 .
Extending a regular curvature profile ¶ Suppose a continuous F : ( 0 , ∞ ) → ( 0 , ∞ ) F:(0,\infty)\to(0,\infty) F : ( 0 , ∞ ) → ( 0 , ∞ ) satisfies
C P ( η ) ≤ F ( a ) C_P(\eta)\le F(a) C P ( η ) ≤ F ( a ) for every centered regular measure with covariance at most
I I I and curvature at least a I aI a I . If ν \nu ν has positive covariance A A A and
density proportional to exp ( − x T B x / 2 − V ( x ) ) \exp(-x^TBx/2-V(x)) exp ( − x T B x /2 − V ( x )) , with B ≻ 0 B\succ0 B ≻ 0 and
extended-valued convex V V V , then, for every δ > 0 \delta>0 δ > 0 and every locally
Lipschitz finite-energy f f f ,
Var ν f ≤ F ( δ ) E ν [ ∇ f T ( A + δ B − 1 ) ∇ f ] . \operatorname{Var}_\nu f\le F(\delta)
\mathbb E_\nu[\nabla f^T(A+\delta B^{-1})\nabla f]. Var ν f ≤ F ( δ ) E ν [ ∇ f T ( A + δ B − 1 ) ∇ f ] . First take a centered, possibly nonsmooth measure η \eta η of covariance at most
I I I and curvature at least a I aI a I . Convolving with N ( 0 , t I ) N(0,tI) N ( 0 , t I ) gives a smooth
positive density with potential Hessian
t − 1 I − t − 2 Cov ( X ∣ X + t Z = y ) t^{-1}I-t^{-2}\operatorname{Cov}(X\mid X+\sqrt tZ=y) t − 1 I − t − 2 Cov ( X ∣ X + t Z = y ) .
The conditional law has curvature at least ( a + t − 1 ) I (a+t^{-1})I ( a + t − 1 ) I ; the constant-curvature
Brascamp--Lieb inequality Brascamp & Lieb, 1976 Bakry et al. , 2014, Theorem 4.9.1 bounds this conditional covariance
by ( a + t − 1 ) − 1 I (a+t^{-1})^{-1}I ( a + t − 1 ) − 1 I . Thus the Hessian lies between
a / ( 1 + a t ) I a/(1+at)I a / ( 1 + a t ) I and t − 1 I t^{-1}I t − 1 I . This use of Brascamp--Lieb on an extended-valued
convex potential is the same nonsmooth form used in the polynomial proof.
Dividing the convolved vector by 1 + t \sqrt{1+t} 1 + t makes its covariance at most I I I
and curvature at least a t = a ( 1 + t ) / ( 1 + a t ) a_t=a(1+t)/(1+at) a t = a ( 1 + t ) / ( 1 + a t ) . The regular profile gives
C P ≤ F ( a t ) C_P\le F(a_t) C P ≤ F ( a t ) , and F ( a t ) → F ( a ) F(a_t)\to F(a) F ( a t ) → F ( a ) . Weak convergence passes the inequality
to compact smooth tests. The clipping and cutoff argument of
Lemma 7.1 extends it to every finite-energy locally
Lipschitz test. To handle varying constants directly, use the common bound
F ( a ) + ζ F(a)+\zeta F ( a ) + ζ for all sufficiently small t t t , then let ζ ↓ 0 \zeta\downarrow0 ζ ↓ 0 .
Now put M = A + δ B − 1 M=A+\delta B^{-1} M = A + δ B − 1 and Y = M − 1 / 2 ( X − E X ) Y=M^{-1/2}(X-\mathbb EX) Y = M − 1/2 ( X − E X ) .
Its covariance is M − 1 / 2 A M − 1 / 2 ⪯ I M^{-1/2}AM^{-1/2}\preceq I M − 1/2 A M − 1/2 ⪯ I . Since
M ⪰ δ B − 1 M\succeq\delta B^{-1} M ⪰ δ B − 1 , inversion gives B ⪰ δ M − 1 B\succeq\delta M^{-1} B ⪰ δ M − 1 ,
so the potential of Y Y Y has curvature at least δ I \delta I δ I .
Applying the extended profile to
y ↦ f ( E X + M 1 / 2 y ) y\mapsto f(\mathbb EX+M^{1/2}y) y ↦ f ( E X + M 1/2 y ) gives the stated gradient form.
No matrices were commuted. Strong convexity supplies Gaussian tails, so
polynomial tests have finite energy as well.
We record elementary estimates used uniformly in the depth. Concavity and
g ( 0 ) = 1 g(0)=1 g ( 0 ) = 1 imply g ( c x ) ≤ c g ( x ) g(cx)\le c g(x) g ( c x ) ≤ c g ( x ) for c ≥ 1 c\ge1 c ≥ 1 ; induction gives
ℓ r ( x ) ≥ 1 , ℓ r ( c x ) ≤ c ℓ r ( x ) , ℓ r ( 2308 d 2 ) ≤ 10 ℓ r ( d ) ( r , d ≥ 1 ) . (1) \ell_r(x)\ge1,\quad \ell_r(cx)\le c\ell_r(x),\quad
\ell_r(2308d^2)\le10\ell_r(d)\quad(r,d\ge1). \tag{1} ℓ r ( x ) ≥ 1 , ℓ r ( c x ) ≤ c ℓ r ( x ) , ℓ r ( 2308 d 2 ) ≤ 10 ℓ r ( d ) ( r , d ≥ 1 ) . ( 1 ) For the last assertion,
e + 2308 d 2 ≤ 2309 ( e + d ) 2 e+2308d^2\le2309(e+d)^2 e + 2308 d 2 ≤ 2309 ( e + d ) 2 and log 2309 < 8 \log2309<8 log 2309 < 8 give
g ( 2308 d 2 ) ≤ 10 g ( d ) g(2308d^2)\le10g(d) g ( 2308 d 2 ) ≤ 10 g ( d ) ; apply the scaling estimate to the remaining compositions.
Furthermore
x g ′ ( x ) g ( x ) ≤ 1 2 ( x > 0 ) . {xg'(x)\over g(x)}\le\tfrac12\quad(x>0). g ( x ) x g ′ ( x ) ≤ 2 1 ( x > 0 ) . Indeed, setting t = e + x t=e+x t = e + x , the required inequality is
t log t − 2 t + 2 e ≥ 0 t\log t-2t+2e\ge0 t log t − 2 t + 2 e ≥ 0 ; its derivative is log t − 1 ≥ 0 \log t-1\ge0 log t − 1 ≥ 0 for t ≥ e t\ge e t ≥ e ,
and its value at e e e is e e e . Integration of this logarithmic derivative and
composition yield
log ℓ r ( u ) ℓ r ( v ) ≤ 2 − r log ( u / v ) ( u ≥ v > 0 , r ≥ 0 ) . (2) \log{\ell_r(u)\over\ell_r(v)}\le2^{-r}\log(u/v)
\quad(u\ge v>0, r\ge0). \tag{2} log ℓ r ( v ) ℓ r ( u ) ≤ 2 − r log ( u / v ) ( u ≥ v > 0 , r ≥ 0 ) . ( 2 ) A coarse coefficient improvement ¶ Suppose for some r ≥ 1 r\ge1 r ≥ 1 and Γ ≥ 1 \Gamma\ge1 Γ ≥ 1 we have the curvature profile
F ( a ) = Γ 2 ℓ r ( a − 1 ) 2 F(a)=\Gamma^2\ell_r(a^{-1})^2 F ( a ) = Γ 2 ℓ r ( a − 1 ) 2 for every regular covariance contraction.
Write ℓ = ℓ r \ell=\ell_r ℓ = ℓ r , b 0 = 1 b_0=1 b 0 = 1 , and b s = ℓ ( s ) s / ( s + 1 ) 2 b_s=\ell(s)^s/(s+1)^2 b s = ℓ ( s ) s / ( s + 1 ) 2 for s ≥ 1 s\ge1 s ≥ 1 .
Monotonicity gives
∑ k = 1 s b k b s − k b s ≤ 16 , b d b d − 1 ≥ ℓ ( d ) d 2 ( d + 1 ) 2 ≥ ℓ ( d ) / 4. (3) \sum_{k=1}^s{b_kb_{s-k}\over b_s}\le16,
\qquad {b_d\over b_{d-1}}\ge\ell(d){d^2\over(d+1)^2}
\ge\ell(d)/4. \tag{3} k = 1 ∑ s b s b k b s − k ≤ 16 , b d − 1 b d ≥ ℓ ( d ) ( d + 1 ) 2 d 2 ≥ ℓ ( d ) /4. ( 3 ) The ratio statement includes d = 1 d=1 d = 1 by the definition of b 0 b_0 b 0 .
For the convolution sum, replace all logarithmic factors by ℓ ( s ) \ell(s) ℓ ( s ) and
split the index set at s / 2 s/2 s /2 . The reciprocal square from the larger index
contributes at most 4 / ( s + 1 ) 2 4/(s+1)^2 4/ ( s + 1 ) 2 ; summing the other reciprocal squares on each
half gives a bound smaller than 16, including the endpoint s − k = 0 s-k=0 s − k = 0 .
We prove, simultaneously for all isotropic log-concave measures,
c d ≤ R d b d , R = 2 12 Γ . (4) c_d\le R^db_d,\qquad R=2^{12}\Gamma. \tag{4} c d ≤ R d b d , R = 2 12 Γ. ( 4 ) Here c d = K d / d ! c_d=\sqrt{K_d}/d! c d = K d / d ! has the Appell normalization of
Theorem 7.1 . Degree one follows from c 1 = 1 c_1=1 c 1 = 1 and R ≥ 4 R\ge4 R ≥ 4 .
For the induction step take a compactly supported isotropic initial law and
f = P d [ T ] f=P_d[T] f = P d [ T ] , T ≠ 0 T\ne0 T = 0 . Use the global covariance-adapted localization
Lemma 120.1 and its derivative hierarchy
Lemma 120.2 , both proved in the polynomial dossier.
Explicitly, A t A_t A t is the posterior covariance, Λ t = ∫ 0 t A s − 1 d s \Lambda_t=\int_0^tA_s^{-1}ds Λ t = ∫ 0 t A s − 1 d s
is its curvature matrix, h j ( t ) = E t D j f h_j(t)=\mathbb E_tD^jf h j ( t ) = E t D j f , and
N j ( t ) = E l o c ⟨ h j , A t ⊗ j h j ⟩ , L j ( t ) = E l o c E t ⟨ D j f − h j , A t ⊗ j ( D j f − h j ) ⟩ . N_j(t)=\mathbb E_{\rm loc}\langle h_j,A_t^{\otimes j}h_j\rangle,
\quad L_j(t)=\mathbb E_{\rm loc}\mathbb E_t
\langle D^jf-h_j,A_t^{\otimes j}(D^jf-h_j)\rangle. N j ( t ) = E loc ⟨ h j , A t ⊗ j h j ⟩ , L j ( t ) = E loc E t ⟨ D j f − h j , A t ⊗ j ( D j f − h j )⟩ . The hierarchy gives N j ′ ≤ 4 j 2 N j + 9 j 2 L j N_j'\le4j^2N_j+9j^2L_j N j ′ ≤ 4 j 2 N j + 9 j 2 L j almost everywhere.
The lower-degree inductive bounds apply to every whitened posterior.
The Appell expansion, followed by the L 2 L^2 L 2 triangle inequality over its full
output tensor direct sum and over localization randomness, gives
L j ( t ) ≤ ∑ k = 1 d − j R k b k N j + k ( t ) . (5) \sqrt{L_j(t)}\le\sum_{k=1}^{d-j}R^kb_k\sqrt{N_{j+k}(t)}. \tag{5} L j ( t ) ≤ k = 1 ∑ d − j R k b k N j + k ( t ) . ( 5 ) Each additional derivative index carries the same covariance factor as the
original derivative indices; this is an affine change of variables, not a
scalar covariance bound.
Put Q = ( d ! ) 2 ∥ T ∥ H S 2 Q=(d!)^2\|T\|_{\rm HS}^2 Q = ( d ! ) 2 ∥ T ∥ HS 2 , w s = R 2 s b s 2 w_s=R^{2s}b_s^2 w s = R 2 s b s 2 , and
M ( t ) = max 1 ≤ j ≤ d N j ( t ) / ( Q w d − j ) M(t)=\max_{1\le j\le d}N_j(t)/(Qw_{d-j}) M ( t ) = max 1 ≤ j ≤ d N j ( t ) / ( Q w d − j ) .
Appell centering implies N d ( 0 ) = Q N_d(0)=Q N d ( 0 ) = Q , N j ( 0 ) = 0 N_j(0)=0 N j ( 0 ) = 0 for j < d j<d j < d , and M ( 0 ) = 1 M(0)=1 M ( 0 ) = 1 .
Equations (3),(5) give L j ≤ 256 Q w d − j M L_j\le256Qw_{d-j}M L j ≤ 256 Q w d − j M , with L d = 0 L_d=0 L d = 0 .
Integrating the hierarchy gives
M ( t ) ≤ 1 + 2308 d 2 ∫ 0 t M ( s ) d s M(t)\le1+2308d^2\int_0^tM(s)ds M ( t ) ≤ 1 + 2308 d 2 ∫ 0 t M ( s ) d s , hence M ( t ) ≤ e 2308 d 2 t M(t)\le e^{2308d^2t} M ( t ) ≤ e 2308 d 2 t .
At τ = 1 / ( 2308 d 2 ) \tau=1/(2308d^2) τ = 1/ ( 2308 d 2 ) , for 0 ≤ s ≤ τ 0\le s\le\tau 0 ≤ s ≤ τ ,
G 1 ( s ) : = E l o c E s [ ∇ f T A s ∇ f ] = N 1 ( s ) + L 1 ( s ) ≤ 257 e Q R 2 d − 2 b d − 1 2 . (6) G_1(s):=\mathbb E_{\rm loc}\mathbb E_s[\nabla f^TA_s\nabla f]
=N_1(s)+L_1(s)\le257eQ R^{2d-2}b_{d-1}^2. \tag{6} G 1 ( s ) := E loc E s [ ∇ f T A s ∇ f ] = N 1 ( s ) + L 1 ( s ) ≤ 257 e Q R 2 d − 2 b d − 1 2 . ( 6 ) Here are the variance and terminal-transfer identities needed in both
coefficient inductions. The localization martingales for f , f 2 f,f^2 f , f 2 give
V ′ ( t ) = − E l o c ∣ Cov t ( f , ξ t ) ∣ 2 ≥ − V ( t ) V'(t)=-\mathbb E_{\rm loc}|\operatorname{Cov}_t(f,\xi_t)|^2\ge-V(t) V ′ ( t ) = − E loc ∣ Cov t ( f , ξ t ) ∣ 2 ≥ − V ( t ) ,
where V ( t ) = E l o c Var t f V(t)=\mathbb E_{\rm loc}\operatorname{Var}_t f V ( t ) = E loc Var t f and
ξ t = A t − 1 / 2 ( X − m t ) \xi_t=A_t^{-1/2}(X-m_t) ξ t = A t − 1/2 ( X − m t ) is isotropic. Bessel’s inequality supplies the last
bound. Therefore V ( τ ) ≥ e − τ Var 0 f V(\tau)\ge e^{-\tau}\operatorname{Var}_0f V ( τ ) ≥ e − τ Var 0 f .
Also
Λ τ − 1 ⪯ τ − 2 ∫ 0 τ A s d s . (7) \Lambda_\tau^{-1}\preceq\tau^{-2}\int_0^\tau A_sds. \tag{7} Λ τ − 1 ⪯ τ − 2 ∫ 0 τ A s d s . ( 7 ) To verify (7) test on v v v and expand
∫ 0 τ ∣ A s 1 / 2 v − τ A s − 1 / 2 Λ τ − 1 v ∣ 2 d s \int_0^\tau|A_s^{1/2}v-\tau A_s^{-1/2}\Lambda_\tau^{-1}v|^2ds ∫ 0 τ ∣ A s 1/2 v − τ A s − 1/2 Λ τ − 1 v ∣ 2 d s .
The result is
∫ 0 τ v T A s v d s − τ 2 v T Λ τ − 1 v \int_0^\tau v^TA_svds-\tau^2v^T\Lambda_\tau^{-1}v ∫ 0 τ v T A s v d s − τ 2 v T Λ τ − 1 v .
For s ≤ τ s\le\tau s ≤ τ , the posterior martingale of each fixed product
∂ i f ∂ j f \partial_if\partial_jf ∂ i f ∂ j f and the F s \mathcal F_s F s -measurability of A s A_s A s give
E l o c E τ [ ∇ f T A s ∇ f ] = G 1 ( s ) . (8) \mathbb E_{\rm loc}\mathbb E_\tau[\nabla f^TA_s\nabla f]=G_1(s). \tag{8} E loc E τ [ ∇ f T A s ∇ f ] = G 1 ( s ) . ( 8 ) All quantities are integrable because the initial support is compact.
The profile-inflation lemma at time τ \tau τ , together with (7),(8), consequently
gives for any δ > 0 \delta>0 δ > 0
Var 0 f ≤ e τ Γ 2 ℓ r ( δ − 1 ) 2 ( G 1 ( τ ) + δ τ − 2 ∫ 0 τ G 1 ( s ) d s ) . (9) \operatorname{Var}_0f\le e^\tau\Gamma^2\ell_r(\delta^{-1})^2
\left(G_1(\tau)+\delta\tau^{-2}\int_0^\tau G_1(s)ds\right). \tag{9} Var 0 f ≤ e τ Γ 2 ℓ r ( δ − 1 ) 2 ( G 1 ( τ ) + δ τ − 2 ∫ 0 τ G 1 ( s ) d s ) . ( 9 ) This identity retains the correlation between the earlier covariance and the
terminal derivative products.
Choose δ = τ \delta=\tau δ = τ . Equations (1),(6),(9) bound the variance by
2 e 2 ⋅ 257 ⋅ 100 Γ 2 ℓ r ( d ) 2 Q R 2 d − 2 b d − 1 2 2e^2\cdot257\cdot100\Gamma^2\ell_r(d)^2 Q R^{2d-2}b_{d-1}^2 2 e 2 ⋅ 257 ⋅ 100 Γ 2 ℓ r ( d ) 2 Q R 2 d − 2 b d − 1 2 .
By (3) this is at most Q R 2 d b d 2 Q R^{2d}b_d^2 Q R 2 d b d 2 because
16 ⋅ 2 e 2 ⋅ 257 ⋅ 100 < 2 24 = R 2 / Γ 2 16\cdot2e^2\cdot257\cdot100<2^{24}=R^2/\Gamma^2 16 ⋅ 2 e 2 ⋅ 257 ⋅ 100 < 2 24 = R 2 / Γ 2
(use e 2 < 8 e^2<8 e 2 < 8 for the strict comparison). This proves (4) at degree d d d for
compact initial laws. Condition any isotropic law on increasing centered balls,
then center and whiten. All fixed moments converge, as do the recursively
specified Appell coefficients, so the degree-d d d inequality passes to the limit.
This completes induction over all measures. Finally, a centered covariance
contraction inherits the estimate from its isotropic whitening, since the map
T ↦ ( Cov ν ) ⊗ d / 2 T T\mapsto(\operatorname{Cov}\nu)^{\otimes d/2}T T ↦ ( Cov ν ) ⊗ d /2 T contracts Hilbert--Schmidt
norm. Singular covariances are treated on their affine support. Thus (4) holds
for every centered covariance contraction as well.
The coefficient improvement with summable additional loss ¶ There exist universal K ≥ 1 K\ge1 K ≥ 1 and an integer r 0 ≥ 2 r_0\ge2 r 0 ≥ 2 such that the following
holds. Suppose r ≥ r 0 r\ge r_0 r ≥ r 0 , Γ ≥ K r 2 \Gamma\ge Kr^2 Γ ≥ K r 2 , and
C P ( ν ) ≤ Γ 2 ℓ r ( a − 1 ) 2 C_P(\nu)\le\Gamma^2\ell_r(a^{-1})^2 C P ( ν ) ≤ Γ 2 ℓ r ( a − 1 ) 2 for every centered regular measure of
covariance at most I I I and curvature at least a I aI a I .
Then all centered log-concave covariance contractions satisfy
c d ≤ R d ℓ r ( d ) d ( d + 1 ) 2 ( d ≥ 1 ) , R = ( 1 + r − 2 ) Γ . c_d\le R^d{\ell_r(d)^d\over(d+1)^2}\quad(d\ge1),
\qquad R=(1+r^{-2})\Gamma. c d ≤ R d ( d + 1 ) 2 ℓ r ( d ) d ( d ≥ 1 ) , R = ( 1 + r − 2 ) Γ. Put α = r − 2 \alpha=r^{-2} α = r − 2 and D r = ⌈ 32 r 2 ⌉ D_r=\lceil32r^2\rceil D r = ⌈ 32 r 2 ⌉ . Use the same b s , w s , Q , N j , L j b_s,w_s,Q,N_j,L_j b s , w s , Q , N j , L j
and induct on degree, starting the hierarchy step only at d ≥ D r d\ge D_r d ≥ D r .
The estimates preceding (6) remain valid with this new R R R , as they use only
the inductive lower-degree bounds. But keeping the zero initial conditions in
the integrated hierarchy improves them to
N j ( t ) ≤ 2304 d 2 t e 2308 d 2 t Q w d − j ( j < d ) , N d ( t ) ≤ Q e 4 d 2 t . (10) N_j(t)\le2304d^2t e^{2308d^2t}Qw_{d-j}\quad(j<d),
\qquad N_d(t)\le Qe^{4d^2t}. \tag{10} N j ( t ) ≤ 2304 d 2 t e 2308 d 2 t Q w d − j ( j < d ) , N d ( t ) ≤ Q e 4 d 2 t . ( 10 ) Indeed variation of constants gives
N j ( t ) ≤ 2304 j 2 Q w d − j ∫ 0 t e 4 j 2 ( t − s ) + 2308 d 2 s d s N_j(t)\le2304j^2Qw_{d-j}\int_0^t
e^{4j^2(t-s)+2308d^2s}ds N j ( t ) ≤ 2304 j 2 Q w d − j ∫ 0 t e 4 j 2 ( t − s ) + 2308 d 2 s d s , whose integrand is at most e 2308 d 2 t e^{2308d^2t} e 2308 d 2 t .
For 0 < η < 1 0<\eta<1 0 < η < 1 , set τ = η / d 2 \tau=\eta/d^2 τ = η / d 2 . In (5) at j = 1 j=1 j = 1 , separate the last term
k = d − 1 k=d-1 k = d − 1 , which contains N d N_d N d , from all others, to which the first bound of
(10) applies. Using (3), for 0 ≤ t ≤ τ 0\le t\le\tau 0 ≤ t ≤ τ ,
L 1 ( t ) ≤ Q 1 / 2 R d − 1 b d − 1 ( e 2 η + 16 2304 η e 1154 η ) . \sqrt{L_1(t)}\le Q^{1/2}R^{d-1}b_{d-1}
\left(e^{2\eta}+16\sqrt{2304\eta}\,e^{1154\eta}\right). L 1 ( t ) ≤ Q 1/2 R d − 1 b d − 1 ( e 2 η + 16 2304 η e 1154 η ) . It follows that
G 1 ( t ) ≤ A ( η ) Q R 2 d − 2 b d − 1 2 , A ( η ) = 2304 η e 2308 η + ( e 2 η + 16 2304 η e 1154 η ) 2 . (11) G_1(t)\le A(\eta)QR^{2d-2}b_{d-1}^2,
\quad A(\eta)=2304\eta e^{2308\eta}
+(e^{2\eta}+16\sqrt{2304\eta}\,e^{1154\eta})^2. \tag{11} G 1 ( t ) ≤ A ( η ) Q R 2 d − 2 b d − 1 2 , A ( η ) = 2304 η e 2308 η + ( e 2 η + 16 2304 η e 1154 η ) 2 . ( 11 ) In particular there exist universal M , η 0 > 0 M,\eta_0>0 M , η 0 > 0 such that
log [ e η ( 1 + η ) A ( η ) ] ≤ M η \log[e^\eta(1+\eta)A(\eta)]\le M\sqrt\eta log [ e η ( 1 + η ) A ( η )] ≤ M η for
0 < η ≤ η 0 0<\eta\le\eta_0 0 < η ≤ η 0 . This follows directly by expanding the finitely many
exponential factors at zero; alternatively the quotient by η \sqrt\eta η
extends continuously to zero and is bounded on any sufficiently short compact
interval.
Apply (9) with δ = η τ \delta=\eta\tau δ = η τ . Since τ ≤ η \tau\le\eta τ ≤ η ,
δ − 1 = d 2 / η 2 \delta^{-1}=d^2/\eta^2 δ − 1 = d 2 / η 2 , and the integral term adds at most η \eta η times the
bound for G 1 G_1 G 1 , the induction closes whenever
e η ( 1 + η ) A ( η ) ℓ r ( d 2 / η 2 ) 2 ℓ r ( d ) 2 ( d + 1 ) 4 d 4 ≤ ( 1 + α ) 2 . (12) e^\eta(1+\eta)A(\eta)
{\ell_r(d^2/\eta^2)^2\over\ell_r(d)^2}
{(d+1)^4\over d^4}\le(1+\alpha)^2. \tag{12} e η ( 1 + η ) A ( η ) ℓ r ( d ) 2 ℓ r ( d 2 / η 2 ) 2 d 4 ( d + 1 ) 4 ≤ ( 1 + α ) 2 . ( 12 ) Choose a universal 0 < c < min ( 1 , η 0 ) 0<c<\min(1,\eta_0) 0 < c < min ( 1 , η 0 ) so small that
M c ≤ 1 / 2 M\sqrt c\le1/2 M c ≤ 1/2 , and set η = c α 2 \eta=c\alpha^2 η = c α 2 .
The first factor of (12) is then at most e α / 2 e^{\alpha/2} e α /2 for every
0 < α ≤ 1 0<\alpha\le1 0 < α ≤ 1 . For all d ≥ 1 d\ge1 d ≥ 1 ,
g ( d 2 / η 2 ) ≤ 2 g ( d ) + 2 log ( 1 / η ) ≤ [ 2 + 2 log ( 1 / η ) ] g ( d ) . g(d^2/\eta^2)\le2g(d)+2\log(1/\eta)
\le[2+2\log(1/\eta)]g(d). g ( d 2 / η 2 ) ≤ 2 g ( d ) + 2 log ( 1/ η ) ≤ [ 2 + 2 log ( 1/ η )] g ( d ) . The first inequality follows from
e + d 2 / η 2 ≤ ( e + d ) 2 / η 2 e+d^2/\eta^2\le(e+d)^2/\eta^2 e + d 2 / η 2 ≤ ( e + d ) 2 / η 2 ; the second uses g ( d ) ≥ 1 g(d)\ge1 g ( d ) ≥ 1 .
Apply (2) to the remaining r − 1 r-1 r − 1 compositions:
log ℓ r ( d 2 / η 2 ) ℓ r ( d ) ≤ 2 − ( r − 1 ) log [ 2 + 2 log ( 1 / η ) ] ≤ α / 8 (13) \log{\ell_r(d^2/\eta^2)\over\ell_r(d)}
\le2^{-(r-1)}\log[2+2\log(1/\eta)]\le\alpha/8 \tag{13} log ℓ r ( d ) ℓ r ( d 2 / η 2 ) ≤ 2 − ( r − 1 ) log [ 2 + 2 log ( 1/ η )] ≤ α /8 ( 13 ) for all r r r above one universal r 0 r_0 r 0 . The uniformity in d d d is explicit:
d d d has disappeared from the upper bound. Existence of such r 0 r_0 r 0 follows
from r 2 2 − ( r − 1 ) log [ 2 + 2 log ( r 4 / c ) ] → 0 r^2 2^{-(r-1)}\log[2+2\log(r^4/c)]\to0 r 2 2 − ( r − 1 ) log [ 2 + 2 log ( r 4 / c )] → 0 .
For d ≥ D r d\ge D_r d ≥ D r , also 4 log ( 1 + 1 / d ) ≤ 4 / d ≤ α / 8 4\log(1+1/d)\le4/d\le\alpha/8 4 log ( 1 + 1/ d ) ≤ 4/ d ≤ α /8 .
Together (12)'s left side is at most
e α / 2 + α / 4 + α / 8 = e 7 α / 8 ≤ ( 1 + α ) 2 e^{\alpha/2+\alpha/4+\alpha/8}=e^{7\alpha/8}\le(1+\alpha)^2 e α /2 + α /4 + α /8 = e 7 α /8 ≤ ( 1 + α ) 2 ,
since 2 log ( 1 + α ) ≥ α 2\log(1+\alpha)\ge\alpha 2 log ( 1 + α ) ≥ α on [ 0 , 1 ] [0,1] [ 0 , 1 ] .
It remains to initialize the entire range d < D r d<D_r d < D r independently of the
hierarchy. The universal polynomial theorem gives c d ≤ 3 2 d d ! c_d\le32^dd! c d ≤ 3 2 d d ! .
If Γ ≥ 128 D r \Gamma\ge128D_r Γ ≥ 128 D r , then
3 2 d d ! ≤ ( 128 D r ) d ( d + 1 ) 2 ≤ R d b d ; 32^dd!\le{(128D_r)^d\over(d+1)^2}\le R^db_d; 3 2 d d ! ≤ ( d + 1 ) 2 ( 128 D r ) d ≤ R d b d ; use d ! ≤ D r d d!\le D_r^d d ! ≤ D r d , ( d + 1 ) 2 ≤ 4 d (d+1)^2\le4^d ( d + 1 ) 2 ≤ 4 d , and ℓ r ( d ) ≥ 1 \ell_r(d)\ge1 ℓ r ( d ) ≥ 1 .
As D r ≤ 33 r 2 D_r\le33r^2 D r ≤ 33 r 2 , choose K ≥ 128 ⋅ 33 K\ge128\cdot33 K ≥ 128 ⋅ 33 .
This initializes every smaller degree before the step at d ≥ D r d\ge D_r d ≥ D r .
Conditioning, whitening, fixed-order moment convergence and affine contraction
remove compactness and isotropy exactly as in the preceding induction.
Closing the depth induction ¶ The universal polynomial estimate gives
c k ≤ 3 2 k k ! ≤ 12 8 k k k / ( k + 1 ) 2 c_k\le32^kk!\le128^kk^k/(k+1)^2 c k ≤ 3 2 k k ! ≤ 12 8 k k k / ( k + 1 ) 2 .
Use Theorem 7.2 with ε = 1 \varepsilon=1 ε = 1 , R = 2 40 R=2^{40} R = 2 40 ,
and ℓ ( k ) = k \ell(k)=k ℓ ( k ) = k . Let h = g ( a − 1 ) ≥ 1 h=g(a^{-1})\ge1 h = g ( a − 1 ) ≥ 1 and choose dyadic d d d with
max ( 2 , h ) ≤ d < 2 max ( 2 , h ) ≤ 6 h \max(2,h)\le d<2\max(2,h)\le6h max ( 2 , h ) ≤ d < 2 max ( 2 , h ) ≤ 6 h .
Because log ( a − 1 ) ≤ h \log(a^{-1})\le h log ( a − 1 ) ≤ h when a − 1 ≥ 1 a^{-1}\ge1 a − 1 ≥ 1 ,
max ( 1 , a − 1 / ( d + 1 ) ) ≤ e \max(1,a^{-1/(d+1)})\le e max ( 1 , a − 1/ ( d + 1 ) ) ≤ e .
The comparison therefore gives
C P ( ν ) ≤ 36 e 2 85 h 2 ≤ ( 2 48 ) 2 ℓ 1 ( a − 1 ) 2 . C_P(\nu)\le36e\,2^{85}h^2\le(2^{48})^2\ell_1(a^{-1})^2. C P ( ν ) ≤ 36 e 2 85 h 2 ≤ ( 2 48 ) 2 ℓ 1 ( a − 1 ) 2 . This is the depth-one profile and makes no use of the final KLS conclusion.
Suppose a profile is available at depth r r r with constant Γ ≥ 2 48 \Gamma\ge2^{48} Γ ≥ 2 48 .
The coarse coefficient improvement gives radius R = 2 12 Γ ≥ 2 40 R=2^{12}\Gamma\ge2^{40} R = 2 12 Γ ≥ 2 40 .
Apply the comparison again with ε = 1 \varepsilon=1 ε = 1 and the same dyadic choice.
By (1), ℓ r ( d ) ≤ 6 ℓ r ( h ) = 6 ℓ r + 1 ( a − 1 ) \ell_r(d)\le6\ell_r(h)=6\ell_{r+1}(a^{-1}) ℓ r ( d ) ≤ 6 ℓ r ( h ) = 6 ℓ r + 1 ( a − 1 ) .
Consequently
C P ( ν ) ≤ 36 e 2 29 Γ 2 ℓ r + 1 ( a − 1 ) 2 ≤ ( 2 30 Γ ) 2 ℓ r + 1 ( a − 1 ) 2 . C_P(\nu)\le36e\,2^{29}\Gamma^2\ell_{r+1}(a^{-1})^2
\le(2^{30}\Gamma)^2\ell_{r+1}(a^{-1})^2. C P ( ν ) ≤ 36 e 2 29 Γ 2 ℓ r + 1 ( a − 1 ) 2 ≤ ( 2 30 Γ ) 2 ℓ r + 1 ( a − 1 ) 2 . Finite induction supplies preliminary constants
Γ ^ r = 2 48 + 30 ( r − 1 ) \widehat\Gamma_r=2^{48+30(r-1)} Γ r = 2 48 + 30 ( r − 1 ) at every fixed depth.
For r ≥ r 0 r\ge r_0 r ≥ r 0 , put α r = r − 2 \alpha_r=r^{-2} α r = r − 2 and suppose Γ ≥ K r 2 \Gamma\ge Kr^2 Γ ≥ K r 2 .
The sharp coefficient lemma supplies R = ( 1 + α r ) Γ R=(1+\alpha_r)\Gamma R = ( 1 + α r ) Γ .
Use the curvature comparison with ε = α r \varepsilon=\alpha_r ε = α r ;
its additional hypothesis is R ≥ 2 40 r 4 R\ge2^{40}r^4 R ≥ 2 40 r 4 .
Choose dyadic d d d with h / α r ≤ d < 2 h / α r h/\alpha_r\le d<2h/\alpha_r h / α r ≤ d < 2 h / α r .
Since r ≥ 2 r\ge2 r ≥ 2 , this automatically has d ≥ 2 d\ge2 d ≥ 2 .
Then max ( 1 , a − 1 / ( d + 1 ) ) ≤ e α r \max(1,a^{-1/(d+1)})\le e^{\alpha_r} max ( 1 , a − 1/ ( d + 1 ) ) ≤ e α r and, by (2),
log ℓ r ( d ) ℓ r ( h ) ≤ 2 − r log ( 2 / α r ) ≤ α r \log{\ell_r(d)\over\ell_r(h)}
\le2^{-r}\log(2/\alpha_r)\le\alpha_r log ℓ r ( h ) ℓ r ( d ) ≤ 2 − r log ( 2/ α r ) ≤ α r for all sufficiently large r r r , after increasing the fixed r 0 r_0 r 0 .
Taking square roots of the comparison gives
C P ( ν ) 1 / 2 ≤ 4 ( 1 + α r ) 3 / 2 e 3 α r / 2 Γ ℓ r + 1 ( a − 1 ) ≤ 4 e 3 / r 2 Γ ℓ r + 1 ( a − 1 ) . C_P(\nu)^{1/2}
\le4(1+\alpha_r)^{3/2}e^{3\alpha_r/2}
\Gamma\ell_{r+1}(a^{-1})
\le4e^{3/r^2}\Gamma\ell_{r+1}(a^{-1}). C P ( ν ) 1/2 ≤ 4 ( 1 + α r ) 3/2 e 3 α r /2 Γ ℓ r + 1 ( a − 1 ) ≤ 4 e 3/ r 2 Γ ℓ r + 1 ( a − 1 ) . Thus define Γ r + 1 = 4 e 3 / r 2 Γ r \Gamma_{r+1}=4e^{3/r^2}\Gamma_r Γ r + 1 = 4 e 3/ r 2 Γ r for r ≥ r 0 r\ge r_0 r ≥ r 0 .
To justify every application, choose the initial constant large enough that
Γ r 0 ≥ Γ ^ r 0 , 4 r − r 0 Γ r 0 ≥ max ( K r 2 , 2 40 r 4 ) for all r ≥ r 0 . (14) \Gamma_{r_0}\ge\widehat\Gamma_{r_0},\qquad
4^{r-r_0}\Gamma_{r_0}\ge\max(Kr^2,2^{40}r^4)
\quad\hbox{for all }r\ge r_0. \tag{14} Γ r 0 ≥ Γ r 0 , 4 r − r 0 Γ r 0 ≥ max ( K r 2 , 2 40 r 4 ) for all r ≥ r 0 . ( 14 ) One finite choice exists because each polynomial divided by 4 r 4^r 4 r is bounded.
The recurrence implies Γ r ≥ 4 r − r 0 Γ r 0 \Gamma_r\ge4^{r-r_0}\Gamma_{r_0} Γ r ≥ 4 r − r 0 Γ r 0 , so (14)
pays both thresholds before the next step is applied.
Also
Γ r = 4 r − r 0 Γ r 0 exp ( 3 ∑ j = r 0 r − 1 j − 2 ) ≤ C 0 4 r . \Gamma_r=4^{r-r_0}\Gamma_{r_0}
\exp\!\left(3\sum_{j=r_0}^{r-1}j^{-2}\right)\le C_0 4^r. Γ r = 4 r − r 0 Γ r 0 exp ( 3 j = r 0 ∑ r − 1 j − 2 ) ≤ C 0 4 r . For r < r 0 r<r_0 r < r 0 , take Γ r = Γ ^ r \Gamma_r=\widehat\Gamma_r Γ r = Γ r and enlarge C 0 C_0 C 0 to cover
these finitely many values. Every profile used to obtain the next coefficient
bound has already been established at the previous finite depth, for all
regular covariance contractions. The profile-inflation lemma supplies its
extension to every posterior needed in that coefficient induction.
Dependencies and fences. The argument uses
Lemma 7.1 , Theorem 7.1 (including its
proved localization and hierarchy lemmas), and Theorem 7.2 .
Brascamp--Lieb is an established literature input Brascamp & Lieb, 1976 Bakry et al. , 2014, Theorem 4.9.1 . There are
no bounded_by edges on this node. In the brief’s threshold test, (14) is the
explicit discharge for an exponentially growing sequence; it is not a
construction of an admissible bounded sequence. The argument neither proves
CMH nor provides universal-time occupation estimates. No infinite-depth limit
or dimension-free KLS conclusion is taken.
Song, Z., & Zhang, X. (2026). An O(4\log^* n) Bound for the KLS Constant . https://arxiv.org/abs/2610.01447v1 Brascamp, H. J., & Lieb, E. H. (1976). On Extensions of the Brunn–Minkowski and Prékopa–Leindler Theorems, Including Inequalities for Log Concave Functions, and with an Application to the Diffusion Equation. Journal of Functional Analysis , 22 (4), 366–389. 10.1016/0022-1236(76)90004-5 Bakry, D., Gentil, I., & Ledoux, M. (2014). Analysis and Geometry of Markov Diffusion Operators (Vol. 348). Springer. 10.1007/978-3-319-00227-9