Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Song–Zhang v2: the improved inner iteration

Part of the second version of Song–Zhang, Chapter Song–Zhang, second version: repeated refinement with summable losses; the reading order is on the full proofs page.

Overview. This dossier reconstructs the operator argument of Section 6 of Song & Zhang, 2026. A restricted inverse operator gives a low-energy gradient. Normalization pays for its skew linear moment. Averaging an orbit then reduces the startup losses to cubic order in the inverse radius. A joint dyadic frame controls all subsequent losses. The resulting comparison iterates a static coefficient radius; the additive conversion from that radius to CPC_P is used only once at the end.

Dependencies. We use Lemma 7.1, Theorem 7.1, Theorem 7.2, the fixed-depth profiles of Theorem 7.3, and the three new foundations Lemma 10.1, Lemma 10.2, Proposition 10.1. The latter have their full proofs in the separate foundations dossier and require independent review. No BKL result or consequence of KLS is used.

Fix a centered regular measure of covariance Σ⪯I\Sigma\preceq I and Hessian D2W⪰aI>0D^2W\succeq aI>0. Use the analytic operator HH, B=H−1B=H^{-1} on centered functions, P+f=f−EfP_+f=f-\mathbb Ef, D=P+∇H−1/2D=P_+\nabla H^{-1/2}, Lf=E[Xf]Lf=\mathbb E[Xf], and T=P+∇B=DB1/2\mathcal T=P_+\nabla B=DB^{1/2}. Operators act componentwise on finite families and append ordered derivative slots. Put P=CP=λ−1P=C_P=\lambda^{-1}. All inverse powers, gradients and second derivatives below are justified by the form-domain and Bochner results in Lemma 7.1.

Restricted operator and normalized hierarchy

The defect in inverse normalization also gives the identity

uj+1=βj+1−1/2P+∇uj+P+∇zj,∥∇zj∥2=χj/βj+1.(1)u^{j+1}=\beta_{j+1}^{-1/2}P_+\nabla u^j+P_+\nabla z^j, \qquad\|\nabla z^j\|^2=\chi_j/\beta_{j+1}. \tag{1}

Indeed take zj=B1/2Fj+1−βj+1−1/2ujz^j=B^{1/2}F_{j+1}-\beta_{j+1}^{-1/2}u^j and expand its Dirichlet norm. The first term has symmetric newest derivative slots. Thus a fresh adjacent swap has norm at most 2χj/λ2\sqrt{\chi_j/\lambda}. An older swap is acted on by βT\sqrt\beta\mathcal T, with its scalar normalizer fixed to that of the original whole family, and never by an uncontrolled differentiation.

For Appell testing put Qkh=E[Akh]Q_kh=\mathbb E[\mathcal A_kh]. Its norm is at most k!ckk!c_k. Integration by parts and the derivative rule for Appell polynomials give Qk+1/(k+1)=Pk+1QkTQ_{k+1}/(k+1)=P_{k+1}Q_k\mathcal T and

Pk(βjQ1uj−1)=∏i=0k−1βj−ik!Qkuj−k.(2)P_k(\sqrt{\beta_j}Q_1u^{j-1}) =\frac{\prod_{i=0}^{k-1}\sqrt{\beta_{j-i}}}{k!}Q_ku^{j-k}. \tag{2}

These hold on full output direct sums. The removed mean gradient contributes zero since every positive-degree Appell polynomial is centered.

The initial restart and a mesoscopic orbit

A coarse coefficient seed needed below follows without a new localization argument. At the fixed depth r∗r_* of the static transfer, the certified v1 profile bounds CPC_P by Γ2ℓr∗(a−1)2\Gamma^2\ell_{r_*}(a^{-1})^2 for one universal Γ\Gamma. Repeated Poincaré inequalities on Appell derivatives give ck≤CP(k−1)/2c_k\le C_P^{(k-1)/2}: start with covariance normalization at degree one, and use Kk≤CPk2Kk−1K_k\le C_P k^2K_{k-1}. Apply Proposition 10.1. Since the fixed iterates of gg satisfy ℓr∗(d)≤Cg(d)\ell_{r_*}(d)\le Cg(d), this proves

ck≤[Cseedlog⁡(e+k)]k−1(k≥1).(3)c_k\le[C_{\rm seed}\log(e+k)]^{k-1}\qquad(k\ge1). \tag{3}

The constant is uniform over all log-concave covariance contractions. In particular choose a universal cˉ3≥c3\bar c_3\ge c_3. Letwin gives K2≤8K_2\le8.

Averaged orbit and its static radius

Joint losses and finite comparison

Closing the depth recurrence

Fences respected. No bounded_by edge is proposed for this node. The brief’s small-degree initialization requirement is met by the static transfer starting at degree two and the fixed seed (3). All large-radius requirements are universal and are imposed before starting the depth induction; their common constant is GG. This profile is still curvature dependent, and its dimension transfer still depends on dimension. No CMH, occupation or trace antecedent is discharged. There is no infinite-depth limit and no use of a KLS-equivalent coefficient assertion without its explicit profile premise.

References
  1. Song, Z., & Zhang, X. (2026). An O(1) Bound for the KLS Constant. https://arxiv.org/abs/2610.01447v2