Normalization and the simplex kernel. Under an invertible linear map L L L ,
the canonical kernel and covariance transform as τ ↦ L τ L ⊤ \tau\mapsto L\tau L^\top τ ↦ Lτ L ⊤
and Σ ↦ L Σ L ⊤ \Sigma\mapsto L\Sigma L^\top Σ ↦ L Σ L ⊤ . Indeed a moment potential ϕ \phi ϕ transforms
to ϕ ( L ⊤ y ) − log ∣ det L ∣ \phi(L^\top y)-\log|\det L| ϕ ( L ⊤ y ) − log ∣ det L ∣ ; its gradient and Hessian give these formulas.
The resulting normalized gate is conjugate to G G G by the orthogonal matrix
( L Σ L ⊤ ) − 1 / 2 L Σ 1 / 2 (L\Sigma L^\top)^{-1/2}L\Sigma^{1/2} ( L Σ L ⊤ ) − 1/2 L Σ 1/2 . We may therefore make each base factor
isotropic and regular by a block linear transformation, preserving the axis.
For a factor of dimension k k k , put N = k + 1 N=k+1 N = k + 1 , let
P = ( P 1 , … , P N ) P=(P_1,\ldots,P_N) P = ( P 1 , … , P N ) be uniform on the probability simplex, and write
H = { v ∈ R N : ∑ i v i = 0 } H=\{v\in\mathbb R^N:\sum_i v_i=0\} H = { v ∈ R N : ∑ i v i = 0 } and e = ( 1 , … , 1 ) \mathbf e=(1,\ldots,1) e = ( 1 , … , 1 ) .
An isotropic copy is
V = N ( N + 1 ) ( P − e / N ) ∈ H . V=\sqrt{N(N+1)}(P-\mathbf e/N)\in H. V = N ( N + 1 ) ( P − e / N ) ∈ H . Its canonical kernel on H H H is
T = ( N + 1 ) C ( P ) , C ( p ) = diag ( p ) − p p ⊤ . T=(N+1)C(P),\qquad C(p)=\operatorname{diag}(p)-pp^\top. T = ( N + 1 ) C ( P ) , C ( p ) = diag ( p ) − p p ⊤ . For completeness, the Dirichlet moment-potential computation underlying
Theorem 17.3 can be verified here without its inequality. On H H H take
ϕ ( y ) = N log ∑ i e y i / N \phi(y)=N\log\sum_i e^{y_i/N} ϕ ( y ) = N log ∑ i e y i / N plus a normalizing constant. Its gradient is
p − e / N p-\mathbf e/N p − e / N and its Hessian on H H H is C ( p ) / N C(p)/N C ( p ) / N , where
p i = e y i / N / ∑ j e y j / N p_i=e^{y_i/N}/\sum_j e^{y_j/N} p i = e y i / N / ∑ j e y j / N . The gradient is a diffeomorphism onto the
centered open simplex. The determinant lemma gives a principal minor
det C [ N − 1 ] = ∏ i p i \det C_{[N-1]}=\prod_i p_i det C [ N − 1 ] = ∏ i p i ; since ker C = R e \ker C=\mathbb R\mathbf e ker C = R e ,
det H C = N ∏ i p i \det_H C=N\prod_i p_i det H C = N ∏ i p i . Also e − ϕ e^{-\phi} e − ϕ is proportional to ∏ i p i \prod_i p_i ∏ i p i
because ∑ i y i = 0 \sum_i y_i=0 ∑ i y i = 0 . Change of variables therefore gives constant target
density. This finite smooth strictly convex potential is a canonical moment
potential, by uniqueness as used in the certified cone dossier. Scaling by
N ( N + 1 ) \sqrt{N(N+1)} N ( N + 1 ) gives T T T . The sum of the factor potentials similarly pushes
forward to the product law and has block diagonal Hessian. Consequently the
canonical base kernel is T K = ⨁ j T j T_K=\bigoplus_j T_j T K = ⨁ j T j . All these kernels are bounded
on their compact bases.
The three block moments. The uniform simplex integral, obtained by iterating
the elementary beta integral, is
E ∏ i = 1 N P i r i = ( N − 1 ) ! ∏ i r i ! ( N − 1 + ∑ i r i ) ! ( r i nonnegative integers ) . \mathbb E\prod_{i=1}^N P_i^{r_i}
=\frac{(N-1)!\prod_i r_i!}{(N-1+\sum_i r_i)!}
\qquad(r_i\text{ nonnegative integers}). E i = 1 ∏ N P i r i = ( N − 1 + ∑ i r i )! ( N − 1 )! ∏ i r i ! ( r i nonnegative integers ) . It gives E V = 0 \mathbb EV=0 E V = 0 and E V V ⊤ = I H \mathbb E VV^\top=I_H E V V ⊤ = I H . Permutation invariance
implies that an invariant endomorphism of H H H is scalar: extend it by zero on
R e \mathbb R\mathbf e R e , and invariance forces all diagonal entries to be equal
and all off-diagonal entries to be equal. In particular the three expectations
below are scalar, and their scalars follow by taking traces:
E T 2 = a k I H , E [ V V ⊤ T ] = E [ T V V ⊤ ] = b k I H , E [ ∣ V ∣ 2 V V ⊤ ] = c k I H . \mathbb ET^2=a_k I_H,\qquad
\mathbb E[VV^\top T]=\mathbb E[TVV^\top]=b_k I_H,\qquad
\mathbb E[|V|^2VV^\top]=c_k I_H. E T 2 = a k I H , E [ V V ⊤ T ] = E [ T V V ⊤ ] = b k I H , E [ ∣ V ∣ 2 V V ⊤ ] = c k I H . Here is the full scalar computation. With Q 2 = ∑ i P i 2 Q_2=\sum_iP_i^2 Q 2 = ∑ i P i 2 and
Q 3 = ∑ i P i 3 Q_3=\sum_iP_i^3 Q 3 = ∑ i P i 3 , the displayed simplex integral gives
E Q 2 = 2 N + 1 , E Q 3 = 6 ( N + 1 ) ( N + 2 ) , E Q 2 2 = 4 N + 20 ( N + 1 ) ( N + 2 ) ( N + 3 ) . \mathbb E Q_2=\frac2{N+1},\qquad
\mathbb E Q_3=\frac6{(N+1)(N+2)},\qquad
\mathbb E Q_2^2=\frac{4N+20}{(N+1)(N+2)(N+3)}. E Q 2 = N + 1 2 , E Q 3 = ( N + 1 ) ( N + 2 ) 6 , E Q 2 2 = ( N + 1 ) ( N + 2 ) ( N + 3 ) 4 N + 20 . Since Tr C 2 = Q 2 − 2 Q 3 + Q 2 2 \operatorname{Tr} C^2=Q_2-2Q_3+Q_2^2 Tr C 2 = Q 2 − 2 Q 3 + Q 2 2 ,
V ⊤ T V = N ( N + 1 ) 2 ( Q 3 − Q 2 2 ) V^\top TV=N(N+1)^2(Q_3-Q_2^2) V ⊤ T V = N ( N + 1 ) 2 ( Q 3 − Q 2 2 ) , and
∣ V ∣ 4 = N 2 ( N + 1 ) 2 ( Q 2 − 1 / N ) 2 |V|^4=N^2(N+1)^2(Q_2-1/N)^2 ∣ V ∣ 4 = N 2 ( N + 1 ) 2 ( Q 2 − 1/ N ) 2 , division by k = N − 1 k=N-1 k = N − 1 yields
a k = 2 ( N + 1 ) N + 3 , b k = 2 N ( N + 1 ) ( N + 2 ) ( N + 3 ) , c k = ( N + 1 ) ( N 2 + 7 N − 6 ) ( N + 2 ) ( N + 3 ) , a_k=\frac{2(N+1)}{N+3},\qquad
b_k=\frac{2N(N+1)}{(N+2)(N+3)},\qquad
c_k=\frac{(N+1)(N^2+7N-6)}{(N+2)(N+3)}, a k = N + 3 2 ( N + 1 ) , b k = ( N + 2 ) ( N + 3 ) 2 N ( N + 1 ) , c k = ( N + 2 ) ( N + 3 ) ( N + 1 ) ( N 2 + 7 N − 6 ) , which are the stated expressions. These moment evaluations include k = 1 k=1 k = 1 ,
so they cover every interval after centering and scaling.
Assembly of the cone. Write U = ( V 1 , … , V q ) U=(V_1,\ldots,V_q) U = ( V 1 , … , V q ) , m = n − 1 m=n-1 m = n − 1 , and
T K = ⨁ j T j T_K=\bigoplus_j T_j T K = ⨁ j T j . The base is isotropic. By Proposition 17.1 ,
with S ∼ Γ ( β , 1 ) S\sim\Gamma(\beta,1) S ∼ Γ ( β , 1 ) independent of U U U ,
Σ = β ⊕ β ( β + 1 ) I m , τ = S ( 1 U ⊤ U B ) , B = U U ⊤ + β T K . \Sigma=\beta\oplus\beta(\beta+1)I_m,\qquad
\tau=S\begin{pmatrix}1&U^\top\\ U&B\end{pmatrix},\qquad
B=UU^\top+\beta T_K. Σ = β ⊕ β ( β + 1 ) I m , τ = S ( 1 U U ⊤ B ) , B = U U ⊤ + β T K . All expectations below are finite by boundedness on the base and
E S 2 = β ( β + 1 ) \mathbb ES^2=\beta(\beta+1) E S 2 = β ( β + 1 ) . Matrix multiplication shows that the transverse
block of E [ τ Σ − 1 τ ] \mathbb E[\tau\Sigma^{-1}\tau] E [ τ Σ − 1 τ ] is
( β + 1 ) I m + E B 2 (\beta+1)I_m+\mathbb EB^2 ( β + 1 ) I m + E B 2 , whereas its axis entry is β + n \beta+n β + n .
The off-axis vector is zero: it is invariant under independent permutations
of the vertices of every factor, and the only invariant vector in each H j H_j H j
is zero. The same invariance forces off-diagonal factor blocks of
E B 2 \mathbb EB^2 E B 2 to vanish (average permutations in just one factor).
Expand
B 2 = ∣ U ∣ 2 U U ⊤ + β ( U U ⊤ T K + T K U U ⊤ ) + β 2 T K 2 . B^2=|U|^2UU^\top+\beta(UU^\top T_K+T_KUU^\top)+\beta^2T_K^2. B 2 = ∣ U ∣ 2 U U ⊤ + β ( U U ⊤ T K + T K U U ⊤ ) + β 2 T K 2 . In a factor of dimension k k k , independence and E ∣ V i ∣ 2 = k i \mathbb E|V_i|^2=k_i E ∣ V i ∣ 2 = k i
show that the first expectation is ( c k + m − k ) I k (c_k+m-k)I_k ( c k + m − k ) I k . The remaining two terms
have expectations 2 β b k I k 2\beta b_k I_k 2 β b k I k and β 2 a k I k \beta^2a_k I_k β 2 a k I k . Dividing the cone
block by β ( β + 1 ) \beta(\beta+1) β ( β + 1 ) proves the claimed λ k \lambda_k λ k and axis spectrum.
The sharp gap. Put t = β − n ≥ 0 t=\beta-n\ge0 t = β − n ≥ 0 and d = n − k − 1 ≥ 0 d=n-k-1\ge0 d = n − k − 1 ≥ 0 . Substitution of
the three scalar expressions, followed by expansion, gives the identity
( k + 3 ) ( k + 4 ) β ( β + 1 ) ( 2 − λ k ( n , β ) ) = 4 ( k + 3 ) t 2 + ( 5 k 2 + 27 k + 28 + 8 ( k + 3 ) d ) t + 4 d ( ( k + 3 ) n + k + 1 ) . \begin{aligned}
&(k+3)(k+4)\beta(\beta+1)(2-\lambda_k(n,\beta))\\
&\quad=4(k+3)t^2+
\bigl(5k^2+27k+28+8(k+3)d\bigr)t
+4d\bigl((k+3)n+k+1\bigr).
\end{aligned} ( k + 3 ) ( k + 4 ) β ( β + 1 ) ( 2 − λ k ( n , β )) = 4 ( k + 3 ) t 2 + ( 5 k 2 + 27 k + 28 + 8 ( k + 3 ) d ) t + 4 d ( ( k + 3 ) n + k + 1 ) . Every coefficient is strictly positive for k ≥ 1 k\ge1 k ≥ 1 and n ≥ k + 1 n\ge k+1 n ≥ k + 1 .
The right side is nonnegative, and vanishes exactly when t = d = 0 t=d=0 t = d = 0 .
Thus a transverse block reaches two exactly when β = n \beta=n β = n and k = n − 1 k=n-1 k = n − 1 ,
which means there is just one factor. The axis gap is
2 − ( 1 + n / β ) = ( β − n ) / β 2-(1+n/\beta)=(\beta-n)/\beta 2 − ( 1 + n / β ) = ( β − n ) / β . Together these prove all asserted strict
inequalities and equality subspaces.