Repository · Full text

Strong average self-concordance of the
log-determinant metric at linear scale

Read PDF

HTML version 1 Added

Papers are listed without authors and are not intended for submission or formal publication.

Contents

Strong average self-concordance of the
log-determinant metric at linear scale

Abstract

Strong average self-concordance controls the change in the local norm of a Gaussian step when an arbitrary positive-semidefinite form is added to its proposal precision. For the log-determinant metric G=∇2(−logdet) on real symmetric q×q positive-definite matrices, we prove this property for every scaled metric s⁢G with s≥(q+1)/2. This replaces the quadratic sufficient scale of Kook and Vempala (COLT 2024) by a linear one. The admissible proposal radius depends only on the error tolerance, and the guarantee includes feasibility and a two-sided bound on the change in squared local norm. The auxiliary form can correlate matrix entries. We normalize the proposal to a centered Gaussian matrix with covariance bounded by the identity. A paired-trace estimate of Song and Zhang (arXiv 2026) bounds its quartic trace moment, and Gaussian Poincaré reduces the cubic-trace variance to the same fourth moment. Both are O⁡(q3) uniformly over these covariance contractions, yielding the linear scale. We also determine the exact strong self-concordance threshold (q+1)/2, making this the optimal scale for requiring strong self-concordance and strong average self-concordance simultaneously.

Note. This paper was generated entirely by AI, including MiMo, using an automated research pipeline developed by Chenghua Liu and Hanyu Li.

1 Introduction

A Dikin proposal measures a Gaussian displacement in the metric at its starting point; the reverse proposal measures that same displacement at its endpoint. Average self-concordance controls the difference between these two squared local lengths. When the proposal metric is a sum of components, a useful statement must control each component under the Gaussian step generated by the full sum. Strong average self-concordance (SASC) imposes precisely this requirement by allowing an arbitrary positive-semidefinite addition to the proposal precision [1].

For the log-determinant barrier on q×q positive-definite matrices, this addition changes the structure of the random matrix. Without it, a congruence transformation produces a Frobenius-isotropic Gaussian matrix, whose independent coordinates permit direct trace moment calculations. With it, the transformed covariance is still bounded by the identity, but the matrix entries may be correlated. Kook and Vempala [1, Theorem 3.6] prove ordinary average self-concordance at scale q, whereas their SASC guarantee uses the ambient dimension n=q⁡(q+1)/2 as the scale. They identify the loss of entry independence in Remark G.7 and ask whether the quadratic SASC scale can be improved.

Covariance domination alone does not immediately supply the trace estimates needed at linear scale. For example, replacing Tr⁡(A4) by ‖A‖F4 for a centered Gaussian matrix with covariance at most the identity gives an expectation bound n⁡(n+2)=O⁡(q4). After factoring out the target norm-change scale r2/n, this term is multiplied by t2=r2/(n⁢s), where s is the metric scale and r the proposal radius. The resulting bound is of order r2⁢q2/s, losing a factor of q relative to the desired linear scale. The loss is already visible for the isotropic Gaussian: its quartic trace has order q3, although the fourth moment of its Frobenius norm has order q4. A successful estimate must retain the trace structure while allowing arbitrary correlations.

We prove SASC for every s≥(q+1)/2. The normalized covariance can be diagonalized in an orthonormal basis of symmetric matrices, even though that basis need not be the coordinate basis. The identity ∑aFa2=(q+1)⁢Iq/2 holds in every such basis. Combined with the paired-trace estimate of Song and Zhang [3, Lemma 5.8], it gives the required O⁡(q3) fourth moment uniformly over covariance contractions. Gaussian Poincaré then bounds the variance of Tr⁡(A3) by that same fourth moment. These two estimates control the leading cubic term and the quartic remainder in the norm change. The quartic estimate also ensures that the proposed endpoint stays inside the cone.

The scale has a second interpretation. Strong self-concordance (SSC), introduced in the sampling setting by Laddha, Lee, and Vempala [2], is a deterministic bound on the normalized metric derivative. We compute its Hilbert–Schmidt norm exactly and show that SSC holds if and only if s≥(q+1)/2, attaining the lower endpoint of [1, Lemma E.23]. Hence (q+1)/2 is the minimum scale at which SSC and SASC hold together. This necessity comes from SSC; the SASC-only question concerns uniformity of the proposal radius across matrix orders and is discussed after the SSC calculation.

1.1 Metric and main theorem

We consider the log-determinant barrier on the positive semidefinite (PSD) cone. For an integer q≥1, let Sq be the space of real symmetric q×q matrices and let S+⁣+q be its positive-definite cone. The Frobenius inner product is ⟨U,V⟩F=Tr⁡(U⊤⁢V), which equals Tr⁡(U⁢V) for symmetric matrices. Write n=dimSq=q⁡(q+1)/2 and cq=(q+1)/2, so that n=q⁢cq. For X∈S+⁣+q, set

ϕ(X)=−logdetX,GX[U,V]=Tr(X−1UX−1V),gs(X)=sGX,s>0.(1)

We identify bilinear forms with their self-adjoint Riesz operators for the Frobenius inner product. In particular, ‖H‖gs⁢(X)2=s⁢Tr⁡(X−1⁢H⁢X−1⁢H). The notation ‖⋅‖HS denotes the Hilbert–Schmidt norm of an operator on Sq, whereas ‖⋅‖F and ‖⋅‖op denote matrix norms.

The distinction between operator and matrix norms matters for SSC, which measures all directions of the metric derivative together. We use the following normalizations. For a Hessian metric, ordinary self-concordance means |D3⁢f⁢(X)⁢[H,H,H]|≤2⁢‖H‖∇2f⁢(X)3.

Definition 1.1 (Strong self-concordance).

A self-concordant positive-definite C1 Hessian metric g is strongly self-concordant (SSC) if

‖g(X)−1/2Dg(X)[H]g(X)−1/2‖HS≤2‖H‖g⁡(X)(2)

for every base point X and every direction H.

Definition 1.2 (Uniform strong average self-concordance).

Let F be a family of continuous positive-definite metrics, each defined on S+⁣+q for some q. The family is uniformly strongly average self-concordant if, for every 0<ε<1, there is rε>0 with the following property. For every g∈F, every base point X∈S+⁣+q, every positive-semidefinite bilinear form g¯⁢(X) on Sq, and every 0<r≤rε, the proposal

H∼NSq⁢(0,r2n⁢(g⁡(X)+g¯⁢(X))−1),Z=X+H,(3)

satisfies

P[Z∈S+⁣+q,‖H‖g⁡(Z)2−‖H‖g⁡(X)2≤2⁢ε⁢r2n]≥1−ε.(4)

A metric is called SASC when its singleton family has this property. Setting g¯⁢(X)=0 gives the corresponding average self-concordance condition.

The covariance in (3) is relative to the Frobenius inner product. The radius is independent of the matrix order, the metric in the family, the base point, and the auxiliary form. Only the pointwise value g¯⁢(X) enters the proposal; it may be singular, and no smoothness or conditioning assumption is imposed. The norm change is measured using g, not g+g¯. Proposals outside the cone are counted as failures, so the event does not require evaluating g outside its domain.

Theorem 1.3 (Linear SASC scaling and the exact SSC threshold).

For every integer q≥1 and every s≥cq=(q+1)/2, the proposal (3) with g=gs satisfies (4) uniformly over X∈S+⁣+q and g¯⁢(X)⪰0, whenever 0<ε<1 and

0<r≤rε:=ε3/2864.(5)

In fact, the stronger two-sided estimate holds:

P[Z∈S+⁣+q,|‖H‖gs⁢(Z)2−‖H‖gs⁢(X)2|≤2⁢ε⁢r2n]≥1−ε.(6)

Thus {gs:q≥1,s≥(q+1)/2} is uniformly SASC.

Moreover, gs is SSC if and only if s≥(q+1)/2. Consequently, for each fixed q,

min⁡{s>0:gs⁢ is both SSC and SASC}=q+12.(7)

The theorem controls the norm change of the log-determinant component itself. This differs from controlling the complete Metropolis ratio, where the quadratic and determinant terms may cancel. Song and Zhang [3, Lemmas 5.3 and 5.15] use such cancellation in their analysis of spectrahedral walks; their covariance-eigenmode argument in Lemma 5.9 is also a technical ingredient of the trace estimates used here. Chen and Kook [4, Theorem 1 and Corollary 3] already obtain O~⁢(n2) total-variation mixing on truncated PSD cones for a hybrid metric containing a linearly scaled log-determinant component, through first-order regularity. Their separate fourth-order criterion concerns ordinary ASC [4, Definition 5 and Lemma 6]. Thus a linear SASC bound answers the local regularity question without asserting a further mixing-time improvement. Numbered references to [1] use the published COLT version; those to [3, 4] use arXiv version 1.

2 Gaussian trace estimates

We first remove the base point and the auxiliary precision from the moment calculation. For X≻0, the invertible congruence operator TX(H)=X−1/2HX−1/2 satisfies GX=TX∗⁢TX. Normalization leaves a Gaussian covariance contraction on the full Frobenius space, as observed in [1, Remark G.7]. The subsequent estimates depend only on that contraction, so they apply uniformly to every auxiliary form in Theorem 1.3.

Lemma 2.1 (Normalized proposal covariance).

Let H have the law (3) for g=gs, and set

t=rn⁢s,A=t−1⁢TX⁢(H).(8)

Then A is a centered Gaussian in Sq, and its covariance operator Q satisfies 0⪯Q⪯I.

Proof.

Writing G¯X⪰0 for the Riesz operator of g¯⁢(X), the covariance transformation rule gives

Q=s⁢TX⁢(s⁢TX∗⁢TX+G¯X)−1⁢TX∗⪯s⁢TX⁢(s⁢TX∗⁢TX)−1⁢TX∗=I.

The inequality follows from inverse monotonicity of positive-definite operators, and the last identity follows from invertibility of TX. The displayed covariance is positive semidefinite. ∎

Remark 2.2 (Coordinate normalization).

An orthonormal basis of Sq is Ei⁢i and (Ei⁢j+Ej⁢i)/2 for i<j, where Ei⁢j is the elementary matrix. Let fsvec denote coordinates in this basis, so its off-diagonal coordinates are 2⁢Hi⁢j. If svec⁡(H) instead stores each off-diagonal entry once, define LX by LX⁢svec⁡(H)=fsvec⁡(TX⁢H). The raw-coordinate matrix of GX is then GX=LX⊤⁢LX, and

Cov⁡(fsvec⁡(A))=s⁢LX⁢(s⁢GX+G¯X)−1⁢LX⊤⪯I,

where G¯X is the raw-coordinate matrix of g¯⁢(X). All Hilbert–Schmidt norms use orthonormal coordinates; the raw upper-triangular coordinate map is not a Frobenius isometry.

Diagonalizing a covariance operator gives independent scalar Gaussian coefficients multiplying symmetric matrices that need not commute. In a fourth-moment expansion, the difficult terms have the alternating order Ba⁢Bb⁢Ba⁢Bb. The following R=Iq specialization of [3, Lemma 5.8] bounds them by the adjacent terms Ba2⁢Bb2, whose sum is controlled by ∑aBa2.

Lemma 2.3 (Paired traces).

Let B1,…,Bm∈Sq, and suppose B:=∑aBa2⪯Ψ for a PSD matrix Ψ. Then

∑a,b=1m|Tr⁡(Ba⁢Bb⁢Ba⁢Bb)|≤Tr⁡(B2)≤Tr⁡(Ψ2).(9)
Proof.

For symmetric U,V, Frobenius Cauchy–Schwarz gives

|Tr⁡(U⁢V⁢U⁢V)|=|⟨V⁢U,U⁢V⟩F|≤‖U⁢V‖F2=Tr⁡(U2⁢V2).(10)

Summing with U=Ba, V=Bb proves the first inequality in (9). For the second, both Ψ−B and Ψ+B are PSD, and therefore

Tr⁡(Ψ2)−Tr⁡(B2)=Tr⁡((Ψ−B)⁢(Ψ+B))≥0.

The nonnegativity follows by writing the trace of a product of PSD matrices M,N as Tr⁡(M1/2⁢N⁢M1/2). ∎

The paired-trace inequality turns covariance domination into a fourth-moment bound without discarding matrix multiplication inside the trace. The cubic trace is odd, hence centered; differentiating it in Gaussian coordinates produces A2, so Gaussian Poincaré reduces its variance to the same fourth moment. This also covers singular covariance operators.

Proposition 2.4 (Gaussian trace bounds).

Let A be a centered Gaussian in Sq with covariance 0⪯Q⪯I. Then

E⁢Tr⁡(A4)≤3⁢n⁢cq≤3⁢q3,E⁢(Tr⁡(A3))2≤27⁢n⁢cq≤27⁢q3.(11)
Proof.

Diagonalize Q in a Frobenius-orthonormal basis F1,…,Fn, and write

A=∑a=1nλaξaFa,0≤λa≤1,ξaindependent N(0,1).

The covariance basis may rotate arbitrarily with the auxiliary form. The following sum-of-squares identity is unchanged by that rotation:

∑a=1nFa2=cq⁢Iq.(12)

In the basis of Remark 2.2, the diagonal terms sum to Iq, and the off-diagonal terms sum to (q−1)⁢Iq/2. For another orthonormal basis, writing its change-of-basis matrix as O, the identity ∑aOa⁢b⁢Oa⁢c=δb⁢c shows that the sum of squares is unchanged.

Set Ba=λa⁢Fa and B=∑aBa2. As in the covariance-eigenbasis argument of [3, Lemma 5.9], 0≤λa≤1 gives

0⪯B⪯∑aFa2=cq⁢Iq,Tr⁡(B2)≤q⁢cq2=n⁢cq.(13)

We now expand the fourth moment in the independent coefficients, retaining the order of the matrix factors. Using E⁡(ξa⁢ξb⁢ξc⁢ξd)=δa⁢b⁢δc⁢d+δa⁢c⁢δb⁢d+δa⁢d⁢δb⁢c, the three Wick pairings give

E⁢Tr⁡(A4)=2⁢Tr⁡(B2)+∑a,bTr⁡(Ba⁢Bb⁢Ba⁢Bb).(14)

Lemma 2.3 bounds the last sum in absolute value by Tr⁡(B2), proving E⁢Tr⁡(A4)≤3⁢n⁢cq.

For the second estimate, let p⁡(A)=Tr⁡(A3). Its expectation is zero because the law of A is invariant under A↦−A. Gaussian Poincaré in the independent coordinates ξa gives

E⁢p⁢(A)2=Var⁡p⁡(A)≤9⁢E⁢∑a=1nλa⁢⟨A2,Fa⟩F 2
≤9⁢E⁢‖A2‖F2=9⁢E⁢Tr⁡(A4)≤27⁢n⁢cq.(15)

Here ∂ξap⁡(A)=3⁢λa⁢Tr⁡(A2⁢Fa), and the second inequality uses Parseval and λa≤1. Finally, n⁢cq=q⁢(q+1)2/4≤q3 for q≥1.

The polynomial Gaussian Poincaré inequality used here is recalled with a self-contained proof in Appendix A.1. ∎

The q3 orders in (11) cannot be reduced uniformly over covariance contractions. For the Frobenius-isotropic Gaussian W, whose covariance is the identity,

E⁢Tr⁡(W4)=q⁡(2⁢q2+5⁢q+5)4,E⁢(Tr⁡(W3))2=3⁢q⁢(4⁢q2+9⁢q+7)4.(16)

These identities, derived in Appendix A.2, locate the factor of q saved over a Frobenius-norm estimate. The appendix also shows by convexity that the isotropic quartic moment is maximal under Q⪯I. The cubic-trace variance bound above uses Gaussian Poincaré, so it does not require a covariance-monotonicity assertion for a sixth-degree polynomial.

3 Uniform control of the local norm change

At scale s≥cq, the moment bounds of Proposition 2.4, multiplied by t2=r2/(n⁢s), are independent of q. To turn them into a statement about the endpoint metric, we need both a remainder bound and feasibility. The deterministic estimate below applies while t⁢‖A‖op≤1/2; in the probabilistic argument, the same quartic event that bounds its remainder will enforce this condition.

Lemma 3.1 (Metric-change remainder).

If A∈Sq, t>0, and t⁢‖A‖op≤1/2, then

|Ft⁢(A)|≤2⁢t⁢|Tr⁡(A3)|+8⁢t2⁢Tr⁡(A4),Ft⁢(A)=Tr⁡[A2⁢((Iq+t⁢A)−2−Iq)].(17)
Proof.

For |y|≤1/2,

y2⁢((1+y)−2−1)=−2⁢y3+y4⁢3+2⁢y(1+y)2,0≤3+2⁢y(1+y)2=21+y+1(1+y)2≤8.

Apply this identity to y=t⁢μi, where μi are the eigenvalues of A, divide by t2, and sum. This gives

Ft⁢(A)=−2⁢t⁢Tr⁡(A3)+t2⁢∑i=1qμi4⁢3+2⁢t⁢μi(1+t⁢μi)2.

The triangle inequality and the coefficient bound prove (17). ∎

Proof of the SASC assertion in Theorem 1.3.

Fix q,s,X,g¯⁢(X),ε,r as in the theorem, and define t,A by (8). Lemma 2.1 gives Cov⁡(A)⪯I. Proposition 2.4, together with s≥cq, yields

t2⁢E⁢Tr⁡(A4)≤3⁢r2⁢cqs≤3⁢r2,t2⁢E⁢(Tr⁡(A3))2≤27⁢r2⁢cqs≤27⁢r2.(18)

The cubic term is centered and is controlled by its second moment; the quartic term is nonnegative and is controlled by its first moment. Accordingly, define

E3={t|Tr(A3)|≤ε2},E4={t2Tr(A4)≤ε16}.

Chebyshev’s inequality and Markov’s inequality imply

P⁡(E3c)≤4⁢t2ε2⁢E⁢(Tr⁡(A3))2≤108⁢r2ε2≤ε8,(19)
P⁡(E4c)≤16⁢t2ε⁢E⁢Tr⁡(A4)≤48⁢r2ε≤ε218,(20)

where the final inequalities use r2≤ε3/864. Their sum is less than ε for 0<ε<1.

Before applying the remainder estimate, we use E4 to verify that the endpoint metric is defined. Since n⁢s≥1 and r<1, we have t≤r<1. On E4,

(t⁢‖A‖op)4≤t4⁢Tr⁡(A4)=t2⁢(t2⁢Tr⁡(A4))≤ε16≤116.

Hence Iq+t⁢A⪰Iq/2, and Z=X1/2⁢(Iq+t⁢A)⁢X1/2≻0.

On E3∩E4, substituting H=t⁢X1/2⁢A⁢X1/2 and using that A commutes with (Iq+t⁢A)−1 gives

Δ:=‖H‖gs⁢(Z)2−‖H‖gs⁢(X)2
=s⁢t2⁢Tr⁡[A2⁢((Iq+t⁢A)−2−Iq)]=r2n⁢Ft⁢(A).(21)

Lemma 3.1 now yields

|Ft⁢(A)|≤2⁢t⁢|Tr⁡(A3)|+8⁢t2⁢Tr⁡(A4)≤ε+ε2<2⁢ε.

Thus (6), and hence (4), holds with probability at least 1−ε/8−ε2/18>1−ε. All dependence on the matrix order and scale entered through cq/s≤1 in (18); the base point and auxiliary form entered only through Q⪯I. This proves the asserted uniformity. ∎

4 The exact strong self-concordance threshold

The probabilistic argument gives SASC at every scale s≥cq. Necessity for the joint assertion is determined by the deterministic SSC inequality. After the same congruence normalization, the metric derivative acts on symmetric matrices by the anticommutator U↦−(K⁢U+U⁢K). Its full spectrum gives the exact Hilbert–Schmidt norm; the extremal direction is a scalar dilation of the base point.

Proposition 4.1 (Normalized derivative identity).

For X≻0, H∈Sq, and K=X−1/2HX−1/2,

‖gs(X)−1/2Dgs(X)[H]gs(X)−1/2‖HS2=(q+2)‖K‖F2+(TrK)2.(22)

Consequently, gs is SSC if and only if s≥(q+1)/2.

Proof.

Write U~=TX⁢(U) and V~=TX⁢(V). Differentiating (1) gives

D⁢gs⁢(X)⁢[H]⁢[U,V]=−s⁢⟨K⁢U~+U~⁢K,V~⟩F.

Define LK⁢(U)=−(K⁢U+U⁢K). The map WX=s⁢TX satisfies gs⁢(X)=WX∗⁢WX and D⁢gs⁢(X)⁢[H]=WX∗⁢LK⁢WX. Therefore JX=WXgs(X)−1/2 is orthogonal, and the normalized derivative is JX∗⁢LK⁢JX. Its Hilbert–Schmidt norm equals that of LK.

Diagonalize K, with eigenvalues k1,…,kq. In the associated orthonormal symmetric-matrix basis, the eigenvalues of LK are −2⁢ki and −(ki+kj) for i<j. It follows that

‖LK‖HS2=4⁢∑iki2+∑i<j(ki+kj)2
=(q+2)⁢∑iki2+(∑iki)2≤2⁢(q+1)⁢‖K‖F2.

This proves (22), with equality in the final inequality for K=Iq/q. Since ‖H‖gs⁢(X)=s⁢‖K‖F, (2) holds for every direction exactly when 2⁢(q+1)≤2⁢s.

It remains to check the ordinary self-concordance prerequisite at these scales. For fs=s⁢ϕ,

|D3⁢fs⁢(X)⁢[H,H,H]|=2⁢s⁢|Tr⁡(K3)|≤2⁢s⁢‖K‖F3≤2⁢s3/2⁢‖K‖F3=2⁢‖H‖gs⁢(X)3,

where |Tr⁡(K3)|≤∑i|ki|3≤(∑iki2)3/2, and s≥(q+1)/2≥1. Conversely, if s<(q+1)/2, the direction H=X/q already violates the SSC inequality. ∎

Proposition 4.1 and the SASC estimate in Section 3 prove Theorem 1.3. The extremal SSC direction also explains why (7) is an exact threshold: the derivative normalization is fixed, so reducing s makes that direction violate the inequality. SASC permits a different tradeoff because its admissible radius can change with a constant rescaling. If a family F is uniformly SASC with radius ρε, then, for any fixed c>0, the family c⁢F is uniformly SASC with radius c⁢ρε. Indeed, put u=r/c and g^=g¯/c. Then

r2n⁢(c⁢g+g¯)−1=u2n⁢(g+g^)−1,

and both the norm change and its bound for g multiply by c. In particular, under Definition 1.2, every s⁢G at a fixed matrix order is SASC with a radius that may depend on q and s. A necessary scaling law for SASC alone would therefore have to concern a profile s=sq with a radius uniform in q. The theorem supplies a linear sufficient profile; the lower bound proved here is the SSC obstruction.

The same scale can be expressed through the barrier’s gradient parameter. Relative to the Frobenius inner product, ∇fs⁢(X)=−s⁢X−1 and gs⁢(X)−1⁢(V)=s−1⁢X⁢V⁢X, so

⟨∇fs(X),gs(X)−1∇fs(X)⟩F=sq.

Thus the gradient parameter of −slogdet is s⁢q: it is n⁢q=Θ⁡(q3) at scale s=n, and equals n=Θ⁡(q2) at scale s=(q+1)/2. Thus the SASC guarantee is attained with gradient parameter equal to the ambient dimension. This is a consequence of the local scaling calculation; the mixing results cited in the introduction use their own additional geometric and acceptance estimates.

Appendix A Auxiliary Gaussian calculations

A.1 Polynomial Gaussian Poincaré inequality

Let ξ be a vector of independent standard Gaussian variables. For every polynomial f, the inequality Var⁡f⁡(ξ)≤E⁢‖∇f⁢(ξ)‖22 follows from a finite Hermite expansion. Define the probabilists’ Hermite polynomials by ez⁢u−z2/2=∑k≥0Hk⁢(u)⁢zk/k!. Gaussian integration of the product of two generating functions gives E⁢Hj⁢(ξ)⁢Hk⁢(ξ)=k!⁢δj⁢k, and differentiation in u gives Hk′=k⁢Hk−1. Their leading coefficients are one, so their products Hα form a basis for multivariate polynomials. If f⁡(ξ)=∑αaα⁢Hα⁢(ξ), orthogonality, independence, and the derivative identity imply

Var⁡f=∑|α|≥1aα2⁢α!,E⁢‖∇f‖22=∑|α|≥1|α|⁢aα2⁢α!.

Here α is a multi-index, |α|=∑iαi, and α!=∏iαi!. Comparison of the finite sums proves the required inequality.

A.2 Exact isotropic trace moments

For identity covariance, the coordinate independence permits exact formulas for the two moments in (16). These formulas distinguish the actual trace scale from the larger moments of the Frobenius norm and verify the q3 order used in the main proof. Let W have identity covariance in Sq. Its diagonal entries are independent N⁡(0,1), and its upper-triangular off-diagonal entries are independent N⁡(0,1/2), independently of the diagonal entries.

For an orthonormal basis (Fa), direct multiplication in the basis of Remark 2.2 gives

∑aFa⁢U⁢Fa=12⁢(U⊤+Tr⁡(U)⁢Iq)(U∈Rq×q).(23)

The left side is unchanged by an orthogonal change of basis. Applying (14) with B=(q+1)⁢Iq/2 and then (23), we obtain

E⁢Tr⁡(W4)=q⁢(q+1)22+12⁢∑a(Tr⁡(Fa2)+(Tr⁡Fa)2)
=q⁢(q+1)22+n+q2=q⁡(2⁢q2+5⁢q+5)4.

Here ∑aTr⁡(Fa2)=n, and Parseval applied to Iq gives ∑a(Tr⁡Fa)2=q.

For the cubic moment, write di=Wi⁢i, xi⁢j=Wi⁢j=xj⁢i for i≠j, and Si=∑j≠ixi⁢j2. Expanding the trace and using H3⁢(d)=d3−3⁢d gives

Tr⁡(W3)=∑iH3⁢(di)+3⁢∑idi⁢(1+Si)+6⁢∑i<j<kxi⁢j⁢xj⁢k⁢xk⁢i.

The expansion separates diagonal cubic fluctuations, diagonal terms weighted by off-diagonal squares, and triangle monomials. These three terms are pairwise orthogonal in Gaussian L2. The first two are orthogonal because E⁢H3⁢(di)=E⁡[di⁢H3⁢(di)]=0. Conditional on the off-diagonal entries, both have mean zero, while the triangle term is fixed. Within the second sum, distinct indices have zero cross moment because E⁢di⁢dj=0. Distinct triangle monomials in the last sum also have zero cross moment, since their product contains an edge variable to an odd power. Finally,

E⁢H3⁢(di)2=6,E⁢Si=q−12,Var⁡Si=q−12.

It follows that

E⁢(Tr⁡(W3))2=6⁢q+9⁢q⁢((q+1)24+q−12)+368⁢(q3)
=3⁢q⁢(4⁢q2+9⁢q+7)4,

where the triangle sum is empty for q<3.

The isotropic quartic moment is also the maximum over 0⪯Q⪯I. To see this without an entrywise comparison, let P⁡(A)=Tr⁡(A4). For symmetric A,H,

D2⁢P⁢(A)⁢[H,H]=8⁢Tr⁡(A2⁢H2)+4⁢Tr⁡(A⁢H⁢A⁢H)≥4⁢Tr⁡(A2⁢H2)≥0

by (10). Thus P is convex. Given a centered Gaussian A with covariance Q⪯I, take an independent centered Gaussian Y with covariance I−Q. Then A+Y has the same law as W, and conditional Jensen gives

E⁢Tr⁡(A4)≤E⁢Tr⁡((A+Y)4)=E⁢Tr⁡(W4).

This covariance monotonicity argument applies to the convex quartic trace. The cubic-trace second-moment bound used in the main proof instead follows from Gaussian Poincaré.

References

  • [1] Y. Kook and S. S. Vempala. Gaussian cooling and Dikin walks: The interior-point method for logconcave sampling. In Proceedings of the 37th Conference on Learning Theory (COLT), volume 247 of Proceedings of Machine Learning Research, pages 3137–3240, 2024. https://proceedings.mlr.press/v247/kook24b.html.
  • [2] A. Laddha, Y. T. Lee, and S. S. Vempala. Strong self-concordance and sampling. In Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing (STOC), pages 1212–1222, 2020. https://doi.org/10.1145/3357713.3384272.
  • [3] Z. Song and L. Zhang. A general framework for Metropolis-adjusted Dikin walks: Dimension-square mixing on polytopes and log-det walks on spectrahedra. arXiv:2608.25273v1, 2026. https://arxiv.org/abs/2608.25273v1.
  • [4] Y. Chen and Y. Kook. On two proofs of d2 mixing of weighted Dikin walks. arXiv:2608.28566v1, 2026. https://arxiv.org/abs/2608.28566v1.