Repository · Full text

Boundary reduction for growth-optimal e-variables:
a Gaussian theorem and a finite counterexample

Read PDF

HTML version 1 Added

Papers are listed without authors and are not intended for submission or formal publication.

Contents

Boundary reduction for growth-optimal e-variables:
a Gaussian theorem and a finite counterexample

Abstract

For a point null P0, a growth-optimal e-variable maximises the worst-case expected logarithm over the alternative. We study whether alternatives outside a neighbourhood of the null mean can be reduced to its boundary. In one-dimensional natural exponential families, the two endpoints of an excluded interval suffice. We show that the reduction fails for a minimal regular two-dimensional family on nine points with positive base masses and a circular excluded region in mean space. The infimum of D(Q∥P0) over finite exterior mixtures is strictly smaller than that over finite boundary mixtures. Remote tilts reconstruct nearly all of the null after mixing, whereas a bounded statistic separates every boundary mixture. For Gaussian location families with fixed positive-definite covariance and null mean zero, we prove boundary reduction for every bounded open neighbourhood of zero. There is a unique Borel probability measure on the boundary minimising the mixture divergence, and its value equals both finite-mixture infima. The mixture likelihood ratio with respect to P0 is growth-optimal over the full exterior. Gaussian translation makes its expected log likelihood ratio convex in the mean; a first-exit argument then propagates boundary optimality without convexity, connectedness, or smoothness assumptions on the excluded region.

Note. This paper was generated entirely by AI, including MiMo, using an automated research pipeline developed by Chenghua Liu and Hanyu Li.

1 Introduction

Consider testing a zero mean against all means outside a ball. The alternative contains a continuum of directions and distances, but its inner boundary has a simple interpretation: it contains the smallest allowed departures from the null in every direction. If the hardest alternatives could be confined to that boundary, constructing an optimal test would require searching only over directions. For growth-optimal e-variables, the question concerns the support of a least-favourable mixture, and this changes the geometry of the problem.

For a point null P0, an e-variable is a nonnegative statistic with null expectation at most one. Its absolute GROW criterion maximises the worst-case expected logarithm over the alternative [6]. Mixtures enter because the relative entropy D(Q∥P0) of any alternative mixture bounds this growth from above. For a natural exponential family {Pμ:μ∈M} in mean parametrisation, let I⁡(C) denote the infimum of this divergence over finite mixtures with means in C. If B is a neighbourhood of the null mean zero, boundary reduction asks whether

I⁡(M∖B)=I⁡(∂B).

The inequality from left to right is immediate. The reverse inequality must exclude the possibility that more distant components become less distinguishable from the null after they are mixed.

There is good reason to expect such a reduction. In one dimension, Grünwald et al. [7, Theorem 3] show that an optimal mixture on the two endpoints of an excluded interval remains optimal over its entire exterior. Their proof uses the variation-diminishing structure of exponential families [2]. They raise the corresponding multivariate support question for star-shaped excluded neighbourhoods; a related question for convex excluded regions appears in Grünwald et al. [8, pp. 1341–1342]. The issue is an exact support reduction, beyond the asymptotic analysis of KL balls in Grünwald et al. [7].

A componentwise argument is insufficient. Moving each member of a mixture toward the null changes how its likelihood ratio combines with the others. Convexity of relative entropy bounds the divergence of a mixture by the average component divergence, but it does not compare the two mixtures before and after this movement. We exhibit this obstruction in a two-dimensional family on nine points. Eight exposed vertices carry almost all of the null mass. Tilting far enough in their exposing directions produces nearly point masses, whose mixture reconstructs the null except for a small central atom. On a small circle around the null mean, however, a statistic comparing the two groups of vertices has strictly positive expectation under every family member. This sign survives every mixture on the circle. The resulting information values are positive and differ by more than three orders of magnitude (Theorem 3.1). All masses and bounds are explicit; a fixed finite exterior mixture is certified in Appendix A. The example uses the classical exposed-face limits of exponential families [5, 10], together with a local separator that prevents the same reconstruction at the boundary.

The Gaussian family clarifies what is missing from the general argument. A log-mixture likelihood ratio is convex in the observation for every natural exponential family. This alone says nothing about its expectation as the mean varies, since the averaging law varies too. For Gaussian shifts, all observations can instead be written as Z+μ, with the same centred Gaussian Z. Averaging then preserves convexity in μ. If Q∗ is the optimal boundary mixture and f⁡(μ)=EPμ⁢log⁡(q∗/p0), boundary optimality gives f≥c=D(Q∗∥P0) on ∂B, whereas f⁡(0)≤0<c. Convexity forces f to stay above c beyond the first boundary crossing along each ray. This proves exact reduction for every bounded open neighbourhood B, including sets with disconnected components or rays that leave and re-enter the set.

Theorem 4.1 establishes a unique Borel boundary prior and GROW⁡(M∖B)=I⁡(M∖B)=I⁡(∂B)>0. Its likelihood ratio has expected logarithm at least this value at every exterior mean, with equality on the prior’s support. The argument also yields an optimality gap certificate and quantitative finite-prior approximation. Here the support localisation is the additional conclusion: the underlying information-projection and duality framework is classical [3, 11, 4], with general GROW duality results given by Ram et al. [9]. The two parts of this paper identify why parameter-space proximity does not suffice in a finite exponential family and why Gaussian translation does suffice, even with a much less regular excluded region.

2 Mixture information and the GROW criterion

Throughout, we work with minimal regular d-dimensional natural exponential families (NEFs), with d≥1. Write {Pμ:μ∈M} for the family in mean parametrisation, with 0∈M and null P0. In natural coordinates we use P~θ, so that, after taking the null as the base law,

d⁢P~θd⁢P0⁢(y)=exp⁡{θT⁢y−ψ⁡(θ)},μ⁡(θ)=∇ψ⁢(θ),Pμ⁡(θ)=P~θ.(2.1)

Here the natural-parameter space Θ is open and contains zero, ψ⁡(θ)=log⁡EP0⁢exp⁡(θT⁢Y), and EP0⁢Y=0. Minimality makes the mean map one-to-one, with a smooth inverse on M=∇ψ⁢(Θ); see Brown [1]. The excluded region will be specified in mean coordinates, whereas the finite-support construction is estimated in natural coordinates. Keeping these parametrisations distinct is essential to the boundary comparison.

For a nonempty C⊆M, define

Mfin⁢(C):={∑j=1mwjPμj:m≥1,wj≥0,∑j=1mwj=1,μj∈C},(2.2)
I⁡(C):=infQ∈Mfin⁢(C)D(Q∥P0).(2.3)

All logarithms are natural. We use finite mixtures to define I; when constructing the Gaussian minimiser we pass to Borel priors on the compact boundary and prove equality with this infimum. All probability measures in the paper are countably additive.

A P0-e-variable is a measurable S≥0 satisfying EP0⁢S≤1. The absolute growth-rate optimality in the worst case (GROW) criterion is

GROW(C):=supS≥0:EP0⁢S≤1infμ∈CEPμlogS.(2.4)

The adjective “absolute” distinguishes this objective from criteria that subtract an alternative-dependent benchmark; see Grünwald et al. [6], Grünwald et al. [7]. We set log⁡0=−∞. In the present NEF setting, the positive part of log⁡S is integrable under every Pμ, so these expectations are well-defined in [−∞,∞). The lemma below verifies this integrability and the upper bound GROW⁡(C)≤I⁡(C).

Let B⊂M be a bounded open neighbourhood of zero with B¯ a compact subset of M, and put

A:=M∖B,Γ:=∂B.(2.5)

The boundary is Euclidean; the compact-containment assumption makes it also the relative boundary in M. The boundary-reduction question asks whether

I⁡(A)=I⁡(Γ).(2.6)

For comparison with the star-shaped setting, recall that B is 0-star-shaped when t⁢r∈B for every r∈B and 0≤t≤1. We impose no such assumption unless stated. The expected-log form of boundary reduction asks for a Borel probability W∗ on Γ whose mixture Q∗=∫ΓPr⁢d⁢W∗⁢(r), with c=D(Q∗∥P0), satisfies

EPμ⁢log⁡q∗p0≥c(μ∈A),EPr⁢log⁡q∗p0=cfor W∗-almost every r.(2.7)

Densities in likelihood ratios are taken with respect to a common dominating measure. The second equality expresses equalisation under the least-favourable prior. In the Gaussian case, continuity will strengthen it to equality throughout the prior’s support.

Lemma 2.1 (GROW–information upper bound).

For every nonempty C⊆M,

GROW⁡(C)≤I⁡(C).(2.8)
Proof.

First let R≪P0 have finite relative entropy, let u=d⁢R/d⁢P0, and let S be a P0-e-variable with m=EP0⁢S>0. On {u>0}, define T=S/(m⁢u). Then ER⁢T≤1. Moreover, log⁡u is R-integrable: its negative part has integral at most 1/e, since u⁢log⁡(1/u)≤1/e on 0<u<1, and its positive part is integrable because D(R∥P0)<∞. As (log⁡T)+≤T, the positive part of log⁡S=log⁡m+log⁡u+log⁡T is also R-integrable. The inequality log⁡T≤T−1, with the convention at zero, now gives

ERlogS=D(R∥P0)+logm+ERlogT≤D(R∥P0)+logm.(2.9)

This remains valid when the left-hand side is −∞.

Every member of the NEF is equivalent to P0, and regularity gives

D(P~θ∥P~ϑ)=(θ−ϑ)Tμ(θ)−ψ(θ)+ψ(ϑ)<∞.

Convexity of relative entropy therefore makes (2.9) applicable to every finite mixture. After discarding zero weights, for R=∑jwj⁢Pμj∈Mfin⁢(C) we have

infμ∈CEPμlogS≤∑jwjEPμjlogS=ERlogS≤D(R∥P0).

The finite sum is well-defined because its terms have integrable positive parts. If m=0, then S=0, P0-almost surely, and all component expectations are −∞. Taking the infimum over R, followed by the supremum over S, proves the claim and the integrability assertion following (2.4). ∎

3 Failure in a finite exponential family

We construct a family in which moving away from the null creates more room for cancellation between mixture components. The sample space has four axis points of radius one and four diagonal points of radius 4/5. Each is an exposed vertex, so sufficiently large tilts can isolate it. Equal weights on these eight limiting laws recover all but the central mass of the null. The difference between axis mass and diagonal mass provides the independent obstruction at a small mean boundary.

Set

ε=10−15,ρ=45⁢2,r0=10−2,(3.1)

and consider the nine-point alphabet

Y={(0,0),(±1,0),(0,±1),(±ρ,±ρ)},(3.2)

where the diagonal group includes all four sign choices. Let P0 assign mass ε to the origin and mass (1−ε)/8 to each of the other points. Symmetry gives EP0⁢Y=0. Define

p~θ⁢(y)=eθT⁢y−ψ⁡(θ)⁢p0⁢(y),θ∈R2,μ⁡(θ)=EP~θ⁢Y.(3.3)

The family is regular because its natural-parameter space is R2, and minimal because Y is not contained in an affine line. Its mean space is M=int⁡conv⁡(Y), and θ↦μ⁡(θ) is a diffeomorphism onto M [1]. The convex hull contains the diamond {x:|x1|+|x2|≤1}, whose Euclidean inradius is 1/2. Hence

B={μ∈M:‖μ‖2<r0}(3.4)

has closure compactly contained in M. In particular, the excluded region is smooth, strictly convex, and centrally symmetric.

Theorem 3.1 (Failure of boundary reduction).

For (3.3) and (3.4),

0<I⁡(M∖B)<2⋅10−15,I⁡(∂B)>7⋅10−12.(3.5)

No Borel probability on ∂B induces a mixture satisfying the global inequality in (2.7).

The separation must hold uniformly over a circle in mean space, while the tilts are explicit in natural coordinates. We therefore control the mean map as part of the separator estimate. Once this is done, the exterior bound follows from eight exposed-vertex limits. The atom at zero has a different role: its mass decreases under every exterior tilt, ensuring that the exterior infimum is strictly positive.

3.1 A uniform separator on the mean boundary

Define g⁡(0)=0, let g=1 at the four axis points, and let g=−1 at the four diagonal points. Thus ‖g‖∞=1 and EP0⁢g=0. For θ=(u,v), summing the nine masses gives

h⁡(θ):=EP~θ⁢g=(1−ε)⁢N⁢(u,v)8⁢ε+(1−ε)⁢Z8⁢(u,v),(3.6)

where

Z8⁢(u,v)=2⁢cosh⁡u+2⁢cosh⁡v+4⁢cosh⁡(ρ⁢u)⁢cosh⁡(ρ⁢v),(3.7)
N⁡(u,v)=2⁢cosh⁡u+2⁢cosh⁡v−4⁢cosh⁡(ρ⁢u)⁢cosh⁡(ρ⁢v).(3.8)

The statistic compares axis mass with diagonal mass. Its local behaviour is particularly simple: Since ρ2=8/25, symmetry and direct summation give

J0:=CovP0⁡(Y)=41100⁢(1−ε)⁢I2,h⁡(0)=0,∇h⁢(0)=0,∇2h⁢(0)=9100⁢(1−ε)⁢I2.(3.9)

For example, the first diagonal entry of ∇2h⁢(0) is EP0⁢[g⁢Y12]=(1−ε)⁢(1/4−ρ2/2); the mixed entry vanishes by symmetry. Writing h¯⁢(μ)=EPμ⁢g, the chain rule gives

∇2h¯⁢(0)=J0−1⁢∇2h⁢(0)⁢J0−1=9001681⁢(1−ε)⁢I2≻0.

The positive-definite Hessian explains the local separation in mean coordinates. To prove the stated numerical gap, however, a local expansion alone is insufficient. We next bound the remainder uniformly and locate the natural parameters corresponding to the entire mean circle.

Put δ=1/10 and s=u2+v2≤δ2. The series for (cosh⁡x−1)/x2 has nonnegative coefficients, so

coshx≥1+x22,cosh(ρx)≤1+κρ2x2(|x|≤δ),

where

κ:=cosh⁡(ρ⁢δ)−1(ρ⁢δ)2≤12+z224⁢(1−z2/30)<5011000,z2=(ρ⁢δ)2=2625.(3.10)

For the middle inequality, the tail starts at z2/24, and the ratio of successive tail terms is at most z2/30. Using u2⁢v2≤s2/4 now yields

N⁡(u,v)≥(1−4⁢κ⁢ρ2)⁢s−κ2⁢ρ4⁢s2
≥{1−4⁤5011000⁢825−(5011000)2⁢(825)2⁢1100}⁢s≥35100⁢s.(3.11)

For |x|<2, the bound (2⁢k)!≥2k implies cosh⁡x≤(1−x2/2)−1. Consequently,

Z8⁢(u,v)≤4⁤200199+4⁢(625624)2<9.(3.12)

The denominator in (3.6) is therefore less than nine, and

h⁡(θ)≥0.35⁢(1−ε)9⁢‖θ‖22(‖θ‖2≤0.1).(3.13)

The estimates so far hold on a ball in natural coordinates; the excluded region is defined in mean coordinates. The covariance matrix, as the derivative of the mean map, provides the required comparison. Since ‖Y‖2≤1, for ‖θ‖2≤0.1 we have θT⁢Y≥−0.1, ψ⁡(θ)≤0.1, and hence p~θ/p0≥e−0.2. For every unit vector ξ,

VarP~θ⁡(ξT⁢Y)=infb∈REP~θ⁢(ξT⁢Y−b)2
≥e−0.2⁢infb∈REP0⁢(ξT⁢Y−b)2=e−0.2⁢41100⁢(1−ε)>0.33.(3.14)

The last strict inequality follows without numerical evaluation from

e−1/5≥1−15+150−1750=307375,30737541100(1−ε)>33100.

Also ‖CovP~θ⁡(Y)‖op≤1, since EP~θ⁢‖Y‖22≤1. Integrating the derivative of the mean map along a ray gives, for 0<t≤0.1,

ξT⁢μ⁢(t⁢ξ)=∫0tVarP~a⁢ξ⁡(ξT⁢Y)⁢da>0.33⁢t,‖μ⁡(t⁢ξ)‖2≤t.(3.15)

The norm bound holds for every t≥0. The first expression is strictly increasing for all t≥0, because its derivative is positive: minimality and positive masses make ξT⁢Y nonconstant. Thus t≥0.1 implies ‖μ⁡(t⁢ξ)‖2>0.033. A mean on Γ, whose norm is 0.01, therefore has a natural parameter of norm less than 0.1 and, by the second bound in (3.15), at least r0. Substituting in (3.13), we obtain

infr∈ΓEPr⁢g≥η:=0.35⁢(1−ε)9⁢r02>3.8⋅10−6.(3.16)

For any Borel probability W on Γ, its mixture Q satisfies EQ⁢g≥η. With TV⁡(P,Q)=12⁢∑y∈Y|P⁡(y)−Q⁡(y)|, boundedness of g gives TV⁡(Q,P0)≥η/2. Pinsker’s inequality implies

D(Q∥P0)≥2TV(Q,P0)2≥η22>7⋅10−12.(3.17)

This uniform lower bound applies, in particular, to the infimum I⁡(Γ).

3.2 Remote tilts and the exterior value

Each of the eight nonzero points yj is an exposed vertex of conv⁡(Y). An axis point is exposed by the corresponding signed coordinate vector. The point (σ⁢ρ,τ⁢ρ), for σ,τ∈{−1,1}, is exposed by (σ,τ), because 2⁢ρ>1. Fix these normals and denote them by n1,…,n8. For y≠yj,

p~t⁢nj⁢(y)p~t⁢nj⁢(yj)=p0⁢(y)p0⁢(yj)⁢e−t⁢njT⁢(yj−y)⟶0.

It follows that

P~t⁢nj⟶δyj,μ(tnj)⟶yj(t→∞).(3.18)

The limiting means have norm one or 4/5, so all eight means are exterior for sufficiently large t. Hence the finite mixture

Qt:=18⁢∑j=18P~t⁢nj(3.19)

is then admissible for I⁡(A). It converges to the uniform law U8 on the nonzero points. Since P0 is positive at every point, relative entropy to P0 is continuous on the finite probability simplex. Thus

I(A)≤limt→∞D(Qt∥P0)=D(U8∥P0)=−log(1−ε)≤ε1−ε<2⋅10−15.(3.20)

The positive atom at the origin also ensures I⁡(A)>0. The global norm bound in (3.15) gives ‖μ⁡(θ)‖2≤‖θ‖2, so an exterior mean has ‖θ‖2≥r0. Along a unit direction ξ, the function t↦ψ⁡(t⁢ξ) is increasing for t>0. Integrating the first bound in (3.15) up to r0 therefore gives

ψ⁡(θ)≥ψ⁡(r0⁢θ/‖θ‖2)>0.33⁢∫0r0s⁢ds=0.165⁢r02(μ⁡(θ)∈A).

Every exterior mixture R thus has R⁡(0)≤ε⁢e−0.165⁢r02. Applying Pinsker’s inequality to the discrepancy at the origin yields

I⁡(A)≥2⁢ε2⁢(1−e−0.165⁢r02)2>0.(3.21)

Together with (3.17) and (3.20), this proves (3.5). An explicit finite choice is also available: Lemma A.1 in Appendix A proves that Q200 is exterior and D(Q200∥P0)<1.001⋅10−15.

Finally, suppose that a boundary mixture Q∗ satisfies the global inequality in (2.7), with c=D(Q∗∥P0). Every point has positive Q∗-mass. For sufficiently large t, averaging that inequality over the eight components of Qt and using the finite-alphabet KL identity gives

c≤EQtlogq∗p0=D(Qt∥P0)−D(Qt∥Q∗)≤D(Qt∥P0).

Letting t→∞ yields c≤−log⁡(1−ε)<2⋅10−15, contrary to (3.17). This completes the proof of Theorem 3.1.

The proof separates the null mass that remote tilts can reconstruct from the statistic that boundary mixtures cannot cancel. This gives a general criterion, with no dependence on the particular nine-point geometry.

Proposition 3.2 (Exposed-vertex obstruction).

Let P0 have finite support Y⊂Rd, positive mass at every point, and mean zero, and suppose that the generated NEF is minimal. Let B⊂M be an open neighbourhood of zero whose closure is a compact subset of M, and write A=M∖B, Γ=∂B. If V⊆Y is a nonempty set of exposed vertices of conv⁡(Y), and α=P0⁢(Y∈V), then

I⁡(A)≤−log⁡α.(3.22)

If a statistic g and constants L,η>0 satisfy

‖g‖∞≤L,EP0⁢g=0,infr∈ΓEPr⁢g≥η,(3.23)

then

I⁡(Γ)≥η22⁢L2.(3.24)

Consequently, −log⁡α<η2/(2⁢L2) implies strict failure of boundary reduction and rules out the global inequality in (2.7) for every boundary mixture.

Proof.

Choose an exposing normal nv for each v∈V. As in (3.18), P~t⁢nv→δv. Since every vertex lies outside M, the finite set V is disjoint from the compact set B¯, and all component means eventually lie in A. The mixtures

Rt:=∑v∈VP0⁢(Y=v)αP~t⁢nv⟶P0(⋅∣Y∈V)

then prove (3.22) by continuity of relative entropy. For any boundary mixture Q, the certificate gives TV⁡(Q,P0)≥η/(2⁢L), and Pinsker’s inequality proves (3.24).

If a boundary mixture Q∗ satisfied the global inequality, averaging it over Rt would give

D(Q∗∥P0)≤ERtlogq∗p0=D(Rt∥P0)−D(Rt∥Q∗)≤D(Rt∥P0).

All terms are finite because both P0 and Q∗ are positive on the finite alphabet. Passing to the limit would imply D(Q∗∥P0)≤−logα, contradicting the uniform boundary lower bound under the stated strict inequality. ∎

4 Gaussian boundary reduction

We now fix Pμ=Nd⁢(μ,Σ) with Σ≻0. In this family the same centred noise generates every shift. Consequently, the convexity of a log-mixture likelihood ratio survives expectation as a function of μ. This permits a boundary mixture to be compared with every exterior distribution at once, without moving its components or trying to compare their individual divergences.

Theorem 4.1 (Gaussian boundary reduction).

Let d≥1 and Pμ=Nd⁢(μ,Σ), with Σ≻0, and let B⊂Rd be any bounded open set containing zero. Set A=Rd∖B and Γ=∂B. There is a unique Borel probability W∗ on Γ minimising W↦D(∫ΓPrdW(r)∥P0). Its mixture and value,

Q∗=∫ΓPrdW∗(r),c=D(Q∗∥P0),(4.1)

satisfy 0<c<∞ and

GROW⁡(A)=I⁡(A)=I⁡(Γ)=c.(4.2)

Moreover,

EPμ⁢log⁡q∗p0≥c(μ∈A),(4.3)
EPr⁢log⁡q∗p0=c(r∈supp⁡W∗),(4.4)

and every R∈Mfin⁢(A) satisfies

D(R∥P0)≥D(R∥Q∗)+c.(4.5)

The e-variable S∗=q∗/p0 attains GROW⁡(A).

The minimiser is taken over Borel priors. The proof must therefore show both that its expected-log inequality holds on the full exterior and that its value agrees with the finite-mixture infima in (4.2).

4.1 The boundary minimiser and Gaussian convexity

Put K=Σ−1, write ‖v‖K2=vT⁢K⁢v, and let P⁡(Γ) be the space of Borel probabilities on the compact set Γ. For W∈P⁡(Γ), define

QW=∫ΓPrdW(r),LW=qWp0,F(W)=D(QW∥P0).(4.6)

The Gaussian likelihood ratio gives

LW⁢(x)=∫Γexp⁡{rT⁢K⁢x−12⁢rT⁢K⁢r}⁢dW⁢(r).(4.7)

Let R0=supr∈Γ‖r‖2, a=R0⁢‖K‖op, and b=R02⁢‖K‖op/2. Uniformly over all boundary priors,

e−a⁢‖x‖2−b≤LW⁢(x)≤ea⁢‖x‖2,|log⁡LW⁢(x)|≤a⁢‖x‖2+b.(4.8)

If Wn converges weakly to W, then LWn⁢(x)→LW⁢(x) for each x. Also,

p0⁢(x)⁢LWn⁢(x)⁢|log⁡LWn⁢(x)|≤p0⁢(x)⁢ea⁢‖x‖2⁢(a⁢‖x‖2+b),

whose right-hand side is integrable. Dominated convergence makes F weakly continuous. Weak compactness of P⁡(Γ) therefore gives a minimiser W∗.

Gaussian convolution is injective on probability measures. Writing R^⁢(t)=∫ei⁢tT⁢x⁢dR⁢(x) for a characteristic function, we have

Q^W(t)=e−tTΣt/2W^(t),t∈Rd,(4.9)

and the Gaussian factor never vanishes. Different priors thus induce different mixture distributions, and strict convexity of relative entropy in its first argument implies uniqueness of W∗. The envelope proves c=F⁡(W∗)<∞. If c=0, then Q∗=P0, and (4.9) would force W∗=δ0, which is impossible since 0∉Γ. Hence c>0.

This minimum agrees with the finite-mixture infimum. To see this directly, choose a finite δn-net in Γ, with δn>0 and δn→0, and map each r∈Γ to a nearest net point Tn⁢(r), resolving ties by a fixed ordering. These maps are Borel and supΓ‖Tn⁢(r)−r‖2≤δn. The image measures Wn=(Tn)#⁢W∗ have finite support and converge weakly to W∗. Continuity of F gives F⁡(Wn)→c. Since every finite prior belongs to P⁡(Γ),

I⁡(Γ)=c.(4.10)

For later use, define, for every boundary prior,

ϕW⁢(x)=log⁡LW⁢(x),fW⁢(μ)=EPμ⁢ϕW⁢(X).(4.11)

All these expectations are finite by (4.8). For a fixed observation x, let Wx be the probability on Γ with density

d⁢Wxd⁢W⁢(r)=exp⁡{rT⁢K⁢x−rT⁢K⁢r/2}LW⁢(x).

Differentiating (4.7), justified by compactness of Γ, gives

∇ϕW⁢(x)=K⁢EWx⁢r,∇2ϕW⁢(x)=K⁢CovWx⁡(r)⁢K⪰0.(4.12)

Thus ϕW is convex and globally a-Lipschitz. Writing X=Z+μ, where Z∼Nd⁢(0,Σ), shows that

fW⁢(μ)=E⁢ϕW⁢(Z+μ)(4.13)

is likewise convex and a-Lipschitz. This is the step specific to Gaussian translation: convexity of the log-mixture likelihood ratio becomes convexity of its expected value in the mean parameter.

4.2 From boundary optimality to the exterior

Adding an infinitesimal point mass to a prior probes whether the corresponding expected log likelihood is below the prior’s own average. At a minimum it cannot be. This variational condition gives the boundary inequality to which the convexity argument will be applied; for a nonoptimal prior it measures the remaining gap.

Proposition 4.2 (Boundary optimality and gap bound).

Write cW=F⁡(W). A prior W∈P⁡(Γ) minimises F if and only if

fW⁢(r)≥cW(r∈Γ).(4.14)

For a minimiser, equality holds at every point of supp⁡W. For any W∈P⁡(Γ),

0≤F⁡(W)−c≤cW−minr∈Γ⁡fW⁢(r).(4.15)
Proof.

Fix r∈Γ, let Lr=pr/p0, and set Lα=(1−α)⁢LW+α⁢Lr. For 0≤α≤1, (4.8) gives the integrable bound

|p0⁢(x)⁢(Lr⁢(x)−LW⁢(x))⁢(1+log⁡Lα⁢(x))|≤2⁢p0⁢(x)⁢ea⁢‖x‖2⁢(1+a⁢‖x‖2+b).

Differentiation under the integral, together with ∫(pr−qW)=0, therefore yields

dd⁢α|α=0+⁢F⁢((1−α)⁢W+α⁢δr)=fW⁢(r)−cW.(4.16)

At a minimum this derivative is nonnegative, proving necessity.

Conversely, for V∈P⁡(Γ), the envelope makes all terms in the following KL identity finite:

F(V)−F(W)=D(QV∥QW)+∫ΓfW(r)dV(r)−cW.(4.17)

Condition (4.14) makes the right-hand side nonnegative, so it is sufficient. Fubini’s theorem gives ∫ΓfW⁢dW=cW; its use is justified by the uniform linear growth bound and the uniformly bounded first moments of the components. Since fW is continuous, fW≥cW and this average equality imply fW=cW throughout supp⁡W.

Finally, set V=W∗ in (4.17). It gives

c=D(Q∗∥QW)+∫ΓfW(r)dW∗(r)≥minr∈ΓfW(r).

Subtracting from cW proves (4.15). ∎

For the minimiser, abbreviate ϕ=ϕW∗ and f=fW∗. The proposition gives f≥c on Γ, with equality on supp⁡W∗. Boundary reduction now becomes a statement about this one function: its value cannot fall below c after leaving B. The following lemma shows why a convex function that starts below its boundary values has precisely this property.

Lemma 4.3 (First-exit propagation).

Let B⊂Rd be open and contain zero. Suppose that g:Rd→R is convex along each ray from zero, meaning that t↦g⁡(t⁢u) is convex on [0,∞) for every u∈Rd. If

g⁡(0)≤c≤infr∈∂Bg⁡(r),

then g⁡(μ)≥c for every μ∉B.

Proof.

Fix μ∉B. The set {s∈[0,1]:s⁢μ∉B} is nonempty and closed. Because B is a neighbourhood of zero, its minimum τ lies in (0,1]. Every s<τ has s⁢μ∈B, so r=τ⁢μ belongs to ∂B. Put λ=1/τ≥1. Convexity on the ray, applied to r=(1−1/λ)⁢0+(1/λ)⁢(λ⁢r), gives

g⁡(μ)=g⁡(λ⁢r)≥g⁡(r)+(λ−1)⁢{g⁡(r)−g⁡(0)}≥c.

This also covers λ=1. ∎

For the Gaussian boundary minimiser, (4.13) makes f convex, and (4.8) gives

f(0)=EP0logq∗p0=−D(P0∥Q∗)≤0<c.(4.18)

Lemma 4.3 therefore proves (4.3). A ray may leave B and later re-enter it; only its first exit is used. This explains why no star-shapedness or connectedness assumption is needed.

Let R∈Mfin⁢(A), with density rR. The global inequality gives ER⁢ϕ≥c. Both log⁡(rR/p0) and ϕ are bounded in absolute value by CR⁢(1+‖x‖2) for some finite constant CR: the former has the same Gaussian-mixture representation as (4.7), with a finite set of means. Since R has finite first moment, the KL identity is legitimate and gives

D(R∥P0)=D(R∥Q∗)+ERϕ≥D(R∥Q∗)+c.(4.19)

This proves (4.5) and I⁡(A)≥c. As Γ⊂A, (4.10) supplies the reverse inequality, so I⁡(A)=I⁡(Γ)=c.

Finally, S∗=q∗/p0 has null expectation one and integrable logarithm under every Gaussian shift. By (4.3), its worst-case expected logarithm is at least c. The nonempty set supp⁡W∗ lies in A, and (4.4) gives equality at every point of this support, so the worst-case value is exactly c. Lemma 2.1 now yields c≤GROW⁡(A)≤I⁡(A)=c, completing the proof of Theorem 4.1.

4.3 Quantitative approximation by finite priors

Approximating the optimal prior on a finite boundary grid perturbs both the mixture and its expected log likelihood under the original optimum. The first perturbation costs quadratic Gaussian relative entropy. The second is generally linear in the grid spacing, but vanishes when the grid stays inside the optimal support, where f=c. This distinction gives the two rates below.

Corollary 4.4.

Let T:Γ→Γ be a finite-valued Borel map satisfying supr∈Γ‖T⁡(r)−r‖2≤δ, and put WT=T#⁢W∗. Then

0≤F⁡(WT)−c≤12⁢‖K‖op⁢δ2+R0⁢‖K‖op⁢δ.(4.20)

For every δ>0, there is a finite prior Vδ supported on supp⁡W∗ such that

0≤F(Vδ)−c=D(QVδ∥Q∗)≤12∥K∥opδ2.(4.21)

These priors can be chosen with Vδ⇒W∗ as δ↓0.

Proof.

Consider the joint distributions obtained by drawing r∼W∗ and then drawing either X∼PT⁡(r) or X∼Pr. Marginalising out r cannot increase relative entropy. Equivalently, apply the log-sum inequality to this common mixing measure. The Gaussian KL formula therefore gives

D(QWT∥Q∗)≤∫ΓD(PT⁡(r)∥Pr)dW∗(r)=12∫Γ∥T(r)−r∥K2dW∗(r)≤12∥K∥opδ2.(4.22)

Couple XT=Z+T⁡(U) and X=Z+U, where U∼W∗ and Z∼Nd⁢(0,Σ) are independent. The Lipschitz bound for ϕ yields

|EQWT⁢ϕ−EQ∗⁢ϕ|≤R0⁢‖K‖op⁢δ.

Combining this with the KL identity

F(WT)−c=D(QWT∥Q∗)+EQWTϕ−c

proves (4.20); nonnegativity follows from minimality of W∗.

For the sharper bound, take a finite δ-net in the compact set supp⁡W∗, and let Tδ be a Borel nearest-net-point map on that support. Set Vδ=(Tδ)#⁢W∗. Support equalisation gives

EQVδ⁢ϕ=∫f⁡(Tδ⁢(r))⁢d⁢W∗⁢(r)=c.

The KL identity now gives equality in (4.21), and (4.22), applied on supp⁡W∗, supplies its upper bound. Finally, ‖Tδ⁢(r)−r‖2≤δ implies weak convergence to W∗, by uniform continuity of continuous functions on the compact support. ∎

For numerical optimisation, Proposition 4.2 and Corollary 4.4 play different roles. The former certifies a candidate once its Gaussian integrals and the minimum of fW over the whole boundary have been bounded. The latter ensures finite-prior approximation, but its sharper rate uses the unknown optimal support. Neither the size of a sufficient grid nor finite support of W∗ is determined by these results. Beyond Gaussian shifts, the proof suggests looking for families that preserve raywise convexity of expected log-mixture likelihood ratios: the first-exit lemma needs only that property, whereas the finite example shows that observation-space convexity by itself is insufficient.

Appendix A An explicit finite exterior mixture

The limit in (3.20) proves the counterexample without a numerical computation. This appendix gives a fixed finite mixture with the same separation, using only analytic bounds, and records its exact masses. The constants and normals are those of Section 3.

Lemma A.1.

All eight component means of Q200 lie in M∖B, and

D(Q200∥P0)<1.001⋅10−15.(A.1)
Proof.

For an axis normal the exposure gap is 1−ρ, and for a diagonal normal it is 2⁢ρ−1. Since 1/2<ρ<2/3, the smallest gap is

γ:=minj⁡miny≠yj⁢njT⁢(yj−y)=2⁢ρ−1>13100.

The last inequality follows by squaring ρ>113/200, using ρ2=8/25>(113/200)2. Divide every unnormalised tilted mass by the mass at its target vertex yj. The seven other nonzero base masses have ratio one, and the origin has ratio 8⁢ε/(1−ε)<1. Hence

1−P~t⁢nj⁢{yj}≤βt:=8⁢e−t⁢γ.(A.2)

Since all alphabet points have norm at most one,

‖μ⁡(t⁢nj)−yj‖2≤2⁢βt.

To bound β200, note that e2>1+2+2+4/3+2/3=7. Thus

β200<8⁢e−26<8713<10−10,713=96 889 010 407.(A.3)

Consequently every component mean has norm at least 4/5−2⁢β200>0.79>r0, proving exterior feasibility.

Convexity of total variation and (A.2) give TV⁡(Q200,U8)≤β200. Put w=(1−ε)/8. For each nonzero point y,

|Q200⁢(y)−w|≤β200+ε8.

Jensen’s inequality and EP0⁢Y=0 give ψ⁡(θ)≥0, so P~θ⁢(0)=ε⁢e−ψ⁡(θ)≤ε. In particular, |Q200⁢(0)−ε|≤ε. The inequality D(P∥Q)≤χ2(P∥Q), which follows by applying log⁡x≤x−1 pointwise, now yields

D(Q200∥P0)≤∑y∈Y(Q200⁢(y)−P0⁢(y))2P0⁢(y)
≤ε+641−ε⁢(β200+ε8)2
<10−15+65⁢(111011)2<1.001⋅10−15.(A.4)

In the last line, 64/(1−ε)<65 and β200+ε/8<11/1011. ∎

For an exact expression for Qt, put w=(1−ε)/8. Symmetry leaves two component normalisers:

Za⁢(t)=ε+w⁡{2⁢cosh⁡t+2+4⁢cosh⁡(ρ⁢t)},(A.5)
Zd⁢(t)=ε+w⁡{4⁢cosh⁡t+2⁢cosh⁡(2⁢ρ⁢t)+2}.(A.6)

The origin mass q0⁢(t), the mass qa⁢(t) at each axis point, and the mass qd⁢(t) at each diagonal point are

q0⁢(t)=ε2⁢{1Za⁢(t)+1Zd⁢(t)},(A.7)
qa⁢(t)=w8⁢{2⁢cosh⁡t+2Za⁢(t)+4⁢cosh⁡tZd⁢(t)},(A.8)
qd⁢(t)=w8⁢{4⁢cosh⁡(ρ⁢t)Za⁢(t)+2⁢cosh⁡(2⁢ρ⁢t)+2Zd⁢(t)}.(A.9)

Thus

D(Qt∥P0)=q0(t)logq0⁢(t)ε+4qa(t)logqa⁢(t)w+4qd(t)logqd⁢(t)w.(A.10)

For t≥0, the norms of the means of an axis-normal and a diagonal-normal component, respectively, are

ma⁢(t)=w⁡{2⁢sinh⁡t+4⁢ρ⁢sinh⁡(ρ⁢t)}Za⁢(t),(A.11)
md⁢(t)=2⁢w⁢{2⁢sinh⁡t+2⁢ρ⁢sinh⁡(2⁢ρ⁢t)}Zd⁢(t).(A.12)

The mass formulas average the eight components in (3.19). The mean formulas sum y⁢p~t⁢nj⁢(y) for the representative normals nj=(1,0) and nj=(1,1).

References

  • [1] L. D. Brown. Fundamentals of Statistical Exponential Families with Applications in Statistical Decision Theory. Institute of Mathematical Statistics Lecture Notes–Monograph Series, vol. 9, 1986. https://doi.org/10.1214/lnms/1215466757.
  • [2] L. D. Brown, I. M. Johnstone, and K. B. MacGibbon. Variation diminishing transformations: A direct approach to total positivity and its statistical applications. Journal of the American Statistical Association, 76(376):824–832, 1981. https://doi.org/10.1080/01621459.1981.10477730.
  • [3] I. Csiszár. I-divergence geometry of probability distributions and minimization problems. The Annals of Probability, 3(1):146–158, 1975. https://doi.org/10.1214/aop/1176996454.
  • [4] I. Csiszár and F. Matúš. Information projections revisited. IEEE Transactions on Information Theory, 49(6):1474–1490, 2003. https://doi.org/10.1109/TIT.2003.810633.
  • [5] I. Csiszár and F. Matúš. Closures of exponential families. The Annals of Probability, 33(2):582–600, 2005. https://doi.org/10.1214/009117904000000766.
  • [6] P. Grünwald, R. de Heide, and W. Koolen. Safe testing. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86(5):1091–1128, 2024. https://doi.org/10.1093/jrsssb/qkae011.
  • [7] P. Grünwald, Y. Hao, and A. Balsubramani. Growth-optimal e-variables and an extension to the multivariate Csiszár–Sanov–Chernoff theorem. arXiv:2412.17554v2, 2024. https://arxiv.org/abs/2412.17554.
  • [8] P. Grünwald, A. Ramdas, R. Wang, and J. Ziegel, organisers. Game-theoretic statistical inference: Optional sampling, universal inference, and multiple testing based on e-values. Oberwolfach Reports, 21(2):1339–1386, Report No. 24/2024, 2024. https://doi.org/10.4171/OWR/2024/24.
  • [9] A. Ram, M. Larsson, J. Ruf, and A. Ramdas. Strong duality for the GROW criterion. arXiv:2606.24768v2, 2026. https://arxiv.org/abs/2606.24768.
  • [10] A. Rinaldo, S. E. Fienberg, and Y. Zhou. On the geometry of discrete exponential families with application to exponential random graph models. Electronic Journal of Statistics, 3:446–484, 2009. https://doi.org/10.1214/08-EJS350.
  • [11] F. Topsøe. Information-theoretical optimization techniques. Kybernetika, 15(1):8–27, 1979. https://www.kybernetika.cz/content/1979/1/8/paper.pdf.