Boundary reduction for growth-optimal e-variables:
a Gaussian theorem and a finite counterexample
Abstract
For a point null , a growth-optimal e-variable maximises the worst-case expected logarithm over the alternative. We study whether alternatives outside a neighbourhood of the null mean can be reduced to its boundary. In one-dimensional natural exponential families, the two endpoints of an excluded interval suffice. We show that the reduction fails for a minimal regular two-dimensional family on nine points with positive base masses and a circular excluded region in mean space. The infimum of over finite exterior mixtures is strictly smaller than that over finite boundary mixtures. Remote tilts reconstruct nearly all of the null after mixing, whereas a bounded statistic separates every boundary mixture. For Gaussian location families with fixed positive-definite covariance and null mean zero, we prove boundary reduction for every bounded open neighbourhood of zero. There is a unique Borel probability measure on the boundary minimising the mixture divergence, and its value equals both finite-mixture infima. The mixture likelihood ratio with respect to is growth-optimal over the full exterior. Gaussian translation makes its expected log likelihood ratio convex in the mean; a first-exit argument then propagates boundary optimality without convexity, connectedness, or smoothness assumptions on the excluded region.
Note. This paper was generated entirely by AI, including MiMo, using an automated research pipeline developed by Chenghua Liu and Hanyu Li.
1 Introduction
Consider testing a zero mean against all means outside a ball. The alternative contains a continuum of directions and distances, but its inner boundary has a simple interpretation: it contains the smallest allowed departures from the null in every direction. If the hardest alternatives could be confined to that boundary, constructing an optimal test would require searching only over directions. For growth-optimal e-variables, the question concerns the support of a least-favourable mixture, and this changes the geometry of the problem.
For a point null , an e-variable is a nonnegative statistic with null expectation at most one. Its absolute GROW criterion maximises the worst-case expected logarithm over the alternative [6]. Mixtures enter because the relative entropy of any alternative mixture bounds this growth from above. For a natural exponential family in mean parametrisation, let denote the infimum of this divergence over finite mixtures with means in . If is a neighbourhood of the null mean zero, boundary reduction asks whether
The inequality from left to right is immediate. The reverse inequality must exclude the possibility that more distant components become less distinguishable from the null after they are mixed.
There is good reason to expect such a reduction. In one dimension, Grünwald et al. [7, Theorem 3] show that an optimal mixture on the two endpoints of an excluded interval remains optimal over its entire exterior. Their proof uses the variation-diminishing structure of exponential families [2]. They raise the corresponding multivariate support question for star-shaped excluded neighbourhoods; a related question for convex excluded regions appears in Grünwald et al. [8, pp. 1341–1342]. The issue is an exact support reduction, beyond the asymptotic analysis of KL balls in Grünwald et al. [7].
A componentwise argument is insufficient. Moving each member of a mixture toward the null changes how its likelihood ratio combines with the others. Convexity of relative entropy bounds the divergence of a mixture by the average component divergence, but it does not compare the two mixtures before and after this movement. We exhibit this obstruction in a two-dimensional family on nine points. Eight exposed vertices carry almost all of the null mass. Tilting far enough in their exposing directions produces nearly point masses, whose mixture reconstructs the null except for a small central atom. On a small circle around the null mean, however, a statistic comparing the two groups of vertices has strictly positive expectation under every family member. This sign survives every mixture on the circle. The resulting information values are positive and differ by more than three orders of magnitude (Theorem 3.1). All masses and bounds are explicit; a fixed finite exterior mixture is certified in Appendix A. The example uses the classical exposed-face limits of exponential families [5, 10], together with a local separator that prevents the same reconstruction at the boundary.
The Gaussian family clarifies what is missing from the general argument. A log-mixture likelihood ratio is convex in the observation for every natural exponential family. This alone says nothing about its expectation as the mean varies, since the averaging law varies too. For Gaussian shifts, all observations can instead be written as , with the same centred Gaussian . Averaging then preserves convexity in . If is the optimal boundary mixture and , boundary optimality gives on , whereas . Convexity forces to stay above beyond the first boundary crossing along each ray. This proves exact reduction for every bounded open neighbourhood , including sets with disconnected components or rays that leave and re-enter the set.
Theorem 4.1 establishes a unique Borel boundary prior and . Its likelihood ratio has expected logarithm at least this value at every exterior mean, with equality on the prior’s support. The argument also yields an optimality gap certificate and quantitative finite-prior approximation. Here the support localisation is the additional conclusion: the underlying information-projection and duality framework is classical [3, 11, 4], with general GROW duality results given by Ram et al. [9]. The two parts of this paper identify why parameter-space proximity does not suffice in a finite exponential family and why Gaussian translation does suffice, even with a much less regular excluded region.
2 Mixture information and the GROW criterion
Throughout, we work with minimal regular -dimensional natural exponential families (NEFs), with . Write for the family in mean parametrisation, with and null . In natural coordinates we use , so that, after taking the null as the base law,
| (2.1) |
Here the natural-parameter space is open and contains zero, , and . Minimality makes the mean map one-to-one, with a smooth inverse on ; see Brown [1]. The excluded region will be specified in mean coordinates, whereas the finite-support construction is estimated in natural coordinates. Keeping these parametrisations distinct is essential to the boundary comparison.
For a nonempty , define
| (2.2) | ||||
| (2.3) |
All logarithms are natural. We use finite mixtures to define ; when constructing the Gaussian minimiser we pass to Borel priors on the compact boundary and prove equality with this infimum. All probability measures in the paper are countably additive.
A -e-variable is a measurable satisfying . The absolute growth-rate optimality in the worst case (GROW) criterion is
| (2.4) |
The adjective “absolute” distinguishes this objective from criteria that subtract an alternative-dependent benchmark; see Grünwald et al. [6], Grünwald et al. [7]. We set . In the present NEF setting, the positive part of is integrable under every , so these expectations are well-defined in . The lemma below verifies this integrability and the upper bound .
Let be a bounded open neighbourhood of zero with a compact subset of , and put
| (2.5) |
The boundary is Euclidean; the compact-containment assumption makes it also the relative boundary in . The boundary-reduction question asks whether
| (2.6) |
For comparison with the star-shaped setting, recall that is -star-shaped when for every and . We impose no such assumption unless stated. The expected-log form of boundary reduction asks for a Borel probability on whose mixture , with , satisfies
| (2.7) |
Densities in likelihood ratios are taken with respect to a common dominating measure. The second equality expresses equalisation under the least-favourable prior. In the Gaussian case, continuity will strengthen it to equality throughout the prior’s support.
Lemma 2.1 (GROW–information upper bound).
For every nonempty ,
| (2.8) |
Proof.
First let have finite relative entropy, let , and let be a -e-variable with . On , define . Then . Moreover, is -integrable: its negative part has integral at most , since on , and its positive part is integrable because . As , the positive part of is also -integrable. The inequality , with the convention at zero, now gives
| (2.9) |
This remains valid when the left-hand side is .
Every member of the NEF is equivalent to , and regularity gives
Convexity of relative entropy therefore makes (2.9) applicable to every finite mixture. After discarding zero weights, for we have
The finite sum is well-defined because its terms have integrable positive parts. If , then , -almost surely, and all component expectations are . Taking the infimum over , followed by the supremum over , proves the claim and the integrability assertion following (2.4). ∎
3 Failure in a finite exponential family
We construct a family in which moving away from the null creates more room for cancellation between mixture components. The sample space has four axis points of radius one and four diagonal points of radius . Each is an exposed vertex, so sufficiently large tilts can isolate it. Equal weights on these eight limiting laws recover all but the central mass of the null. The difference between axis mass and diagonal mass provides the independent obstruction at a small mean boundary.
Set
| (3.1) |
and consider the nine-point alphabet
| (3.2) |
where the diagonal group includes all four sign choices. Let assign mass to the origin and mass to each of the other points. Symmetry gives . Define
| (3.3) |
The family is regular because its natural-parameter space is , and minimal because is not contained in an affine line. Its mean space is , and is a diffeomorphism onto [1]. The convex hull contains the diamond , whose Euclidean inradius is . Hence
| (3.4) |
has closure compactly contained in . In particular, the excluded region is smooth, strictly convex, and centrally symmetric.
Theorem 3.1 (Failure of boundary reduction).
The separation must hold uniformly over a circle in mean space, while the tilts are explicit in natural coordinates. We therefore control the mean map as part of the separator estimate. Once this is done, the exterior bound follows from eight exposed-vertex limits. The atom at zero has a different role: its mass decreases under every exterior tilt, ensuring that the exterior infimum is strictly positive.
3.1 A uniform separator on the mean boundary
Define , let at the four axis points, and let at the four diagonal points. Thus and . For , summing the nine masses gives
| (3.6) |
where
| (3.7) | ||||
| (3.8) |
The statistic compares axis mass with diagonal mass. Its local behaviour is particularly simple: Since , symmetry and direct summation give
| (3.9) |
For example, the first diagonal entry of is ; the mixed entry vanishes by symmetry. Writing , the chain rule gives
The positive-definite Hessian explains the local separation in mean coordinates. To prove the stated numerical gap, however, a local expansion alone is insufficient. We next bound the remainder uniformly and locate the natural parameters corresponding to the entire mean circle.
Put and . The series for has nonnegative coefficients, so
where
| (3.10) |
For the middle inequality, the tail starts at , and the ratio of successive tail terms is at most . Using now yields
| (3.11) |
For , the bound implies . Consequently,
| (3.12) |
The denominator in (3.6) is therefore less than nine, and
| (3.13) |
The estimates so far hold on a ball in natural coordinates; the excluded region is defined in mean coordinates. The covariance matrix, as the derivative of the mean map, provides the required comparison. Since , for we have , , and hence . For every unit vector ,
| (3.14) |
The last strict inequality follows without numerical evaluation from
Also , since . Integrating the derivative of the mean map along a ray gives, for ,
| (3.15) |
The norm bound holds for every . The first expression is strictly increasing for all , because its derivative is positive: minimality and positive masses make nonconstant. Thus implies . A mean on , whose norm is , therefore has a natural parameter of norm less than and, by the second bound in (3.15), at least . Substituting in (3.13), we obtain
| (3.16) |
For any Borel probability on , its mixture satisfies . With , boundedness of gives . Pinsker’s inequality implies
| (3.17) |
This uniform lower bound applies, in particular, to the infimum .
3.2 Remote tilts and the exterior value
Each of the eight nonzero points is an exposed vertex of . An axis point is exposed by the corresponding signed coordinate vector. The point , for , is exposed by , because . Fix these normals and denote them by . For ,
It follows that
| (3.18) |
The limiting means have norm one or , so all eight means are exterior for sufficiently large . Hence the finite mixture
| (3.19) |
is then admissible for . It converges to the uniform law on the nonzero points. Since is positive at every point, relative entropy to is continuous on the finite probability simplex. Thus
| (3.20) |
The positive atom at the origin also ensures . The global norm bound in (3.15) gives , so an exterior mean has . Along a unit direction , the function is increasing for . Integrating the first bound in (3.15) up to therefore gives
Every exterior mixture thus has . Applying Pinsker’s inequality to the discrepancy at the origin yields
| (3.21) |
Together with (3.17) and (3.20), this proves (3.5). An explicit finite choice is also available: Lemma A.1 in Appendix A proves that is exterior and .
Finally, suppose that a boundary mixture satisfies the global inequality in (2.7), with . Every point has positive -mass. For sufficiently large , averaging that inequality over the eight components of and using the finite-alphabet KL identity gives
Letting yields , contrary to (3.17). This completes the proof of Theorem 3.1.
The proof separates the null mass that remote tilts can reconstruct from the statistic that boundary mixtures cannot cancel. This gives a general criterion, with no dependence on the particular nine-point geometry.
Proposition 3.2 (Exposed-vertex obstruction).
Let have finite support , positive mass at every point, and mean zero, and suppose that the generated NEF is minimal. Let be an open neighbourhood of zero whose closure is a compact subset of , and write , . If is a nonempty set of exposed vertices of , and , then
| (3.22) |
If a statistic and constants satisfy
| (3.23) |
then
| (3.24) |
Consequently, implies strict failure of boundary reduction and rules out the global inequality in (2.7) for every boundary mixture.
Proof.
Choose an exposing normal for each . As in (3.18), . Since every vertex lies outside , the finite set is disjoint from the compact set , and all component means eventually lie in . The mixtures
then prove (3.22) by continuity of relative entropy. For any boundary mixture , the certificate gives , and Pinsker’s inequality proves (3.24).
If a boundary mixture satisfied the global inequality, averaging it over would give
All terms are finite because both and are positive on the finite alphabet. Passing to the limit would imply , contradicting the uniform boundary lower bound under the stated strict inequality. ∎
4 Gaussian boundary reduction
We now fix with . In this family the same centred noise generates every shift. Consequently, the convexity of a log-mixture likelihood ratio survives expectation as a function of . This permits a boundary mixture to be compared with every exterior distribution at once, without moving its components or trying to compare their individual divergences.
Theorem 4.1 (Gaussian boundary reduction).
Let and , with , and let be any bounded open set containing zero. Set and . There is a unique Borel probability on minimising . Its mixture and value,
| (4.1) |
satisfy and
| (4.2) |
Moreover,
| (4.3) | ||||||
| (4.4) |
and every satisfies
| (4.5) |
The e-variable attains .
The minimiser is taken over Borel priors. The proof must therefore show both that its expected-log inequality holds on the full exterior and that its value agrees with the finite-mixture infima in (4.2).
4.1 The boundary minimiser and Gaussian convexity
Put , write , and let be the space of Borel probabilities on the compact set . For , define
| (4.6) |
The Gaussian likelihood ratio gives
| (4.7) |
Let , , and . Uniformly over all boundary priors,
| (4.8) |
If converges weakly to , then for each . Also,
whose right-hand side is integrable. Dominated convergence makes weakly continuous. Weak compactness of therefore gives a minimiser .
Gaussian convolution is injective on probability measures. Writing for a characteristic function, we have
| (4.9) |
and the Gaussian factor never vanishes. Different priors thus induce different mixture distributions, and strict convexity of relative entropy in its first argument implies uniqueness of . The envelope proves . If , then , and (4.9) would force , which is impossible since . Hence .
This minimum agrees with the finite-mixture infimum. To see this directly, choose a finite -net in , with and , and map each to a nearest net point , resolving ties by a fixed ordering. These maps are Borel and . The image measures have finite support and converge weakly to . Continuity of gives . Since every finite prior belongs to ,
| (4.10) |
For later use, define, for every boundary prior,
| (4.11) |
All these expectations are finite by (4.8). For a fixed observation , let be the probability on with density
Differentiating (4.7), justified by compactness of , gives
| (4.12) |
Thus is convex and globally -Lipschitz. Writing , where , shows that
| (4.13) |
is likewise convex and -Lipschitz. This is the step specific to Gaussian translation: convexity of the log-mixture likelihood ratio becomes convexity of its expected value in the mean parameter.
4.2 From boundary optimality to the exterior
Adding an infinitesimal point mass to a prior probes whether the corresponding expected log likelihood is below the prior’s own average. At a minimum it cannot be. This variational condition gives the boundary inequality to which the convexity argument will be applied; for a nonoptimal prior it measures the remaining gap.
Proposition 4.2 (Boundary optimality and gap bound).
Write . A prior minimises if and only if
| (4.14) |
For a minimiser, equality holds at every point of . For any ,
| (4.15) |
Proof.
Fix , let , and set . For , (4.8) gives the integrable bound
Differentiation under the integral, together with , therefore yields
| (4.16) |
At a minimum this derivative is nonnegative, proving necessity.
Conversely, for , the envelope makes all terms in the following KL identity finite:
| (4.17) |
Condition (4.14) makes the right-hand side nonnegative, so it is sufficient. Fubini’s theorem gives ; its use is justified by the uniform linear growth bound and the uniformly bounded first moments of the components. Since is continuous, and this average equality imply throughout .
For the minimiser, abbreviate and . The proposition gives on , with equality on . Boundary reduction now becomes a statement about this one function: its value cannot fall below after leaving . The following lemma shows why a convex function that starts below its boundary values has precisely this property.
Lemma 4.3 (First-exit propagation).
Let be open and contain zero. Suppose that is convex along each ray from zero, meaning that is convex on for every . If
then for every .
Proof.
Fix . The set is nonempty and closed. Because is a neighbourhood of zero, its minimum lies in . Every has , so belongs to . Put . Convexity on the ray, applied to , gives
This also covers . ∎
For the Gaussian boundary minimiser, (4.13) makes convex, and (4.8) gives
| (4.18) |
Lemma 4.3 therefore proves (4.3). A ray may leave and later re-enter it; only its first exit is used. This explains why no star-shapedness or connectedness assumption is needed.
Let , with density . The global inequality gives . Both and are bounded in absolute value by for some finite constant : the former has the same Gaussian-mixture representation as (4.7), with a finite set of means. Since has finite first moment, the KL identity is legitimate and gives
| (4.19) |
This proves (4.5) and . As , (4.10) supplies the reverse inequality, so .
Finally, has null expectation one and integrable logarithm under every Gaussian shift. By (4.3), its worst-case expected logarithm is at least . The nonempty set lies in , and (4.4) gives equality at every point of this support, so the worst-case value is exactly . Lemma 2.1 now yields , completing the proof of Theorem 4.1.
4.3 Quantitative approximation by finite priors
Approximating the optimal prior on a finite boundary grid perturbs both the mixture and its expected log likelihood under the original optimum. The first perturbation costs quadratic Gaussian relative entropy. The second is generally linear in the grid spacing, but vanishes when the grid stays inside the optimal support, where . This distinction gives the two rates below.
Corollary 4.4.
Let be a finite-valued Borel map satisfying , and put . Then
| (4.20) |
For every , there is a finite prior supported on such that
| (4.21) |
These priors can be chosen with as .
Proof.
Consider the joint distributions obtained by drawing and then drawing either or . Marginalising out cannot increase relative entropy. Equivalently, apply the log-sum inequality to this common mixing measure. The Gaussian KL formula therefore gives
| (4.22) |
Couple and , where and are independent. The Lipschitz bound for yields
Combining this with the KL identity
proves (4.20); nonnegativity follows from minimality of .
For the sharper bound, take a finite -net in the compact set , and let be a Borel nearest-net-point map on that support. Set . Support equalisation gives
The KL identity now gives equality in (4.21), and (4.22), applied on , supplies its upper bound. Finally, implies weak convergence to , by uniform continuity of continuous functions on the compact support. ∎
For numerical optimisation, Proposition 4.2 and Corollary 4.4 play different roles. The former certifies a candidate once its Gaussian integrals and the minimum of over the whole boundary have been bounded. The latter ensures finite-prior approximation, but its sharper rate uses the unknown optimal support. Neither the size of a sufficient grid nor finite support of is determined by these results. Beyond Gaussian shifts, the proof suggests looking for families that preserve raywise convexity of expected log-mixture likelihood ratios: the first-exit lemma needs only that property, whereas the finite example shows that observation-space convexity by itself is insufficient.
Appendix A An explicit finite exterior mixture
The limit in (3.20) proves the counterexample without a numerical computation. This appendix gives a fixed finite mixture with the same separation, using only analytic bounds, and records its exact masses. The constants and normals are those of Section 3.
Lemma A.1.
All eight component means of lie in , and
| (A.1) |
Proof.
For an axis normal the exposure gap is , and for a diagonal normal it is . Since , the smallest gap is
The last inequality follows by squaring , using . Divide every unnormalised tilted mass by the mass at its target vertex . The seven other nonzero base masses have ratio one, and the origin has ratio . Hence
| (A.2) |
Since all alphabet points have norm at most one,
To bound , note that . Thus
| (A.3) |
Consequently every component mean has norm at least , proving exterior feasibility.
Convexity of total variation and (A.2) give . Put . For each nonzero point ,
Jensen’s inequality and give , so . In particular, . The inequality , which follows by applying pointwise, now yields
| (A.4) |
In the last line, and . ∎
For an exact expression for , put . Symmetry leaves two component normalisers:
| (A.5) | ||||
| (A.6) |
The origin mass , the mass at each axis point, and the mass at each diagonal point are
| (A.7) | ||||
| (A.8) | ||||
| (A.9) |
Thus
| (A.10) |
For , the norms of the means of an axis-normal and a diagonal-normal component, respectively, are
| (A.11) | ||||
| (A.12) |
The mass formulas average the eight components in (3.19). The mean formulas sum for the representative normals and .
References
- [1] L. D. Brown. Fundamentals of Statistical Exponential Families with Applications in Statistical Decision Theory. Institute of Mathematical Statistics Lecture Notes–Monograph Series, vol. 9, 1986. https://doi.org/10.1214/lnms/1215466757.
- [2] L. D. Brown, I. M. Johnstone, and K. B. MacGibbon. Variation diminishing transformations: A direct approach to total positivity and its statistical applications. Journal of the American Statistical Association, 76(376):824–832, 1981. https://doi.org/10.1080/01621459.1981.10477730.
- [3] I. Csiszár. -divergence geometry of probability distributions and minimization problems. The Annals of Probability, 3(1):146–158, 1975. https://doi.org/10.1214/aop/1176996454.
- [4] I. Csiszár and F. Matúš. Information projections revisited. IEEE Transactions on Information Theory, 49(6):1474–1490, 2003. https://doi.org/10.1109/TIT.2003.810633.
- [5] I. Csiszár and F. Matúš. Closures of exponential families. The Annals of Probability, 33(2):582–600, 2005. https://doi.org/10.1214/009117904000000766.
- [6] P. Grünwald, R. de Heide, and W. Koolen. Safe testing. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86(5):1091–1128, 2024. https://doi.org/10.1093/jrsssb/qkae011.
- [7] P. Grünwald, Y. Hao, and A. Balsubramani. Growth-optimal e-variables and an extension to the multivariate Csiszár–Sanov–Chernoff theorem. arXiv:2412.17554v2, 2024. https://arxiv.org/abs/2412.17554.
- [8] P. Grünwald, A. Ramdas, R. Wang, and J. Ziegel, organisers. Game-theoretic statistical inference: Optional sampling, universal inference, and multiple testing based on e-values. Oberwolfach Reports, 21(2):1339–1386, Report No. 24/2024, 2024. https://doi.org/10.4171/OWR/2024/24.
- [9] A. Ram, M. Larsson, J. Ruf, and A. Ramdas. Strong duality for the GROW criterion. arXiv:2606.24768v2, 2026. https://arxiv.org/abs/2606.24768.
- [10] A. Rinaldo, S. E. Fienberg, and Y. Zhou. On the geometry of discrete exponential families with application to exponential random graph models. Electronic Journal of Statistics, 3:446–484, 2009. https://doi.org/10.1214/08-EJS350.
- [11] F. Topsøe. Information-theoretical optimization techniques. Kybernetika, 15(1):8–27, 1979. https://www.kybernetika.cz/content/1979/1/8/paper.pdf.