Uniform-in-Time Particle Convergence
for Projected Gaussian SVGD
Abstract
Liu, Ghosal, Balasubramanian, and Pillai [12] conjectured uniform-in-time mean-field convergence for projected Gaussian Stein variational gradient descent (SVGD) beyond Gaussian targets. We give a quantitative affirmative answer for the continuous-time, exact-Gaussian-moment system with kernel , assuming , , and . For Gaussian i.i.d. initialization and every , the empirical particle law and Gaussian mean-field law satisfy . The asymptotic orders of are , , and in dimensions one, two, and at least three, respectively; each is sharp in particle number. Thus the time supremum can be placed inside the expectation without losing the Gaussian empirical-measure rate, even at the minimal full-rank sample size. We also obtain high-probability all-time bounds and a last-iterate guarantee to the reverse-KL optimal Gaussian at the same statistical scale. The key deterministic inequality separates initial moment mismatch from the Wasserstein discrepancy of the standardized cloud, using stability in square-root covariance coordinates and exact affine transport. At fixed , we characterize the discrete limiting cloud: its moments are optimal, but its preserved standardized shape leaves a strictly positive full-law error.
Note. This paper was generated entirely by AI, including MiMo, using an automated research pipeline developed by Chenghua Liu and Hanyu Li.
1 Introduction and prior conjecture
This section states the prior conjecture, explains the obstacle to uniform full-law control, and summarizes our resolution. Detailed definitions and results follow in Section 2.
Stein variational gradient descent (SVGD) transports particles by a kernelized deterministic velocity field [11, 10]. Finite-horizon propagation-of-chaos constants can grow with the horizon, and time-averaged bounds do not control a physical-time (fixed-time or last-iterate) law; uniform-in-time full-law control is therefore delicate. Our result concerns a continuous-time Gaussian-score projection with exact Gaussian expectations, not the ordinary unprojected SVGD particle system.
In the arXiv version of Liu et al. [12], Theorem 3.7 gives a Gaussian-target, pointwise-in-time expectation bound with a constant uniform in time under Gaussian i.i.d. initialization. For general targets, their Theorem 4.3 derives the exact-moment affine representation and is followed by a qualitative conjecture of uniform-in-time mean-field convergence. We formulate a stronger quantitative version, making the sampling model explicit.
Open problem 1.1 (Quantitative exact-moment conjecture).
For a possibly non-Gaussian target satisfying and , initialize the projected particle system with i.i.d. samples from a full-rank Gaussian law . At fixed dimension, can one place the time supremum inside the expectation, retain the sharp Gaussian empirical-measure rate, and obtain a physical-time last-iterate bound to the reverse-KL optimal Gaussian?
Exact moment closure alone does not answer this question. The empirical law retains a standardized shape that affine dynamics transports rather than Gaussianizes, while the moment vector field becomes sensitive near the positive-definite boundary. We resolve both issues under global strong convexity and a bounded continuous Hessian, without assuming a third derivative or a Hessian-Lipschitz modulus:
- 1.
a square-root-coordinate energy yields semiglobal all-time stability on every energy sublevel;
- 2.
a deterministic oracle separates moment mismatch from a time-independent standardized-shape discrepancy;
- 3.
Gaussian matching and log-Wishart estimates give the sharp expected -rate, a high-probability upper bound, and the endpoint ; and
- 4.
the exact finite-particle limit retains a positive shape floor.
A scalar counterexample delineates the scope of the first item: the semiglobal Euclidean Lipschitz estimate used here cannot be replaced by one initialization-uniform constant over the full positive-definite cone.
2 Model and main results
This section defines the projected moment and particle dynamics, then states the deterministic oracle, its probabilistic consequences, and the exact finite-particle limit. We close with the assumptions that delimit their scope.
Let be fixed and let
where by strong convexity and, for constants ,
| (2.1) |
We use for the quadratic Wasserstein distance on and for pushforward by . Vector norms are Euclidean; matrix norms carry subscripts or . For , the product norm is . For every , denotes its principal symmetric square root, and
We also write , set and , and define . Unless indexed otherwise, may change from line to line and depend only on the fixed problem data in the relevant statement, never on , or a tail parameter.
For , define the exact Gaussian moments
| (2.2) |
The Gaussian mean-field moment flow is
| (2.3a) | ||||
| (2.3b) | ||||
where and . We write for the solution starting from fixed and .
Given initial points , set
| (2.4) |
With
the projected particle system is
| (2.5) |
This is the Gaussian-score projection associated with ; it is not the ordinary, unprojected SVGD particle system. Define
For Theorems 2.2 and 2.4, the initial points are drawn independently from . Under this Gaussian initialization, all particle paths and standardized quantities are assigned an arbitrary fixed value on the null event . On its complement the paths and Gaussian moments are continuous in , hence so is . Consequently
| (2.6) |
so every random supremum below is measurable.
For , use the constant-free energy representative
| (2.7) |
Then , where . Let be the unique minimizer, put and , and define
| (2.8) |
Section 4 proves that these sublevels are compactly contained in the positive covariance domain.
We first state the deterministic result from which the probabilistic all-time estimates are lifted. Given a deterministic cloud with empirical mean and covariance , choose any whitening matrix satisfying and set
| (2.9) |
The value of does not depend on the whitening: two whitening matrices differ by a left orthogonal factor, and is rotation invariant.
Theorem 2.1 (Deterministic localized oracle).
Let , let be the solution of (2.3), set , and let be the projected particle solution from an arbitrary deterministic cloud with . Put . For every there are finite constants and , depending only on , such that, whenever ,
| (2.10) |
Both the mean-field and particle solutions are global under these conditions. No independence or distributional assumption on the initial cloud is required.
In particular, any sequence of full-rank deterministic or dependent clouds lying in a common sublevel and satisfying converges uniformly in physical time.
For an unambiguous statement at every admissible , let
| (2.11) |
Thus, asymptotically, the first two rates are and , respectively. The regularization inside the logarithms only makes these rates finite for every admissible sample size.
We now specialize the oracle to Gaussian sampling, first in expectation.
Theorem 2.2 (Uniform-in-time particle convergence).
Assume (2.1), fix , , and , let , and draw the initial particles i.i.d. from . Then almost surely, the systems (2.3) and (2.5) have global solutions, and there is a finite constant , depending only on , such that
| (2.12) |
In particular, the same bound holds with the supremum outside the expectation. The constant is independent of and .
The next proposition shows that the expected rate cannot be improved in its -dependence, even before the dynamics starts.
Proposition 2.3 (Sharpness in particle number).
Proof.
At time zero, is the empirical measure of i.i.d. samples from . Applying the invertible affine map to the standard Gaussian empirical matching problem, with and , gives
The expectation on the right is bounded below by : in dimension one this is the two-sided Gaussian order-statistic estimate of Bobkov and Ledoux [5, Corollary 6.14] for ; in dimension two it follows from the nonstandard two-sample Gaussian matching lower bound of Talagrand [16, main result and display (1)] (if is an independent copy of , then ); and for it is the Gaussian matching lower bound of Ledoux and Zhu [9, Theorem 1.1]. Since time zero is included in the supremum, (2.13) follows. Reducing the constant handles the finitely many nonasymptotic sample sizes. ∎
The same decomposition also yields exponential concentration around the dimension-dependent statistical scale.
Theorem 2.4 (High-probability all-time convergence).
Under the assumptions of Theorem 2.2, there are constants , depending only on the fixed problem data, such that for every and , there is an event with on which
| (2.14) |
On , simultaneously for every ,
| (2.15) |
Equivalently, for , the statistical term is with probability at least , after changing the fixed constants.
Combining all-time particle control with convergence of the mean-field flow gives a last-iterate guarantee rather than a time-averaged one.
Corollary 2.5 (Physical-time last iterate).
Under the assumptions of Theorem 2.2, there are constants , depending only on the fixed problem data, such that for all ,
| (2.16) |
Consequently, once , the error is at the statistical scale .
The deterministic dynamics has an exact limit at every fixed particle number; the residual discrepancy is precisely an affine shape error.
Proposition 2.6 (Exact finite-particle limit and shape floor).
Let a deterministic full-rank cloud and mean-field initialization satisfy for some . There exist , , and an invertible matrix with such that, for
| (2.17) |
one has
| (2.18) | ||||
| (2.19) |
Moreover,
| (2.20) |
The constants and depend only on ; the matrix may depend on the initial cloud and the whitening coordinates, while the law does not. Thus the moments converge exactly to , while every full-rank finite cloud retains a strictly positive full-law error.
This shape floor persists for a fixed non-Gaussian sampling law, even as the number of particles grows.
Corollary 2.7 (Non-Gaussian initialization is not Gaussianized).
On one probability space, let be an infinite i.i.d. sequence from a non-Gaussian law having a density, positive-definite covariance , and finite second moment. For each , initialize the particle system with and define by Proposition 2.6 on its almost-sure full-rank event. Write
Put . Then, almost surely,
| (2.21) |
In particular, the limiting error is bounded away from zero; if , it converges to . Thus Gaussianity is a structural consistency condition for a fixed i.i.d. initialization law in Theorems 2.2 and 2.4, not merely a convenient source of concentration. This does not preclude consistent deterministic or -dependent designs covered by the oracle.
Remark 2.8 (Scope).
The probabilistic theorems concern continuous time, exact evaluation of (2.2), fixed dimension, Gaussian i.i.d. initialization, and the projected system (2.5). The deterministic oracle does allow arbitrary full-rank initialization. These are not dimension-free results or theorems for ordinary SVGD, Euler discretization, or Monte Carlo estimators of and . Since every centered -point covariance has rank at most , the restriction is necessary for this unregularized positive-definite dynamics.
3 Proof overview
This section gives the mechanism behind the results and records where each technical step enters. Complete arguments, including the endpoint and rare event calculations, are developed in Sections 4–7.
In square-root coordinates , the energy in (2.7) has second variation
| (3.1) |
which is at least . Its sublevels are compactly contained in . The moment flow can be written as with a positive-definite mobility ; on each sublevel, . This yields exponential convergence to . For two trajectories, a Gronwall bound controls the common compact sublevel until both enter a fixed neighborhood of . There the linearization is dissipative in the -norm, and patching the two regimes proves the all-time stability estimate. Section 4 and Section 5 give the details.
The particle equation is affine. Its empirical moments solve the same closed ODE, and a fundamental matrix gives
| (3.2) |
After whitening the initial cloud, both the empirical law and its moment- matched Gaussian are pushforwards by the same affine map. Consequently,
| (3.3) |
Combining this with semiglobal moment stability proves the deterministic oracle. Exponential moment convergence also makes the affine coefficient integrable, producing the limit and the two-sided shape floor in Proposition 2.6. These arguments are in Section 6.
For Gaussian initialization, concentration of the sample mean and covariance places the empirical initial state in a deterministic common sublevel, while Gaussian empirical matching controls . This proves the high-probability theorem directly from the oracle. For the expectation, the rare nearly singular event cannot be discarded; energy decrease reduces its contribution to a log-determinant moment. Bartlett’s decomposition gives a sum of variables whose second moments remain bounded even when . Cauchy–Schwarz then makes the rare-event contribution exponentially small. Section 7 contains the full localization and matching argument.
4 Reverse-KL geometry in square-root coordinates
This section identifies the gradient-flow structure behind the moment ODE. We first compute its Stein mobility and then use square-root coordinates to obtain strong convexity, compact sublevels, and quantitative convergence.
All matrix tangent and cotangent variables below lie in , equipped with the Frobenius inner product. The product inner product is
For , the exact reverse-KL energy is
| (4.1) |
Lemma 4.1 (Energy gradient and Stein mobility).
Proof.
For and all sufficiently small , Gaussian differentiation (Price’s identity) gives
The growth needed to differentiate under the integral follows from (2.1); differentiation of the log determinant then proves (4.2).
More generally, direct expansion gives the bilinear identity
| (4.6) |
The right-hand side is unchanged when the primed and unprimed variables are interchanged, so is self-adjoint. Taking the two arguments equal gives
which is (4.4); it vanishes only when and because . Substituting and in yields (2.3a) and (2.3b). ∎
Recall that with and that is defined by (2.7). Thus , so the two energies have identical derivatives in these coordinates.
The next lemma supplies the strong convexity and coercivity used throughout the stability analysis.
Lemma 4.2 (Regularity, strong convexity, and coercivity).
The function is on and, for ,
| (4.7) |
Moreover, has compact sublevel sets in and hence a unique minimizer .
Proof.
The Hessian bound implies and . The first two derivatives of are therefore dominated locally in by integrable polynomials in . In particular,
On a compact parameter neighborhood the integrands in the first and second variations are bounded respectively by and . Since is bounded and continuous, dominated convergence gives and the displayed second-variation formula. The first term in (4.7) is at least
and the log-determinant term is nonnegative because .
Let be the unique minimizer of and set . Strong convexity gives, with the eigenvalues of ,
| (4.8) |
The right-hand side diverges if , if an eigenvalue of tends to zero, or if . Thus every sublevel is compactly contained in . Existence follows by compactness and uniqueness by (4.7). Under the bijection , this minimizer corresponds to the unique reverse-KL optimal Gaussian , with . ∎
Let . Its differential is
| (4.9) |
The Sylvester map is invertible on when . The chain rule and Lemma 4.1 therefore transform the moment flow into
| (4.10) |
Here is smooth, symmetric, and positive definite. Lemma 4.2 gives , and therefore . This is the only local ODE regularity used below.
Proposition 4.3 (Sublevel ellipticity and convergence).
For , recall
There are such that on ,
| (4.11) |
For , one may take
| (4.12) |
If , every solution of (4.10) starting in is global, remains in , and satisfies
| (4.13) | ||||
| (4.14) |
For , the only trajectory is , which is global and stationary.
Proof.
For , uniqueness gives , so all claims follow with the stationary trajectory . Assume henceforth that . The bounds (4.11) follow from Lemma 4.2. To prove (4.12), put and . Then
whereas
Thus
Since and on , the congruence gives (4.12): indeed, .
Strong convexity and the fact that the segment from to the interior minimizer remains in imply the Polyak–Lojasiewicz bound
| (4.15) |
Along the flow,
This proves invariance and (4.13). Strong convexity also gives , proving (4.14). Finally, a local solution trapped in the compact set cannot explode or reach the positive-definite boundary, so the standard continuation theorem gives global existence. ∎
5 Semiglobal all-time stability
The next lemma is the deterministic core of the argument. Its constant is uniform in time but is allowed to depend on the initial energy sublevel.
Lemma 5.1 (Semiglobal uniform stability).
For every there is such that any two solutions of (4.10) with satisfy
| (5.1) |
If , there also exist , depending only on , such that
Proof.
The argument has two regimes. Sublevel invariance first gives a common compact set and hence an early-time Gronwall bound; exponential energy decay then brings both trajectories into a neighborhood of where a fixed -norm is contractive. Patching the estimates yields the all-time constant and eventual exponential decay.
The case is immediate. Suppose and set
Since ,
In the fixed norm , write . Then
| (5.2) |
Continuity of gives such that the convex ellipsoid
satisfies, for all and all ,
| (5.3) |
Indeed, continuity is used here to control the symmetric part of :
For , the segment formula and give
Consequently, whenever the segment lies in ,
| (5.4) |
Since every concentric ellipsoid is convex, the usual first-exit argument applied to (5.4) shows that each closed concentric ellipsoid contained in is forward invariant.
If and lie in the same such ellipsoid, then the segment joining them also lies there and
Thus
| (5.5) |
By (4.14), all trajectories starting in enter the ellipsoid of radius no later than
| (5.6) |
Before , all trajectories remain in the convex compact set
Let , where the operator norm is induced by the Euclidean–Frobenius product norm. The segment between the two trajectories remains in , so Gronwall’s inequality gives
| (5.7) |
After , both trajectories remain in the half-radius ellipsoid. Combining (5.5), norm equivalence, and (5.7) gives
Together with (5.7), this proves (5.1) and the claimed eventual exponential decay. In particular, the construction gives the explicit admissible choice
| (5.8) |
∎
Remark 5.2 (Dependence of the stability constant).
Finiteness of requires no global bound on : continuity of on the relevant compact parameter sets suffices. A closed-form upper bound for , however, requires a quantitative modulus of continuity for on those sets. A global bound on is one sufficient, but not necessary, condition for such a modulus.
The synchronous coupling and gives the useful consequence
| (5.9) |
6 Exact affine decomposition and shape error
We next derive the moment closure and affine representation directly from the particle system. The resulting affine lift separates the evolving moment error from a time-independent discrepancy of the standardized shape.
Lemma 6.1 (Exact closure and affine transport).
If , the particle system has a unique global solution. Its empirical moments solve (2.3) from and remain positive definite. If solves
| (6.1) |
then
| (6.2) | ||||
| (6.3) |
Proof.
We first construct the particle solution, thereby avoiding any continuation argument that presupposes positive definiteness of its empirical covariance. Let
be the global solution of (4.10) from , whose existence follows from Proposition 4.3. Along this trajectory define
and let be the global solution of , . Now set
| (6.4) |
Its empirical mean is . Its empirical covariance
solves
The covariance equation in (2.3) can equivalently be written
Uniqueness for this linear matrix equation therefore gives . Likewise, the mean equation is . Differentiating (6.4) now verifies
so the constructed paths solve (2.5) with their own empirical moments.
For completeness, any local particle solution with positive-definite empirical covariance satisfies, by averaging and centering,
These are exactly (2.3); uniqueness of the -flow forces its moments to equal . Its centered particles then solve the same linear equation with the same initial data, proving uniqueness. Finally,
for every finite , so . Since both the global -flow and its linear lift exist on every finite interval, this proves all claims, with . ∎
Henceforth write
| (6.5) |
for this empirical-moment trajectory.
Write the initial particles as
| (6.6) |
and define
| (6.7) |
For , almost surely. Let
| (6.8) |
This is precisely (2.9) with the whitening .
The following estimate relates the standardized shape to an ordinary Gaussian empirical measure without using inverse sample-covariance moments.
Lemma 6.2 (Standardized shape error).
Proof.
Since , the cross term vanishes in
Using the definition of gives
The scalar inequality for proves (6.9).
Let . Coupling by index and using the squared triangle inequality yields
Because is Wishart with degrees of freedom, the first expectation is explicit:
| (6.11) |
For the remaining empirical-matching term, the estimate is the Gaussian bound of Bobkov and Ledoux [5, Corollary 6.14]; the estimate is due to Ledoux [8, Theorem 1]; and for the claimed rate follows from Ledoux and Zhu [9, Theorem 1.1, with = p 2 ]:
The finitely many sample sizes excluded by an asymptotic formulation of these estimates are absorbed by enlarging . Since , this proves (6.10). ∎
To transfer the standardized discrepancy through the particle dynamics, we record a two-sided bound for every invertible affine map.
Lemma 6.3 (Two-sided affine lifting).
For any full-rank initial cloud, use the whitening and standardized law in (2.9). Then, for every ,
| (6.12) |
Proof.
Set . Then , and (6.2) gives
| (6.13) |
The same affine map pushes to . Pushing forward any coupling of and multiplies its quadratic cost by at most , proving the upper bound. Conversely, pull any coupling of the two image laws back through the invertible affine map and use . Taking the infimum over image couplings proves the lower bound because . ∎
Proof of Theorem 2.1.
Lemma 6.1 gives a unique global particle solution and identifies its moments with the -flow from . Since both initial moment states lie in , Proposition 4.3 and Lemma 5.1 give
| (6.14) |
where we used the synchronous Gaussian coupling (5.9). Invariance of gives . The upper bound in Lemma 6.3, followed by the squared triangle inequality, now proves (2.10). ∎
Proof of Proposition 2.6.
The proof first uses exponential moment convergence to show that the affine coefficient is integrable. Its fundamental matrix then converges to an invertible limit, which identifies both the limiting empirical law and its two-sided shape floor.
Write
The coordinate map has invertible differential. Since is an interior minimizer, ; hence
| (6.15) |
The map is Lipschitz on each . For , this follows by synchronously coupling the Gaussians and using the -Lipschitzness of . For , boundedness of , Pinsker’s inequality, and the Gaussian KL formula give, for ,
| (6.16) |
The last inequality holds because the Gaussian KL divergence is bounded by on the compact covariance box containing . Set . Proposition 4.3 and (6.15) therefore imply
| (6.17) |
The fundamental matrix from (6.1) consequently satisfies
Indeed, this follows by Gronwall for and for . Moreover,
exists, is invertible, and . Invertibility is explicit from
Set and . Because ,
so . Passing to the limit in gives . Since the standardized cloud has zero mean and identity covariance, coupling particles by index yields
which proves (2.18). Metric continuity and prove (2.19); applying the two-sided lifting argument to proves (2.20). Finally, is finitely atomic and is non-atomic, so their Wasserstein distance is strictly positive. Equivalently, is orthogonal: the dynamics optimizes the Gaussian moments but can only rotate, not Gaussianize, the standardized initial shape. If is another whitening, with orthogonal, then and . Hence the pushforward in (2.17), and therefore the limiting empirical law, is independent of the chosen whitening coordinates. ∎
Proof of Corollary 2.7.
Let and use the principal empirical whitening in (2.9). The empirical Wasserstein law of large numbers gives almost surely; the sample mean and covariance converge almost surely as well. Hence . Coupling corresponding sample points and then applying the fixed limiting affine map shows
On the same probability-one event, every cloud is full rank and its moment state has finite energy, hence belongs to some sublevel . The two-sided comparison below is independent of that cloud-dependent sublevel. Explicitly, the squared cost of replacing the empirical centering and whitening by their population values tends to zero by the sample second-moment law of large numbers, while the remaining term is at most . Taking the lower limit in the lower bound and the upper limit in the upper bound of (2.20) proves (2.21). Here because would imply that itself is Gaussian. ∎
7 Localization and proof of the main theorem
This section upgrades the deterministic oracle to the expected and high-probability theorems. A fixed well-conditioned event supplies a common energy sublevel, while a log-Wishart estimate controls its rare complement.
Fix and define the good event
| (7.1) |
Lemma 7.1 (Good-event control).
There are constants such that, for all ,
| (7.2) |
There is a deterministic , depending only on the fixed problem data, such that on both and belong to . Furthermore, there is , depending only on , such that
| (7.3) |
Proof.
Since , the Gaussian mean has an exponential tail at the fixed threshold . Furthermore,
Thus
For each fixed pair of coordinates , the centered variable has a subexponential norm bounded by a constant independent of . Scalar Bernstein’s inequality, a union bound over the entries, and fixed-dimensional norm equivalence bound the first probability by ; a Gaussian tail gives the same bound for the second. Together with the mean event in (7.1), and after enlarging for the finitely many small sample sizes, this proves (7.2).
On ,
and
Thus ranges over a fixed compact subset of ; continuity of places it, together with , in a deterministic sublevel .
Proof of Theorem 2.4.
We intersect concentration events for the empirical measure, mean, and covariance. These events place all empirical initial states in one deterministic sublevel, so the oracle applies with constants independent of and .
Let . Matching corresponding atoms shows that the map
is -Lipschitz on . Gaussian concentration and therefore give
| (7.5) |
Indeed, outside an event of probability , , and the squared triangle inequality gives (7.5).
Fixed-dimensional Gaussian concentration for and scalar Bernstein concentration applied to every entry of imply that, for , outside an event of probability ,
| (7.6) |
Here is reduced, if needed, so that the second assertion follows from the usual Bernstein scale ; indeed, throughout the stated range. On the intersection of the good events underlying (7.5) and (7.6), the square-root estimate (7.4) gives
| (7.7) |
The second bound follows from the residual identity and coupling the atoms with . Moreover, all on these events, together with the fixed , lie in one deterministic energy sublevel : the covariance eigenvalues stay in a fixed compact interval, and the sample means stay in a fixed ball because . Here depends only on the fixed problem data and the chosen , not on or . Denote this intersection by ; the preceding bounds give .
The deterministic oracle (2.10), the fact that , and a union bound prove (2.14). For the last iterate, Proposition 4.3 and the synchronous Gaussian coupling give the deterministic estimate
| (7.8) |
Combining this with (2.14) by the squared triangle inequality proves (2.15). Finally take and rename the fixed constants to obtain the confidence-level formulation. ∎
On the good event, Lemmas 5.1 and 7.1, together with (5.9), imply
| (7.9) |
Invariance of also gives a deterministic upper bound for on . Hence Lemmas 6.2 and 6.3 yield
| (7.10) |
The complement of cannot simply be discarded: there the sample covariance can be arbitrarily close to singular. The following lemma is the required weighted tail estimate.
Lemma 7.2 (Weighted control outside the good event).
There are constants , depending only on , such that
| (7.11) |
Proof.
We first dominate the entire path by the initial energy, reduce that energy to polynomial moments and a log determinant, and then use a uniform log-Wishart second moment with Cauchy–Schwarz on .
For a probability law , write . The independent coupling through the origin gives
| (7.12) |
For the particle cloud,
Strong convexity of and energy decrease along the random moment flow give
| (7.13) |
The upper Hessian bound on gives
| (7.14) |
Indeed, Taylor’s theorem and the upper Hessian bound imply , so
the entropy term is , whose positive contribution is exactly one half of . Define
| (7.15) |
All log-determinant quantities here and below are evaluated on the almost-sure full-rank event; on its null complement they inherit the fixed extension specified after (2.5).
We next prove
| (7.16) |
The polynomial terms have uniform moments of every fixed order: the sample mean is Gaussian with uniformly bounded covariance, while is bounded above by a fixed multiple of . For the remaining term, Bartlett’s decomposition [1, Lemma 7.2.1] gives
| (7.17) |
For in a neighborhood of zero,
Differentiating the cumulant generating function of at zero gives the digamma and trigamma identities
and hence the explicit formula
Here , and the second term is uniformly bounded by the standard digamma asymptotic , together with the finitely many bounded small values of . Hence
| (7.18) |
Now
For fixed and , the deterministic logarithms are uniformly bounded. Hence (7.18) and show that
| (7.19) |
Since , the asserted uniform second-moment bound (7.16) follows.
The algebraic endpoint is independent of the dynamics and follows from an affine-span characterization.
Proposition 7.3 (Exact full-rank endpoint).
For any points in , the centered empirical covariance has rank at most and is positive definite if and only if the points affinely span . Consequently is necessary for the unregularized positive-definite dynamics, and it is sufficient almost surely for i.i.d. initialization from any distribution having a density.
Proof.
If , then and . Thus , and rank is equivalent to affine spanning. Under a density, affine dependence of any fixed samples is a Lebesgue-null event. ∎
Remark 7.4 (Logarithmic control at the endpoint).
At , the final Bartlett factor in (7.17) has one degree of freedom. The covariance is positive definite almost surely, but , whereas . This is why logarithmic-energy localization reaches the exact algebraic endpoint.
Proof of Theorem 2.2.
On , the squared triangle inequality with the intermediate Gaussian , followed by (7.9) and (7.10), gives
Lemma 7.2 controls . Since for every fixed , the two event contributions prove (2.12). Almost-sure positive definiteness at time zero follows from Proposition 7.3; global existence follows from Lemma 6.1 and Proposition 4.3. ∎
8 Why a global Euclidean stability constant is impossible
The stability constant in Lemma 5.1 necessarily depends on an energy sublevel. The next proposition rules out a uniform replacement.
Proposition 8.1 (No global Euclidean Lipschitz constant for the scalar flow).
In dimension , there is a smooth, strongly convex potential with globally bounded Hessian for which no finite can satisfy
where is the standard-deviation component of the moment flow started from .
Proof.
Fix and , and take
| (8.1) |
Then . The subspace is invariant. Writing on this subspace gives
| (8.2) |
In fact,
because . Thus is strictly increasing from zero to infinity and has a unique crossing of level one. In particular, on . Fix . For , let be the time at which first reaches , and then freeze time at . This time is finite because is continuous and strictly positive on . The variational identity for a scalar autonomous flow is
Moreover, as . Therefore, with fixed while differentiating,
| (8.3) |
Any common Lipschitz constant for all fixed-time maps would bound the left-hand side of (8.3), a contradiction. If the full moment flow had one initialization-uniform Euclidean Lipschitz constant, its restriction to the invariant subspace would have the same property, so the scalar contradiction also rules out that claim. Notice that we have not differentiated the composite , whose derivative is zero. ∎
This obstruction does not contradict Lemma 5.1: a fixed energy sublevel stays a positive distance away from . It explains why near-singular Wishart samples must be controlled probabilistically through energy rather than by a global derivative bound for the flow.
9 An explicit nonquadratic example
This section verifies the assumptions and identifies the variational optimum for a concrete nonquadratic target. It illustrates that the theorem is not confined to Gaussian potentials.
Example 9.1 (Cosine perturbation of a Gaussian target).
Let
| (9.1) |
Then , so Theorem 2.2 applies. If denotes the variance of , then
| (9.2) |
The moment flow is
| (9.3) |
Up to an additive constant, its reverse-KL energy is
| (9.4) |
and
| (9.5) |
The mobility determinant is . With ,
| (9.6) |
The upper-left entry is at least , and its Schur complement is bounded below as follows:
| (9.7) | ||||
Indeed, and . Hence is strictly convex and coercive, since the perturbation in (9.4) is bounded while diverges at every boundary of . Furthermore,
Let . Then
Thus there is a unique with . The point is stationary and, by strict convexity, is the unique minimizer; equivalently,
| (9.8) |
This gives a fully worked, genuinely nonquadratic target with an explicit scalar characterization of the optimum.
10 Related work and positioning
This section distinguishes the projected exact-moment problem from ordinary SVGD and from time-averaged propagation of chaos. It also identifies the empirical-matching results used only as statistical inputs.
SVGD was introduced by Liu and Wang [11], and its gradient-flow and mean-field viewpoints were developed by Liu [10], Lu et al. [13], and Duncan et al. [6]. Nonasymptotic and finite-particle analyses include Korba et al. [7] and Shi and Mackey [14]. These works concern ordinary SVGD and do not themselves give uniform physical-time stability for the projected system (2.5).
The closest starting point is Liu et al. [12], whose exact-moment closure and affine representation lead to the prior conjecture discussed above. Banerjee et al. [3] prove finite-particle KSD and bounds for ordinary SVGD, including bilinear-plus-Matérn kernels, with long-time statements for time-averaged particle laws. The broad KSD and conclusions of Balasubramanian et al. [2], and their in-probability conclusion, concern time-averaged laws uniformly over deterministic averaging horizons, rather than physical-time last iterates; their finite-feature results have a different scope. Banerjee and Kim [4] obtain physical-time results for unprojected, Langevin-regularized SVGD with Brownian noise, whereas our system is deterministic and uses an exact Gaussian-score projection. Stein and Li [15] analyze generalized bilinear kernels for Gaussian targets. None of these results directly yields the fixed-, non-Gaussian, projected full-law theorem here.
11 Implications and limitations
The deterministic oracle applies to each fixed full-rank cloud, with constants determined by a sublevel containing both initial moment states. A sequence of deterministic, dependent, or variance-reduced designs inherits a uniform guarantee when it lies in one common sublevel and its moment and standardized-shape errors vanish. Gaussian independence is used only to obtain the sharp expected -rate and the stated high-probability upper bound.
At fixed , the affine dynamics converges to a discrete law with moments but preserves standardized shape up to rotation. Thus moment convergence must not be confused with Gaussianization. The nonquadratic target in Section 9 shows that the theory is not limited to Gaussian potentials.
The restrictions are substantive. Constants depend on and deteriorate near the positive-definite boundary. Section 8 proves, for a smooth scalar potential with bounded Hessian, that the Euclidean standard-deviation flow has no single initialization-uniform Lipschitz constant over . This does not exclude other metrics or weighted stability mechanisms. No dimension-free, Euler-discretization, or estimated-moment claim is made. Finally, is the reverse-KL Gaussian approximation, not the non-Gaussian target itself, so approximation bias is separate from particle error.
12 Conclusion
Square-root geometry and affine whitening turn the non-Gaussian projected Gaussian-SVGD problem into a stable moment flow plus a preserved shape term. This yields a deterministic all-time oracle, a sharp expected -rate, a high-probability upper bound at the corresponding scale, validity at the minimal count , and an exact description of the finite-particle limit. The result resolves the strengthened continuous exact-moment conjecture under global strong convexity and a bounded continuous Hessian, without a third-derivative or Hessian-Lipschitz assumption.
References
- [1] T. W. Anderson. An Introduction to Multivariate Statistical Analysis. Third edition, Wiley-Interscience, Hoboken, NJ, 2003.
- [2] K. Balasubramanian, S. Banerjee, and A. Korba. Uniform-in-time propagation-of-chaos for Stein variational gradient descent. arXiv:2607.00149, 2026.
- [3] S. Banerjee, K. Balasubramanian, and P. Ghosal. Improved finite-particle convergence rates for Stein variational gradient descent. In International Conference on Learning Representations, 2025.
- [4] S. Banerjee and D. Kim. Quantitative target convergence and uniform-in-time propagation of chaos for Langevin-regularized SVGD. arXiv:2608.28827, 2026.
- [5] S. Bobkov and M. Ledoux. One-dimensional empirical measures, order statistics, and Kantorovich transport distances. Memoirs of the American Mathematical Society, 261(1259), 2019. doi:10.1090/memo/1259.
- [6] A. Duncan, N. Nüsken, and L. Szpruch. On the geometry of Stein variational gradient descent. Journal of Machine Learning Research, 24(56):1–39, 2023.
- [7] A. Korba, A. Salim, M. Arbel, G. Luise, and A. Gretton. A non-asymptotic analysis for Stein variational gradient descent. In Advances in Neural Information Processing Systems, volume 33, pages 4672–4682, 2020.
- [8] M. Ledoux. On optimal matching of Gaussian samples. Journal of Mathematical Sciences, 238(4):495–522, 2019. doi:10.1007/s10958-019-04253-6.
- [9] M. Ledoux and J.-X. Zhu. On optimal matching of Gaussian samples III. Probability and Mathematical Statistics, 41(2):237–265, 2021. doi:10.37190/0208-4147.41.2.3.
- [10] Q. Liu. Stein variational gradient descent as gradient flow. In Advances in Neural Information Processing Systems, volume 30, pages 3115–3123, 2017.
- [11] Q. Liu and D. Wang. Stein variational gradient descent: A general purpose Bayesian inference algorithm. In Advances in Neural Information Processing Systems, volume 29, pages 2378–2386, 2016.
- [12] T. Liu, P. Ghosal, K. Balasubramanian, and N. S. Pillai. Towards understanding the dynamics of Gaussian–Stein variational gradient descent. In Advances in Neural Information Processing Systems, volume 36, pages 61234–61291, 2023. doi:10.52202/075280-2676; arXiv:2305.14076.
- [13] J. Lu, Y. Lu, and J. Nolen. Scaling limit of the Stein variational gradient descent: The mean field regime. SIAM Journal on Mathematical Analysis, 51(2):648–671, 2019. doi:10.1137/18M1187611.
- [14] J. Shi and L. W. Mackey. A finite-particle convergence rate for Stein variational gradient descent. In Advances in Neural Information Processing Systems, volume 36, pages 26831–26844, 2023. doi:10.52202/075280-1166.
- [15] V. Stein and W. Li. Towards understanding accelerated Stein variational gradient flow: Analysis of generalized bilinear kernels for Gaussian target distributions. arXiv:2509.04008, version 2, 2025.
- [16] M. Talagrand. Scaling and non-standard matching theorems. Comptes Rendus Mathématique, 356(6):692–695, 2018. doi:10.1016/j.crma.2018.04.018.