Rank-dependent lower bounds for quantum chi-squared tomography
Abstract
We prove copy lower bounds for estimating a rank-at-most- quantum state in dimension under two chi-squared losses. For Bures chi-squared divergence, error at most with success probability at least requires copies with collective measurements and with adaptive one-copy measurements. The bounds hold uniformly for , , and , where is universal, and match the upper bounds of Flammia and O’Donnell (Quantum, 2024) up to logarithmic factors. For right-inverse chi-squared divergence, the corresponding rates are and ; we give upper bounds matching exactly in the collective model and up to a logarithmic factor in the one-copy model at constant confidence. The lower bounds quantify how much output mass must be allocated outside an uncertain support. We bound the volume of support subspaces compatible with any successful estimate, then use a Sobolev inequality to relate posterior concentration on these sets to the Fisher information available from measurements. For pure states, the usual covariant measurement with a regularized output attains both collective rates, and we evaluate the exact minimax expected right-inverse chi-squared loss.
Note. This paper was generated entirely by AI, including MiMo, using an automated research pipeline developed by Chenghua Liu and Hanyu Li.
1 Introduction
Low rank reduces the number of parameters of a quantum state, but it need not reduce the cost of estimation under a loss sensitive to small output eigenvalues. Chi-squared tomography illustrates this distinction. An estimate supported on the wrong subspace can have infinite loss even when the estimated and true subspaces are close. A learner must therefore allocate output mass to directions that remain uncertain. We study the statistical cost of this allocation.
Flammia and O’Donnell [1] gave Bures chi-squared estimators using copies with collective measurements and with adaptive one-copy measurements, for rank-at-most- states in dimension and divergence error . They asked whether the collective rate admits a matching rank-dependent lower bound. We prove lower bounds matching both rates up to logarithmic factors. We also analyze right-inverse chi-squared loss, which agrees with Bures chi-squared loss on commuting states but has a different accuracy dependence: its collective copy complexity is , and its adaptive one-copy complexity is , at constant confidence and sufficiently small .
Our hard instances are states maximally mixed on an unknown subspace. For a fixed output estimate, small loss restricts both its tail mass and the orientation of the true support. We show that these restrictions force its set of successful support subspaces to have small volume. For the Bures loss, a convexity argument shows that a scalar tail spectrum maximizes the relevant volume; for the right-inverse loss, the constraint is an ellipsoid whose volume can be evaluated directly. A Sobolev inequality then bounds the probability mass that a posterior can place in a set of that volume. Combining this bound with the Fisher information of the experiment controls success probability without an unbiasedness assumption. On the same family, adaptive one-copy measurements have a Fisher-information trace budget smaller by a factor of order than collective measurements.
These bounds separate quantum chi-squared tomography from more familiar tomographic tasks. Haah, Harrow, Ji, Wu, and Yu [3] gave collective infidelity estimators with copy complexity for error . O’Donnell and Wright [8, 9] developed collective estimators based on Schur–Weyl sampling. Their spectral estimate satisfies for the true spectrum [9, Theorem 1.7]. This guarantee has the true spectrum in the denominator, whereas our loss has the estimate in the denominator. Moreover, estimating the spectrum does not resolve its eigenvectors. Our flat-spectrum instances isolate the latter uncertainty.
Projector states also underlie the trace-distance lower bounds of Scharnhorst, Spilecki, and Wright [11]. Their proof includes a reduction from trace-distance to Bures-distance projector tomography; Bures distance is distinct from the Bures chi-squared divergence studied here. Nayak and Zhou [7] recently proved that trace-norm tomography with adaptive measurements on at most fresh copies at a time has copy complexity at sufficiently small constant accuracy. Their lower bound uses the same flat-support family in local graph coordinates. In the single-copy case, their block-positivity calculation gives the same Fisher-information trace scale as Lemma 3.5; they also use a smoothly truncated Gaussian prior and an adaptive Fisher-information chain rule, but convert the resulting information bound to trace-norm risk via the van Trees inequality. Thus we do not regard support rotations, the single-copy Fisher budget, or adaptive Fisher accounting as specific to the present work. Related multiparameter Fisher-information tradeoffs appear in Gill and Massar [2] and Zhou and Chen [13].
The loss-specific ingredient here is instead the geometry of the set of supports on which a fixed output has small chi-squared loss, together with its conversion to posterior mass by a Sobolev inequality. Using only the trace-norm comparison at error gives the scales collectively and for one-copy measurements. The Bures lower bounds below are larger by , while the right-inverse bounds have a different accuracy exponent. Concurrent work of Keskin, Luo, Majid, and Radzihovsky [6] proves optimal bounded-joint-measurement lower bounds for full-rank trace-norm tomography using a metric Fano inequality and a log-Sobolev comparison between mutual and Fisher information. That result is methodologically adjacent but does not address the low-rank chi-squared losses considered here.
We complement the lower bounds with upper-bound reductions that track subnormalized blocks and all copies consumed in preparing them. The Bures construction follows the spectral-block approach of [1, Section 3.3]; the right-inverse construction uses isotropic regularization. The collective reduction uses the normalized Frobenius estimator of O’Donnell and Wright [8, Theorem 1.2]. For pure states, we analyze a regularized output of the known covariant measurement [4]. We also evaluate its exact minimax expected right-inverse loss, distinguishing the choice of measurement from the Bayes-optimal output rule [12].
1.1 Model, losses, and results
Let be the density matrices on , and let , where and are integers. We write , , and for the Frobenius, operator, and trace norms. Positive definiteness is denoted by ; denotes the positive-semidefinite order. An identity matrix carries a dimension subscript when needed.
For , set and . The Bures (or SLD) and right-inverse losses are
| (1) | ||||
| (2) |
For singular , evaluate the formulas on its support if , and set the loss to infinity otherwise. In an eigenbasis with ,
| (3) | ||||
| (4) |
Both reduce to on commuting states. These are two members of the broader family of quantum chi-squared divergences [10, 14]. A third choice is
with the same support convention. The scalar mean inequalities give
| (5) |
In particular, our Bures lower bounds also apply to ; the right-inverse rates concern specifically.
Definition 1.1 (Learning model).
A learner receives independent copies of an unknown and returns an arbitrary density matrix . Its output need not have rank at most . For a loss , the constant-confidence requirement is
| (6) |
Collective measurements permit any positive operator-valued measure (POVM) on . Adaptive one-copy measurements permit a POVM on each next copy chosen from previous classical outcomes and internal randomness, but no quantum memory coupling different copies. All consumed copies, including rejected projection outcomes, count toward . Measurements and estimators are measurable.
Write and for the minimum copy counts satisfying Equation 6 in the two models. The parameter bounds the divergence itself; the convention is obtained by substitution. The notation and hides universal constants, and and additionally suppress logarithmic factors.
Theorem 1.2 (Main lower bounds).
There are universal constants such that, for all integers , , and ,
| (7) | ||||||
| (8) |
For , the same bounds hold on the smaller family , where ranges over rank- orthogonal projectors.
For pure states, the Bures bound is already , rather than the scale suggested by a local parameter count. Pure states also suffice for the collective right-inverse bound . Allowing a low-rank input does not remove the need for a higher-rank output that assigns mass to uncertain directions.
Theorem 1.3 (Upper bounds).
For and , let . There are learners that succeed with probability at least uniformly on , with the following copy bounds:
| Bures, collective: | (9) | |||
| Bures, adaptive one-copy: | (10) | |||
| Right-inverse, collective: | (11) | |||
| Right-inverse, adaptive one-copy: | (12) |
For pure inputs, collective measurements achieve for and for at constant confidence, without logarithmic factors.
The Bures constructions recover the rates of [1]; we include their Frobenius-to-chi-squared reduction with explicit error and copy accounting. The pure-state construction also explains the different accuracy exponents: the Bures tail contribution is quadratic in the residual mass outside the estimated direction, whereas the right-inverse tail contribution is linear.
1.2 Comparison inequalities
We use the squared Bures distance
The amplitude representation gives the triangle inequality: align an amplitude of the intermediate state with one endpoint, and an amplitude of the other endpoint with that intermediate amplitude, then apply the Frobenius triangle inequality.
Lemma 1.4 (Two comparison inequalities).
For density matrices,
| (13) |
Moreover is the largest classical chi-squared divergence obtainable by measuring the pair of states with the same POVM.
Proof.
Assume first and put . For a POVM element , set and . Hilbert–Schmidt Cauchy–Schwarz yields
Summing gives classical chi-squared at most . Equality holds for the projective measurement in an eigenbasis of .
To compare with Bures distance, diagonalize the positive semidefinite matrix
It satisfies . In its eigenbasis, with eigenvalues , the two measured distributions obey and
Their squared Hellinger distance is therefore , and .
Finally, weighted Cauchy–Schwarz gives, for Hermitian ,
Choose to obtain the trace-norm bound. If is singular and contains the support of , restrict the argument to its support. Otherwise both quantum losses are infinite, and measuring the projection onto also gives infinite classical chi-squared divergence. ∎
The trace-norm inequality gives a first reduction from chi-squared to trace-distance tomography. To obtain the additional factor in the Bures bounds, and the additional in the right-inverse bounds, we must also use the restriction on the estimate’s tail spectrum. The next section measures its effect on the volume of a successful set.
2 Volumes of chi-squared loss balls
Fix integers , put , and write
The hard family consists of , with a rank- orthogonal projector. Let be the invariant probability measure on these projectors. Complex matrices have real dimension ; their Lebesgue measure uses the real and imaginary parts of their entries.
2.1 Grassmannian coordinates
Lemma 2.1 (Grassmann coordinates and a matrix ball).
Let
Then
| (14) |
In the graph chart
| (15) |
Haar measure has density
| (16) |
If is a Haar orthonormal -frame in and denotes its last rows, then is uniform on the matrix ball in Equation 14.
Proof.
One way to derive the chart density is to take a complex Gaussian matrix with square of size . Its column space is Haar and, almost surely, its graph coordinate is . Substituting has real Jacobian . Changing in the remaining Gaussian integral leaves the factor .
The normalizing integral is elementary. Integrate the rows one at a time, using the determinant lemma and
The resulting product is
For completeness, the change of variables maps onto the operator-norm unit ball and has Jacobian . To check this, let be the squared singular values of ; those of are . In complex rectangular singular-value coordinates the radial factor is . The substitution contributes, for each , the exponent . Angular coordinates are unchanged. Thus is uniform on the matrix ball under Equation 16. A Haar frame over the graph subspace is with a Haar ; right-unitary invariance shows its bottom block has the same law as .
Finally,
which proves the two volume bounds. ∎
We next fix an estimate and determine which support subspaces can have small loss relative to it. The first restriction is on the mass outside its leading eigenspace.
Fix a positive definite estimate and let be a projector onto its largest eigenvalues. In this basis write
Lemma 2.2 (Tail mass and subspace distance).
If and , then
| (17) |
The same holds under a guarantee.
Proof.
Let denote root fidelity. For a flat rank- state,
By Lemma 1.4, , so . The variational principle for the sum of the top eigenvalues gives and hence .
The eigenvalues of a compression are individually at most the corresponding top eigenvalues of . Therefore minimizes over all rank- projectors . The triangle inequality implies . If are the principal angles between the two subspaces, then
Use for the last assertion. ∎
2.2 Tail geometry and successful volumes
For and Hermitian , define
Lemma 2.3 (Uniform tail eigenvalues maximize a volume).
For fixed and , the Lebesgue volume of is at most its value at .
Proof.
The variational identity
| (18) |
shows joint convexity in . Also
is a positive map. Thus is increasing in the positive-semidefinite order when : if , then . It follows that
The constraint is the convex Schur-complement constraint . Consequently is jointly convex in .
Let . These are bounded convex bodies: positivity of as a quadratic form and the quartic homogeneity in imply coercivity, and a neighborhood of zero lies inside each body. Joint convexity gives . The Brunn–Minkowski inequality in real dimension therefore says that is concave in . Unitary conjugation of , accompanied by left multiplication of , preserves volume. Diagonalize and average its conjugates by all coordinate permutations; their average is . Concavity gives the claimed volume bound. ∎
The preceding symmetrization removes the dependence on individual tail eigenvalues. This gives a uniform success-volume bound for every output estimate, which is the geometric input to the statistical argument.
Theorem 2.4 (Haar volumes of loss balls).
For and any density matrix ,
| (19) | ||||
| (20) |
Proof.
If is singular, finite loss requires to lie in one fixed proper subspace. This is a Haar-null set. We may assume and work in the basis above. Take a Haar frame for . By Lemma 2.1, is uniform on the operator-norm unit ball, and the true bottom block is .
For the Bures loss, the contribution of this bottom block is
All entrywise contributions outside it are nonnegative. By Lemma 2.2, a successful therefore satisfies
| (21) |
Discarding the restriction only enlarges the numerator volume. By Lemma 2.3, it is enough to consider . In that case
Cauchy–Schwarz on the at most nonzero eigenvalues of gives . Thus Equation 21 implies
The Frobenius ball of squared radius has volume . Divide by and use to obtain .
For the right-inverse loss, the projector identity gives
Matrix Cauchy–Schwarz and the principal angles give
Thus a successful satisfies . The volume of this ellipsoid is
where . Dividing by proves Equation 20. ∎
Corollary 2.5 (Euclidean volumes in a fixed local chart).
Let and restrict in Equation 15 to . For every fixed estimate , let and be the Lebesgue volumes of the two successful parameter sets in this chart. Then
| (22) |
Proof.
On this chart, Equation 16 is at least . Therefore a local Lebesgue volume is at most times the corresponding Haar probability. Take th roots, use , and apply Equations 14 and 2.4. ∎
3 Posterior concentration and the measurement lower bounds
Each possible output succeeds on a set whose volume was bounded in Corollary 2.5. A successful experiment must therefore produce posteriors concentrated on small sets. We first quantify the information needed for that concentration, then compare it with the information a measurement can extract from the flat-state family.
3.1 Posterior mass and Fisher information
For a density on , write
Weak derivatives suffice.
Lemma 3.1 (Fisher information and volume).
Let , let vanish outside a compact set, and let have Lebesgue volume . Then
| (23) |
An absolute constant in place of would give the same conclusions below.
Proof.
The Sobolev inequality, obtained from the Euclidean isoperimetric inequality by coarea (see, e.g., [5]), is
where is the volume of the unit ball in . Apply it to and use Cauchy–Schwarz to obtain
For the last inequality, the cube of side lies in the unit ball, so , and . Hölder’s inequality with gives
which proves the claim. Approximation extends the argument to . ∎
Let be a prior with compactly supported square root . Write the likelihood of a classical outcome as relative to a parameter-independent measure , and put . We use the following regularity conditions: is locally Lipschitz in for -almost every , the local squared gradient energies are integrable in , and differentiation can be passed under . We verify these conditions for our quantum experiments below. Define
This is the usual classical Fisher information matrix where the likelihood is positive, interpreted through weak derivatives at zeros.
Lemma 3.2 (Posterior information identity).
Under the preceding conditions, assume . Then
| (24) |
In particular, almost every posterior has a compactly supported square root whenever the right-hand side is finite.
Proof.
Let be the marginal density. For , the posterior square root is . Expanding its gradient and averaging gives
The cross term vanishes because . Its absolute integrability follows by Cauchy–Schwarz from the two energies on the right. Multiplication by the locally Lipschitz function preserves the property on the compact support of . Outcomes with make no contribution. ∎
Corollary 3.3 (Bayesian success bound).
Under the preceding conditions, let . Suppose that every estimate’s successful parameter set within the prior support has Lebesgue volume at most . Then
| (25) |
Proof.
Apply Lemma 3.1 to each posterior and the success set specified by its output, then average and use Equation 24. Include any internal randomization in . ∎
For the Grassmannian family, the prior must stay inside the chart where the volume comparison is uniform. A truncated Gaussian achieves this with Fisher information of order .
Lemma 3.4 (Local prior).
On , , there is a prior supported on such that
| (26) |
Its square root is compactly supported and belongs to .
Proof.
We give a construction to keep the dimension dependence explicit. Let and let the real coordinates of be independent centered Gaussians of variance , with density . Let be for , linear from to on , and zero for . Set
Using -nets of the complex unit spheres of sizes at most and , respectively,
Here is at most twice the largest over the two nets, and each such complex Gaussian has the usual radial Gaussian tail. Consequently .
The operator norm is -Lipschitz in Frobenius norm, so the squared norm of the weak gradient of the cutoff is at most . Differentiating and using gives
For example, suffices. The cutoff vanishes at the boundary, so extension by zero has the stated property. ∎
3.2 Information budgets of quantum measurements
For the graph family Equation 15, use also the orthonormal complementary frame
A coordinate variation induces
| (27) |
The map is a contraction for the real Frobenius norm.
Lemma 3.5 (Collective and adaptive information budgets).
For the real coordinates of :
| (28) | ||||||
| (29) |
Proof.
We write sums over POVM outcomes; for a continuous POVM these are integrals against a dominating scalar measure. At , a tangent has symmetric logarithmic derivative (SLD) and quantum Fisher information
The same weighted Cauchy–Schwarz argument as in Lemma 1.4 bounds any measured Fisher information in this direction by that quantity. On copies the SLD is the sum of its single-copy versions. The cross terms vanish because , so the bound multiplies by . Summing over an orthonormal real coordinate basis and using the contraction in Equation 27 proves Equation 28.
For a single-copy POVM element, use its blocks in the basis:
In the aligned real tangent coordinates , the sum of squares of the probability derivatives is . Positivity gives and thus . Consequently the Fisher trace in these coordinates is at most
Terms with have and contribute zero. Pullback by the contractive map in Equation 27 cannot increase the Fisher trace.
For adaptation, condition on the previous outcomes. Each conditional likelihood has the same bound. Its score has conditional mean zero, so scores at distinct times are orthogonal in expectation. Their Fisher traces add to at most . ∎
We now justify the likelihood regularity used in Lemma 3.2, including continuous outcomes and zero probabilities. The full transcript of either measurement model, including internal randomness, is the outcome of a fixed POVM on copies. In finite dimension it has an operator-valued density relative to the finite scalar measure , with and almost everywhere. Hence
Here is smooth in . The norm of a smooth matrix-valued function is locally Lipschitz. On every compact chart, its Lipschitz constants are bounded uniformly in , since . The same bounds and finiteness of justify integration of weak derivatives and differentiation of . A weak gradient vanishes almost everywhere on the zero set of a locally Lipschitz function, so the zero-likelihood terms agree with the convention in the budget proof. Thus the information budgets bound exactly the square-root energies used in Equation 24. No positivity assumption on the likelihoods is needed.
3.3 Proof of the rank-dependent lower bounds
Assume first , so . Apply Corollary 3.3 with the prior in Lemma 3.4, the volumes in Corollary 2.5, and the budgets in Lemma 3.5. For any learner on the flat rank- family, its prior-average success probability is at most the corresponding expression below:
| (30) | ||||
| (31) | ||||
| (32) | ||||
| (33) |
The constants absorb only numerical quantities. Since , a small universal choice of makes the term independent of at most . A uniform success guarantee of therefore forces
| (34) |
The omitted case is ; the direct pure-state argument in Section 4 proves the same bounds without a Sobolev inequality in dimension two.
To prove Theorem 1.2 for , choose and use flat rank- inputs. Then is comparable to . If , is also comparable to ; otherwise . Substituting in Equation 34 gives all four bounds in Theorem 1.2. This completes the lower-bound proof.
Thus the separation between the two measurement models arises from their information budgets on the same geometric family: the respective traces are and . Conditioning on previous outcomes preserves the one-copy budget, so classical adaptation retains the factor- separation.
4 Pure states: matching collective bounds
For a pure input, the symmetric subspace gives a direct bound on posterior concentration. This supplies the case , where the preceding Sobolev inequality does not apply, and leads to a matching covariant estimator in every dimension.
Let and let be Haar on pure states. The -copy input lies in the symmetric subspace, of dimension
If every fixed estimate has a successful pure-state set of Haar measure at most , then any POVM has Haar-average success probability at most . Indeed, restricted to the symmetric subspace,
so integration over the success set of outcome , followed by summation, gives at most . The same proof uses integrals for continuous POVMs.
Apply Theorem 2.4 with , and use . Success probability at least implies, for a small universal ,
| (35) |
This proof works also when .
The same symmetry supplies an estimator attaining these rates. Its output is a measured direction with a small isotropic component on the orthogonal complement.
Use the covariant POVM of Hayashi [4] on the symmetric subspace
The Haar moment identity verifies its normalization. On input , the outcome density is . If , then
| (36) |
This follows immediately from the Haar density before the likelihood factor is applied.
For a parameter , output
| (37) |
Direct substitution in the two loss formulas gives
| (38) | ||||
| (39) |
The factor in the tail terms is the statistical cost of spreading output mass over uncertain orthogonal directions.
For , take for Bures chi-squared and . Markov’s inequality and Equation 36 give with probability at least . On that event,
For the right-inverse loss, take and . With probability at least , ; hence
Together with Equation 35, these estimators prove the pure-state claims of Theorem 1.3.
5 Upper-bound constructions
A mixed state can have several spectral scales. Following the block approach of [1, Section 3.3], we refine nested principal blocks and freeze eigenvalues once their scale has been resolved. The analysis charges each matrix entry at the stage when one of its indices is frozen; it does not require the global Frobenius error to decrease after every refinement. We first establish this accounting and then give one-copy and collective submatrix estimators. For right-inverse loss, isotropic regularization provides a separate reduction. Appendix C records the elementary block identities and counterexamples explaining why the spectral weights and final normalization must be tracked.
5.1 Shell errors and normalization
Suppose nested active subspaces are refined successively. At stage , a diagonalized estimate of the active true block satisfies
Freeze the coordinates whose estimated eigenvalues are at least , and allow later unitaries only on the remaining active subspace. The final positive diagonal matrix retains these frozen eigenvalues. The active coordinates can always be placed first, fixing a common coordinate convention for the successive blocks.
Lemma 5.1 (Shell accounting).
For each stage , assign to it all matrix entries with one index in the newly frozen shell and both indices in the active subspace at that stage. The sum of over these ordered entries is at most . Their contribution to
is therefore at most . Every entry outside the final active block is assigned exactly once.
Proof.
Immediately after diagonalization, the estimate has zero cross blocks between the frozen and remaining coordinates. It agrees with on the frozen block. Every later unitary preserves these zero cross blocks and rotates the true cross block unitarily. Its Frobenius norm, and the frozen-block error, are unchanged. Their squared norms are part of the original error . For an assigned entry at least one output eigenvalue is at least , so its weight is at most . ∎
The final active block is handled separately in the construction below. The shell estimate uses the Bures denominator : one frozen index bounds its reciprocal. For right-inverse loss, both indices matter, and we will instead use an isotropic regularization.
After all blocks have been estimated, the following identity controls the normalization of the output.
Lemma 5.2 (Exact normalization formula).
Let have trace one, let , and set . For either as above or ,
| (40) |
In particular, implies and .
Proof.
For the Bures case, use and , then expand the square. For the right-inverse case, expand . Weighted Cauchy–Schwarz applied to gives . Solving this quadratic inequality proves the stated bound on . ∎
5.2 A one-copy submatrix estimator
A covariant one-copy measurement gives an unbiased matrix observation for a subnormalized block. Matrix Bernstein concentration [15, Theorem 1.4] and rank truncation then provide the Frobenius estimate needed at each scale.
Let be a principal block of the current true state, let , and suppose and , where is known. On one full copy, project onto this block. If projection fails, record . If it succeeds, perform the covariant rank-one POVM on and, for outcome , record
| (41) |
This is implementable by choosing a Haar orthonormal basis and measuring in that basis. The unconditional probability density of is .
The second Haar moment gives
| (42) |
In particular, . For the average of independent observations, applying matrix Bernstein to both signs of the centered observations yields
| (43) |
Keep the largest positive eigenvalues of , or all of them if , and set the others to zero. Call the result . If , Weyl’s inequality shows , and hence
Thus we have proved the following submatrix routine, with all projection failures counted as consumed copies.
Lemma 5.3 (One-copy submatrix routine).
Let and . If , the zero estimate suffices. Otherwise the construction above outputs a positive semidefinite estimate with squared Frobenius error at most , with probability at least , using
| (44) |
copies. The rank bound can be replaced by . In the nontrivial case , the displayed bound is at least a positive constant, so rounding the copy count to an integer does not change its order.
For a trace-one state, a normalized output satisfies the same bound. Put and let be the eigenvalues of . Keep the first and replace them by , where makes their sum one; set the rest to zero. If the operator error is at most , Weyl’s inequality gives . At the retained sum is at most one, and at it is at least one, since . Thus a suitable exists. Every retained eigenvalue changes by at most ; every discarded eigenvalue has absolute value at most . The normalized estimate is therefore within operator distance of and has rank at most , giving the same squared Frobenius bound .
5.3 Multiscale refinement for Bures loss
Let , set
| (45) |
Initially the whole space is active. At stage :
- 1.
Estimate the current active block using Lemma 5.3, with error and failure probability .
- 2.
Diagonalize the estimate within the active subspace. Freeze all eigenvalues at least , retaining them as output entries, and put the remaining coordinates first for the next stage.
- 3.
If no coordinates remain, stop. After stage , assign every remaining coordinate the output value .
Let be the resulting diagonal matrix in the final accumulated basis. Return , undoing that accumulated basis change for the original coordinate system.
The mass bound used by the routine is and, for ,
| (46) |
To verify it, condition on all previous estimates being accurate. On the new active subspace, the previous estimate has operator norm below and . Let project onto the support of . Then
Thus the bound is valid conditionally at every stage. Fresh copies at each stage and a union bound on the first failed stage give simultaneous accuracy with probability at least .
To bound the loss of the final estimate, condition on the simultaneous accuracy event. By Lemma 5.1, all frozen shells contribute at most
to . If a final active block remains, its last estimate has eigenvalues in and is within Frobenius-squared error of the true block . Therefore
There is no tail term if all coordinates were frozen. In either case . By Lemma 5.2, normalization gives .
For the copy count, the decreasing mass of the active block compensates for the increasing accuracy required at later stages. For ,
Substitution in Equation 44, with , bounds a stage’s nonlogarithmic cost by
The first stage satisfies the same upper bound. Summing stages proves Equation 10.
5.4 Right-inverse loss by regularization
Lemma 5.4 (Frobenius error to right-inverse chi-squared).
Let be a density matrix and , . Then
| (47) |
In particular, and imply for .
Proof.
Use the squared-norm inequality in the inner product weighted by , splitting at . Since , the first term is at most . The matrices and commute, and their eigenvalues give
This proves Equation 47 and its stated specialization. ∎
Apply the normalized version of Lemma 5.3 with , , and . Its copy cost is
Regularize using Lemma 5.4. This proves Equation 12.
For the unrestricted class , the logarithm can be omitted at constant confidence. In the full-space measurement Equation 41,
Projection onto the convex set of density matrices cannot increase Frobenius distance to . Markov’s inequality and Lemma 5.4 therefore give copies directly.
5.5 Collective measurements and the Frobenius reduction
O’Donnell and Wright [8, Theorem 1.2] give a collective estimator in dimension whose output is a density matrix and satisfies
| (48) |
from copies of any trace-one state . Their output is , where is the observed partition of . To apply this estimate to a subnormalized block, we must learn its mass and count the copies rejected by projection.
Proposition 5.5 (Collective submatrix routine).
Let be a positive semidefinite block of dimension and mass . For and , a collective routine returns a positive semidefinite estimate with squared Frobenius error at most , with probability at least , using
| (49) |
full copies. If , no copies are required.
Proof.
If , the zero estimate already suffices. Otherwise , so integer rounding of the copy counts below is harmless. Project full copies onto the block. If , all projections fail and the zero output is exact, so assume . Let succeed; then . Conditional on , the retained copies have state . Apply Equation 48 to them and output ; output zero for . Conditioning on and expanding at gives
This accounts for the mass estimate and all rejected copies without a separate assumption that is known. Markov’s inequality supplies a Frobenius estimate with success probability at least at cost . For amplification, run independent batches whose individual estimates are within Frobenius distance of with probability at least . Choose an output whose ball of radius contains a strict majority of the outputs, choosing the first output if none qualifies. If a majority are within distance of , a qualifying output exists and every qualifying ball intersects that majority. Its center is therefore within of . A scalar Chernoff bound makes this event have probability at least after batches, proving the claim. ∎
Completion of the collective upper bounds.
Use Proposition 5.5 in the dyadic algorithm, with the same mass bounds, thresholds, and error accounting as before. For , the nonlogarithmic cost at stage is at most
The first stage satisfies the same bound. Summing over at most stages, each with failure probability , proves Equation 9. Later measurements may depend on previous outcomes, which is permitted by the unrestricted collective model.
For right-inverse loss, apply Equation 48 in dimension and use Markov’s inequality to obtain squared Frobenius error with constant success probability from copies. The same independent-batch selection used in Proposition 5.5, applied to normalized estimates, reduces the failure probability to at a factor in cost. The selected estimate remains a density matrix. Finally apply Lemma 5.4. This proves Equation 11 and completes Theorem 1.3. ∎
References
- [1] Steven T. Flammia and Ryan O’Donnell. Quantum chi-squared tomography and mutual information testing. Quantum 8, 1381 (2024). doi:10.22331/q-2024-06-20-1381.
- [2] Richard D. Gill and Serge Massar. State estimation for large ensembles. Physical Review A 61, 042312 (2000). doi:10.1103/PhysRevA.61.042312.
- [3] Jeongwan Haah, Aram W. Harrow, Zhengfeng Ji, Xiaodi Wu, and Nengkun Yu. Sample-optimal tomography of quantum states. IEEE Transactions on Information Theory 63(9), 5628–5641 (2017). doi:10.1109/TIT.2017.2719044.
- [4] Masahito Hayashi. Asymptotic estimation theory for a finite dimensional pure state model. Journal of Physics A: Mathematical and General 31(20), 4633–4655 (1998). doi:10.1088/0305-4470/31/20/006. Corrigendum: 31, 8405 (1998), doi:10.1088/0305-4470/31/41/015.
- [5] Elliott H. Lieb and Michael Loss. Analysis, second edition. American Mathematical Society, 2001.
- [6] Ufuk Keskin, Jason Luo, Mahbod Majid, and Matthew Radzihovsky. Tight lower bounds for state tomography with limited entanglement. arXiv:2609.05718 (2026). https://arxiv.org/abs/2609.05718.
- [7] Ashwin Nayak and Xingyu Zhou. Optimal low-rank quantum state tomography with bounded-sample joint measurements. arXiv:2609.10514 (2026). https://arxiv.org/abs/2609.10514.
- [8] Ryan O’Donnell and John Wright. Efficient quantum tomography. In Proceedings of the 48th Annual ACM Symposium on Theory of Computing (STOC), 899–912 (2016). doi:10.1145/2897518.2897544.
- [9] Ryan O’Donnell and John Wright. Efficient quantum tomography II. In Proceedings of the 49th Annual ACM SIGACT Symposium on Theory of Computing (STOC) (2017). doi:10.1145/3055399.3055454. Preprint: https://arxiv.org/abs/1612.00034.
- [10] Dénes Petz. Monotone metrics on matrix spaces. Linear Algebra and its Applications 244, 81–96 (1996). doi:10.1016/0024-3795(94)00211-8.
- [11] Thilo Scharnhorst, Jack Spilecki, and John Wright. Optimal lower bounds for quantum state tomography. arXiv:2510.07699 (2025). https://arxiv.org/abs/2510.07699.
- [12] Fuyuhiko Tanaka. Generalized Bayesian predictive density operators. arXiv:quant-ph/0602072 (2006). https://arxiv.org/abs/quant-ph/0602072.
- [13] Sisi Zhou and Senrui Chen. Randomized measurements for multiparameter quantum metrology. PRX Quantum 7, 010314 (2026). doi:10.1103/s27y-gbrp.
- [14] Kristan Temme, Michael J. Kastoryano, Mary Beth Ruskai, Michael M. Wolf, and Frank Verstraete. The -divergence and mixing times of quantum Markov processes. Journal of Mathematical Physics 51, 122201 (2010).
- [15] Joel A. Tropp. User-friendly tail bounds for sums of random matrices. Foundations of Computational Mathematics 12, 389–434 (2012). doi:10.1007/s10208-011-9099-z.
Appendix A Exact pure-state risk
For a fixed measurement, the Bayes-optimal output for right-inverse loss is the normalized square root of the posterior second moment. This is the specialization of Tanaka’s generalized Bayesian predictive rule [12]. On pure inputs, the second moment equals the posterior mean state. We evaluate the resulting risk and optimize over all measurements, complementing the constant-success bounds of Section 4.
Theorem A.1 (Minimax expected right-inverse loss for pure states).
For copies of an unknown pure state in dimension and arbitrary collective measurements,
| (50) |
Proof.
Use the Haar prior. Restrict a POVM element to the symmetric subspace; discard elements with zero trace. For , its posterior mean state is
| (51) |
To verify this, use the Haar moment identity on copies. On the first symmetric factors,
where swaps the indicated factors. Partial trace against gives . Finally, supplies the denominator in Equation 51.
Since a pure projector squares to itself, the conditional expected loss is . Since , a singular output has infinite conditional loss. Matrix Cauchy–Schwarz gives
Concavity of implies that its minimum over density matrices is attained at a pure state; its value there is . Thus every measurement has Haar Bayes risk at least the right-hand side of Equation 50.
The covariant POVM of Section 4 has . Returning
achieves the lower bound for every outcome. The procedure is covariant, so its risk is the same for every true pure state. Its Bayes risk is therefore also its worst-case risk, proving minimax equality. For , the Haar posterior mean is , so the same Cauchy–Schwarz bound gives risk at least . The uniform estimate attains it, agreeing with the formula. ∎
For much larger than , the expression is asymptotic to . This agrees with the copy scale, while also showing why a bare infidelity-to-chi-squared conversion can miss a dimension factor.
Appendix B Common centers and confidence bounds
Proposition B.1 (Common center of two flat states).
Let be rank- projectors with principal angles , in dimension at least . Then
| (52) |
The same minimum is obtained if the maximum is replaced by the arithmetic mean. A minimizing state is proportional to on its support.
Proof.
The average loss plus one is
Its minimum is the squared trace of the square root of , by the same Cauchy–Schwarz calculation used above. In a principal-angle decomposition, the eigenvalues of are , including zero eigenvalues when an angle is zero. Therefore the minimum equals the displayed expression. At the proposed minimizer the two losses are equal: the principal-angle decomposition provides a unitary swapping and and leaving fixed. Hence the minimum of the maximum is the same as that of the average.
Use for , and , to obtain the inequality. ∎
For two pure states this radius is exactly . This differs from the quadratic small-angle behavior of squared Bures distance, and is another direct way to see why the two chi-squared definitions cannot be interchanged.
Dependence on confidence.
The main statements use success probability . A simple two-point argument adds a necessary confidence term for . Choose two pure states with angle and write . Their -copy optimal equal-prior discrimination error is
If their two loss balls are disjoint, a learner failing with probability at most on each state would discriminate them with error at most . Thus
| (53) |
For Bures chi-squared, take ; the trace-norm inequality in Lemma 1.4 makes the loss balls disjoint. For right-inverse chi-squared, take and apply Proposition B.1 with . For small , Equation 53 gives, respectively,
Every rank-at-most- class contains these pure states. Combining these bounds with Theorem 1.2 adds to the corresponding dimension-dependent numerator, up to constants. These give necessary confidence terms in addition to the constant-success bounds.
Appendix C Block replacement and spectral weights
The shell argument in Section 5 avoids two pitfalls of local Frobenius refinement. First, replacing a block need not decrease the global error. Write
where and are diagonal. After conjugation by , replace by a new diagonal block and set . With , block orthogonality gives
| (54) |
The cross-block norm is unchanged because . Thus improvement requires comparison with the old error on the replaced block, not with the old global error. For example,
are density matrices. The old global error is and the new first-block error is , but the new global error is , since the old first-block error was zero.
Trace preservation is separate: . For instance, replacing the first entry of by produces trace , even when is the correct true entry. Positive replacement blocks preserve positivity but not normalization; Lemma 5.2 accounts for the final normalization.
Second, Frobenius error alone cannot uniformly bound either chi-squared loss. For fixed , take
Then
| (55) |
as . The eigenvalue thresholds in the Bures construction, and the regularization floor in the right-inverse construction, supply the missing denominator control.