The Holevo Bound
Statement
For a quantum ensemble ε = {px, ρx} prepared with prior probabilities px and read out by an arbitrary POVM {Ey} producing outcome Y given the label X, the accessible information obeys I(X:Y) ≤ χ(ε) = S(ÏÌ„) − ∑x px S(ρx), where ÏÌ„ = ∑x px ρx is the average state and S(·) is the von Neumann entropy. The Holevo quantity χ is itself bounded by log2 d for a d-dimensional system, so n qubits carry at most n bits of extractable classical information.
Why it matters
The Holevo bound is the fundamental limit on how much classical information one can recover from a quantum carrier. Even though a qubit lives in a continuous state space with infinitely many distinguishable-looking preparations, no measurement can reliably extract more than one bit from it. This single inequality kills the naive dream of packing unlimited classical data into a single quantum degree of freedom, and it underwrites the security of quantum key distribution and the capacity theorems of quantum Shannon theory.
Operationally it separates what is encoded (the ensemble) from what is readable (the accessible information), and it does so with a quantity, χ, that depends only on the states and priors — not on the measurement. That measurement-independence is what makes it a genuine capacity bound rather than a statement about one detector.
Assumptions
Derivation
Result
Reading. The classical mutual information any measurement can create between the prepared label and its outcome is capped by the Holevo χ — the entropy of the average state minus the average entropy of the members. Since S(ÏÌ„) ≤ log2 d and the average entropy is ≥ 0, a d-level system yields at most log2 d bits; a qubit (d = 2) yields at most one bit, no matter how cleverly encoded or measured.
Units check. Entropies are dimensionless, measured in bits when the logarithm is base 2 (nats in base e). χ is a difference of entropies, hence also bits; log2 d is bits. Mutual information I(X:Y) is bits. Both sides carry the same unit, so the inequality is dimensionally consistent.
Limiting cases
- Orthogonal pure states. If the ρx are mutually orthogonal pure states, each S(ρx) = 0 and S(ÏÌ„) = H({px}), so χ = H({px}) and the bound is saturated by projective measurement: full classical information is recoverable.
- Identical states. If all ρx = Ï, then ÏÌ„ = Ï and χ = S(Ï) − S(Ï) = 0: no information about X is encoded, so I(X:Y) = 0 for every measurement.
- Pure but non-orthogonal states. Each S(ρx) = 0 so χ = S(ÏÌ„) < H({px}); overlap strictly reduces the extractable information below the classical prior entropy.
- Large alphabet in fixed dimension. As the number of signal states grows with d fixed, χ → log2 d at most; adding more preparations cannot push accessible information past the dimensional ceiling.
Breaks when
- The receiver measures collectively across many copies. Holevo bounds the information per use for single-shot measurement, but the Holevo capacity (HSW theorem) shows χ is achievable only asymptotically with joint measurements over long blocks; for a single copy the accessible information is generally strictly below χ, and treating χ as attainable per-shot is wrong.
- Post-selection or side information changes the ensemble. If outcomes are conditioned on later events (e.g. entanglement-assisted or feed-forward protocols where the label is not classical), the cq-state construction fails and the plain bound no longer constrains the extractable correlations — superdense coding beats "one bit per qubit" precisely because a shared entangled resource sits outside this ensemble.
- Infinite-dimensional systems with unbounded energy. With no energy constraint, S(ÏÌ„) can diverge and log2 d is meaningless; the naive dimensional ceiling disappears and one must impose a mean-energy constraint to recover a finite capacity.
Failure modes
- Confusing χ with achievable single-shot information. Students treat I(X:Y) = χ as generically reachable; it is only an upper bound, saturated in the single-copy setting solely for commuting/orthogonal ensembles.
- Adding member entropies with the wrong sign. Writing χ = S(ÏÌ„) + ∑ px S(ρx) instead of minus. The average entropy is subtracted; mixedness of the signals hurts, it does not help.
- Using classical entropy of the eigenvalues of ÏÌ„ but forgetting off-diagonal coherence. S(ÏÌ„) is the entropy of ÏÌ„'s eigenvalues, not of its diagonal in the preparation basis; using the diagonal overcounts and violates concavity.
- Assuming a rank-1 POVM is always optimal. Optimal accessible-information measurements can require more than d outcomes, but never more than d²; guessing that the computational-basis measurement is optimal underestimates I(X:Y).
- Thinking a qubit "stores infinite information" because its state is continuous. The continuum of amplitudes is not readable; the bound is exactly the statement that read-out, not storage, is limited to one bit.
Discussion
The proof's heart is the identification χ = I(X:Q) for a classical-quantum state, followed by the data-processing inequality when the quantum system is measured. Nothing about the specific POVM enters until the very last step, which is why the bound is measurement-independent: it is a property of the ensemble, computed once, that no detector can beat. The whole argument is really a corollary of strong subadditivity of the von Neumann entropy — arguably the deepest inequality in quantum information — dressed up as a statement about communication.
Physically, the bound quantifies a trade-off between distinguishability and disturbance. Non-orthogonal signal states cannot be perfectly distinguished, and the "cost" of their overlap is booked as the gap between S(ÏÌ„) and the classical prior entropy H({px}). Mixed signal states cost further, through the subtracted average entropy: any noise the sender puts into the individual ρx is information the receiver can never recover. This is the information-theoretic face of the no-cloning theorem and the basis of eavesdropper bounds in BB84 and its relatives — Eve's information about the key is itself Holevo-bounded.
The bound also frames the correct notion of quantum channel capacity. The Holevo–Schumacher–Westmoreland theorem promotes χ from an upper bound to an achievable rate, but only with collective measurements over asymptotically many channel uses and only after maximizing χ over input ensembles. The single-letter χ can even be super-additive for some channels, which is why classical capacity of a general quantum channel is not simply the one-shot Holevo quantity — a subtlety that took two decades and the resolution of additivity conjectures to fully understand.
At the sharpest level, the data-processing step (step 7) is equivalent to strong subadditivity: S(ABC) + S(B) ≤ S(AB) + S(BC). Monotonicity of quantum relative entropy under CPTP maps, S(Î›Ï â€–Î›σ) ≤ S(Ï‖σ), is the modern lens: writing χ = ∑x px S(ρx ‖ ÏÌ„) exhibits Holevo as an average relative entropy to the barycenter, and the measurement map's monotonicity of relative entropy immediately yields the bound. The recent theory of recoverability (Fawzi–Renner) refines the inequality with a remainder term measuring how far the ensemble is from a state the measurement could undo.
Common misconceptions. χ is not the mutual information you obtain — it is the ceiling. It is also not the entanglement of the ensemble (there is no entanglement in a cq-state); it is a correlation measure. And "one bit per qubit" refers to information extracted from an unentangled carrier — superdense coding delivers two bits per transmitted qubit precisely because it spends a pre-shared entangled qubit, which does not violate the bound applied to the actual ensemble sent.
Worked examples
Reading. Despite equiprobable labels (1 bit of prior entropy), the 45° overlap of the two states caps recoverable information at about 0.60 bits per qubit. The optimal (Helstrom-type) measurement approaches this ceiling; the computational-basis measurement alone gives only I ≈ 0.40 bits, so it is sub-optimal.
Reading. Channel noise (f = 0.2) has cut the accessible information from the noiseless ceiling of 1 bit down to about 0.53 bits: the maximally-mixed average keeps S(ÏÌ„) = 1, but the mixedness injected into each signal is subtracted directly. At f = 0 one recovers χ = 1 bit; at f = 1 the signals become identical (ð•€/2) and χ = 0.
Problems
- An ensemble sends |0⟩ and |1⟩ with probabilities p and 1−p. Show χ equals the binary entropy H(p) and state the maximal accessible information.
Solution
Both states are pure and orthogonal, so S(ρx) = 0 and ÏÌ„ = diag(p, 1−p). Its eigenvalues are p and 1−p, so χ = S(ÏÌ„) = −p logâ‚‚ p − (1−p) logâ‚‚(1−p) = H(p). Because the states are orthogonal, the projective measurement in {|0⟩,|1⟩} perfectly distinguishes them, giving I(X:Y) = H(p) = χ: the bound is saturated. For p = ½, χ = 1 bit. - Two states |0⟩ and cosθ|0⟩+sinθ|1⟩ are sent with equal priors. Find χ as a function of θ and evaluate at θ = 60°.
Solution
Pure states ⇒ χ = S(ÏÌ„). The overlap is |⟨ψ0|ψ1⟩| = cosθ. Bloch vectors are ẑ and (sin2θ, 0, cos2θ); their average has length |r| = ½√(sin²2θ + (1+cos2θ)²) = ½√(2+2cos2θ) = |cosθ|. So λ± = ½(1 ± cosθ) and χ = H(½(1+cosθ)). At θ = 60°, cosθ = 0.5, λ± = 0.75, 0.25, χ = −0.75 logâ‚‚0.75 − 0.25 logâ‚‚0.25 = 0.3113 + 0.5 = 0.8113 bits. - Three states are sent with equal priors: |0⟩, |1⟩, and (|0⟩+|1⟩)/√2. Compute χ (in bits).
Solution
All pure, so χ = S(ÏÌ„). ÏÌ„ = â…“|0⟩⟨0| + â…“|1⟩⟨1| + â…“|+⟩⟨+| = [ â…“+â…™ , â…™ ; â…™ , â…“+â…™ ] = [ ½ , â…™ ; â…™ , ½ ]. Eigenvalues ½ ± â…™ = 0.6667, 0.3333. χ = −0.6667 logâ‚‚0.6667 − 0.3333 logâ‚‚0.3333 = 0.3900 + 0.5283 = 0.9183 bits. (Below log₂3 ≈ 1.585 and below log₂2 = 1: the three non-orthogonal states in a 2-D space cannot exceed 1 bit.) - A qutrit (d = 3) is used to send three mutually orthogonal pure states with equal priors. What is χ, and what does this say about the number of bits per qutrit?
Solution
Orthogonal pure states ⇒ S(ρx) = 0 and ÏÌ„ = (1/3)ð•€3, maximally mixed. S(ÏÌ„) = log₂3 = 1.585 bits. So χ = 1.585 bits = log₂3, saturating the dimensional ceiling log₂d. A qutrit therefore carries at most log₂3 ≈ 1.585 bits, achieved here by a simple projective measurement in the signal basis. - A qubit ensemble has ÏÌ„ = ð•€/2 and each signal is a pure state (as in BB84). If a depolarizing channel with parameter f acts before measurement, find the value of f at which χ drops to 0.5 bits.
Solution
As derived in worked example 2, S(ÏÌ„) = 1 (unchanged by symmetric depolarizing) and each signal has Bloch radius 1−f, giving S(ρx) = H(½(1+(1−f))) = H(1−f/2). So χ = 1 − H(1−f/2). Set χ = 0.5 ⇒ H(1−f/2) = 0.5. Solving H(q) = 0.5 gives q ≈ 0.8900 (or 0.1100). Take q = 1−f/2 = 0.8900 ⇒ f/2 = 0.1100 ⇒ f ≈ 0.220. So at about 22% depolarization the accessible-information ceiling halves to 0.5 bits.