physics2u
Tier
⌕ Search ⌘K
Derivation

The Holevo Bound

Statement

For a quantum ensemble ε = {px, ρx} prepared with prior probabilities px and read out by an arbitrary POVM {Ey} producing outcome Y given the label X, the accessible information obeys I(X:Y) ≤ χ(ε) = S(ÏÌ„) − ∑x px Sx), where ÏÌ„ = ∑x px ρx is the average state and S(·) is the von Neumann entropy. The Holevo quantity χ is itself bounded by log2 d for a d-dimensional system, so n qubits carry at most n bits of extractable classical information.

Why it matters

The Holevo bound is the fundamental limit on how much classical information one can recover from a quantum carrier. Even though a qubit lives in a continuous state space with infinitely many distinguishable-looking preparations, no measurement can reliably extract more than one bit from it. This single inequality kills the naive dream of packing unlimited classical data into a single quantum degree of freedom, and it underwrites the security of quantum key distribution and the capacity theorems of quantum Shannon theory.

Operationally it separates what is encoded (the ensemble) from what is readable (the accessible information), and it does so with a quantity, χ, that depends only on the states and priors — not on the measurement. That measurement-independence is what makes it a genuine capacity bound rather than a statement about one detector.

Assumptions
Finite-dimensional Hilbert spaces.In infinite dimensions S(ρ) can diverge and the entropies need not be finite; the chain of inequalities still holds under trace-class conditions but requires limiting arguments rather than the finite sums used here.
The ensemble is fixed and the priors px are classical.If the label register is itself entangled with the system rather than classically correlated, "accessible information" is not defined by a POVM on the system alone and the statement must be recast.
Strong subadditivity of von Neumann entropy holds.The whole proof rests on this deep inequality (a prior result). Drop it and the monotonicity of quantum mutual information under local operations fails, and no bound survives — classical entropy's subadditivity is not enough.
The measurement is a completely-positive trace-preserving readout (Kraus/Stinespring form).If the "measurement" could act non-physically (non-CP maps), data-processing monotonicity would not apply and χ would not bound anything.
Derivation
1
ÏXQ = ∑x px |x⟩⟨x|X ⊗ ρx
Build a classical-quantum state on a label register X plus the quantum system Q. The orthonormal kets |x⟩ make X a faithful classical record of which state was prepared. A
2
S(X) = H({px}),   S(Q) = S(ÏÌ„),   S(XQ) = H({px}) + ∑x px Sx)
Marginals of the cq-state. The joint entropy splits because the blocks pxρx sit in orthogonal label sectors, so S of a block-diagonal state = classical entropy of weights + weighted average of block entropies. B
3
I(X:Q) = S(X) + S(Q) − S(XQ)
Definition of quantum mutual information. This measures the total correlation between label and system before any readout. A
4
I(X:Q) = [H + S(ÏÌ„)] − [H + ∑x px Sx)] = S(ÏÌ„) − ∑x px Sx) ≡ χ(ε)
Substitute step 2 into step 3. The label entropy H({px}) cancels, leaving exactly the Holevo quantity. So χ is the input-side mutual information of the cq-state. B
5
εy(σ) = ∑k My,k σ My,k†,   Ey = ∑k My,k†My,k,   ∑y Ey = ð•€
Represent the POVM {Ey} as a CPTP measurement map that writes the outcome into a fresh classical register Y (Kraus/Stinespring representation, a prior result). Any physical measurement admits such a form. B
6
Λ: Q → Y,   Λ(σ) = ∑y [Tr(Ey σ)] |y⟩⟨y|Y
The measurement acts only on Q, leaving the label register X untouched: it is a local CPTP map (idX ⊗ Λ). This is the key structural fact enabling data processing. B
7
I(X:Y) ≤ I(X:Q)
Quantum data-processing inequality: applying a local CPTP map to one subsystem cannot increase mutual information. This is equivalent to strong subadditivity of von Neumann entropy (the assumed prior result); it is the deepest step and cannot be replaced by any classical argument. C
8
I(X:Y) ≤ χ(ε) = S(ÏÌ„) − ∑x px Sx)
Chain step 7 with step 4. Because after measurement X and Y are both classical registers, I(X:Y) is precisely the classical Shannon mutual information between prepared label and readout — the accessible information. A
Result
I(X:Y) ≤ χ(ε) = S(ÏÌ„) − ∑x px Sx) ≤ log2 d

Reading. The classical mutual information any measurement can create between the prepared label and its outcome is capped by the Holevo χ — the entropy of the average state minus the average entropy of the members. Since S(ÏÌ„) ≤ log2 d and the average entropy is ≥ 0, a d-level system yields at most log2 d bits; a qubit (d = 2) yields at most one bit, no matter how cleverly encoded or measured.

Units check. Entropies are dimensionless, measured in bits when the logarithm is base 2 (nats in base e). χ is a difference of entropies, hence also bits; log2 d is bits. Mutual information I(X:Y) is bits. Both sides carry the same unit, so the inequality is dimensionally consistent.

Limiting cases
  • Orthogonal pure states. If the ρx are mutually orthogonal pure states, each Sx) = 0 and S(ÏÌ„) = H({px}), so χ = H({px}) and the bound is saturated by projective measurement: full classical information is recoverable.
  • Identical states. If all ρx = Ï, then ÏÌ„ = Ï and χ = S(Ï) − S(Ï) = 0: no information about X is encoded, so I(X:Y) = 0 for every measurement.
  • Pure but non-orthogonal states. Each Sx) = 0 so χ = S(ÏÌ„) < H({px}); overlap strictly reduces the extractable information below the classical prior entropy.
  • Large alphabet in fixed dimension. As the number of signal states grows with d fixed, χ → log2 d at most; adding more preparations cannot push accessible information past the dimensional ceiling.
Breaks when
  • The receiver measures collectively across many copies. Holevo bounds the information per use for single-shot measurement, but the Holevo capacity (HSW theorem) shows χ is achievable only asymptotically with joint measurements over long blocks; for a single copy the accessible information is generally strictly below χ, and treating χ as attainable per-shot is wrong.
  • Post-selection or side information changes the ensemble. If outcomes are conditioned on later events (e.g. entanglement-assisted or feed-forward protocols where the label is not classical), the cq-state construction fails and the plain bound no longer constrains the extractable correlations — superdense coding beats "one bit per qubit" precisely because a shared entangled resource sits outside this ensemble.
  • Infinite-dimensional systems with unbounded energy. With no energy constraint, S(ÏÌ„) can diverge and log2 d is meaningless; the naive dimensional ceiling disappears and one must impose a mean-energy constraint to recover a finite capacity.
Failure modes
  • Confusing χ with achievable single-shot information. Students treat I(X:Y) = χ as generically reachable; it is only an upper bound, saturated in the single-copy setting solely for commuting/orthogonal ensembles.
  • Adding member entropies with the wrong sign. Writing χ = S(ÏÌ„) + ∑ px Sx) instead of minus. The average entropy is subtracted; mixedness of the signals hurts, it does not help.
  • Using classical entropy of the eigenvalues of ÏÌ„ but forgetting off-diagonal coherence. S(ÏÌ„) is the entropy of ÏÌ„'s eigenvalues, not of its diagonal in the preparation basis; using the diagonal overcounts and violates concavity.
  • Assuming a rank-1 POVM is always optimal. Optimal accessible-information measurements can require more than d outcomes, but never more than d²; guessing that the computational-basis measurement is optimal underestimates I(X:Y).
  • Thinking a qubit "stores infinite information" because its state is continuous. The continuum of amplitudes is not readable; the bound is exactly the statement that read-out, not storage, is limited to one bit.
Discussion

The proof's heart is the identification χ = I(X:Q) for a classical-quantum state, followed by the data-processing inequality when the quantum system is measured. Nothing about the specific POVM enters until the very last step, which is why the bound is measurement-independent: it is a property of the ensemble, computed once, that no detector can beat. The whole argument is really a corollary of strong subadditivity of the von Neumann entropy — arguably the deepest inequality in quantum information — dressed up as a statement about communication.

Physically, the bound quantifies a trade-off between distinguishability and disturbance. Non-orthogonal signal states cannot be perfectly distinguished, and the "cost" of their overlap is booked as the gap between S(ÏÌ„) and the classical prior entropy H({px}). Mixed signal states cost further, through the subtracted average entropy: any noise the sender puts into the individual ρx is information the receiver can never recover. This is the information-theoretic face of the no-cloning theorem and the basis of eavesdropper bounds in BB84 and its relatives — Eve's information about the key is itself Holevo-bounded.

The bound also frames the correct notion of quantum channel capacity. The Holevo–Schumacher–Westmoreland theorem promotes χ from an upper bound to an achievable rate, but only with collective measurements over asymptotically many channel uses and only after maximizing χ over input ensembles. The single-letter χ can even be super-additive for some channels, which is why classical capacity of a general quantum channel is not simply the one-shot Holevo quantity — a subtlety that took two decades and the resolution of additivity conjectures to fully understand.

At the sharpest level, the data-processing step (step 7) is equivalent to strong subadditivity: S(ABC) + S(B) ≤ S(AB) + S(BC). Monotonicity of quantum relative entropy under CPTP maps, S(Î›Ï â€–Î›σ) ≤ S(Ï‖σ), is the modern lens: writing χ = ∑x px Sx ‖ ÏÌ„) exhibits Holevo as an average relative entropy to the barycenter, and the measurement map's monotonicity of relative entropy immediately yields the bound. The recent theory of recoverability (Fawzi–Renner) refines the inequality with a remainder term measuring how far the ensemble is from a state the measurement could undo.

Common misconceptions. χ is not the mutual information you obtain — it is the ceiling. It is also not the entanglement of the ensemble (there is no entanglement in a cq-state); it is a correlation measure. And "one bit per qubit" refers to information extracted from an unentangled carrier — superdense coding delivers two bits per transmitted qubit precisely because it spends a pre-shared entangled qubit, which does not violate the bound applied to the actual ensemble sent.

Worked examples
1
Two non-orthogonal qubit states, equal priors. ρ0 = |0⟩⟨0|,   ρ1 = |+⟩⟨+|,   p0 = p1 = ½
Pure signal states, so ∑ px Sx) = 0 and χ = S(ÏÌ„). A
2
ÏÌ„ = ½|0⟩⟨0| + ½|+⟩⟨+| = [ ¾   ¼ ;   ¼   ¼ ]
Average the two rank-1 projectors in the computational basis. B
3
λ± = ½ ± ½·(1/√2) = ½(1 ± 0.7071) ⇒ λ+ = 0.8536,   λ− = 0.1464
Eigenvalues of ÏÌ„: for a qubit, λ± = ½(1 ± |r|) with Bloch radius |r| = 1/√2 here (average of Bloch vectors ẑ and x̂ gives length √(¼+¼) = 1/√2). B
4
χ = S(ÏÌ„) = −0.8536 logâ‚‚ 0.8536 − 0.1464 logâ‚‚ 0.1464 = 0.1955 + 0.4060 = 0.6009 bits
Binary entropy of the eigenvalues. A
I(X:Y) ≤ χ ≈ 0.601 bits

Reading. Despite equiprobable labels (1 bit of prior entropy), the 45° overlap of the two states caps recoverable information at about 0.60 bits per qubit. The optimal (Helstrom-type) measurement approaches this ceiling; the computational-basis measurement alone gives only I ≈ 0.40 bits, so it is sub-optimal.

1
Depolarized BB84-style signals. Four pure states {|0⟩,|1⟩,|+⟩,|−⟩}, each sent with p = ¼, then passed through a depolarizing channel of strength f = 0.2: ρx = (1−f)|ψx⟩⟨ψx| + f·(ð•€/2)
Now the signals are mixed, so both terms of χ are non-trivial. B
2
ÏÌ„ = ∑x ¼ ρx = (1−f)·(ð•€/2) + f·(ð•€/2) = ð•€/2 ⇒ S(ÏÌ„) = 1 bit
The four BB84 states are symmetric on the Bloch sphere; their equal mixture is maximally mixed regardless of f. B
3
Each ρx has Bloch radius |r| = 1−f = 0.8 ⇒ λ± = ½(1 ± 0.8) = 0.9, 0.1
Depolarizing shrinks every Bloch vector by (1−f); all four members have identical eigenvalues, hence identical entropy. A
4
Sx) = −0.9 log₂ 0.9 − 0.1 log₂ 0.1 = 0.1368 + 0.3322 = 0.4690 bits
Binary entropy at eigenvalues (0.9, 0.1); same for all x so the average equals this value. A
5
χ = S(ÏÌ„) − ∑x ¼ Sx) = 1 − 0.4690 = 0.5310 bits
Subtract the average member entropy from the average-state entropy. A
I(X:Y) ≤ χ ≈ 0.531 bits

Reading. Channel noise (f = 0.2) has cut the accessible information from the noiseless ceiling of 1 bit down to about 0.53 bits: the maximally-mixed average keeps S(ÏÌ„) = 1, but the mixedness injected into each signal is subtracted directly. At f = 0 one recovers χ = 1 bit; at f = 1 the signals become identical (ð•€/2) and χ = 0.

Problems
  1. An ensemble sends |0⟩ and |1⟩ with probabilities p and 1−p. Show χ equals the binary entropy H(p) and state the maximal accessible information.
    Solution Both states are pure and orthogonal, so Sx) = 0 and ÏÌ„ = diag(p, 1−p). Its eigenvalues are p and 1−p, so χ = S(ÏÌ„) = −p logâ‚‚ p − (1−p) logâ‚‚(1−p) = H(p). Because the states are orthogonal, the projective measurement in {|0⟩,|1⟩} perfectly distinguishes them, giving I(X:Y) = H(p) = χ: the bound is saturated. For p = ½, χ = 1 bit.
  2. Two states |0⟩ and cosθ|0⟩+sinθ|1⟩ are sent with equal priors. Find χ as a function of θ and evaluate at θ = 60°.
    Solution Pure states ⇒ χ = S(ÏÌ„). The overlap is |⟨ψ01⟩| = cosθ. Bloch vectors are ẑ and (sin2θ, 0, cos2θ); their average has length |r| = ½√(sin²2θ + (1+cos2θ)²) = ½√(2+2cos2θ) = |cosθ|. So λ± = ½(1 ± cosθ) and χ = H(½(1+cosθ)). At θ = 60°, cosθ = 0.5, λ± = 0.75, 0.25, χ = −0.75 logâ‚‚0.75 − 0.25 logâ‚‚0.25 = 0.3113 + 0.5 = 0.8113 bits.
  3. Three states are sent with equal priors: |0⟩, |1⟩, and (|0⟩+|1⟩)/√2. Compute χ (in bits).
    Solution All pure, so χ = S(ÏÌ„). ÏÌ„ = â…“|0⟩⟨0| + â…“|1⟩⟨1| + â…“|+⟩⟨+| = [ â…“+â…™ , â…™ ; â…™ , â…“+â…™ ] = [ ½ , â…™ ; â…™ , ½ ]. Eigenvalues ½ ± â…™ = 0.6667, 0.3333. χ = −0.6667 logâ‚‚0.6667 − 0.3333 logâ‚‚0.3333 = 0.3900 + 0.5283 = 0.9183 bits. (Below log₂3 ≈ 1.585 and below log₂2 = 1: the three non-orthogonal states in a 2-D space cannot exceed 1 bit.)
  4. A qutrit (d = 3) is used to send three mutually orthogonal pure states with equal priors. What is χ, and what does this say about the number of bits per qutrit?
    Solution Orthogonal pure states ⇒ Sx) = 0 and ÏÌ„ = (1/3)ð•€3, maximally mixed. S(ÏÌ„) = log₂3 = 1.585 bits. So χ = 1.585 bits = log₂3, saturating the dimensional ceiling log₂d. A qutrit therefore carries at most log₂3 ≈ 1.585 bits, achieved here by a simple projective measurement in the signal basis.
  5. A qubit ensemble has ÏÌ„ = ð•€/2 and each signal is a pure state (as in BB84). If a depolarizing channel with parameter f acts before measurement, find the value of f at which χ drops to 0.5 bits.
    Solution As derived in worked example 2, S(ÏÌ„) = 1 (unchanged by symmetric depolarizing) and each signal has Bloch radius 1−f, giving Sx) = H(½(1+(1−f))) = H(1−f/2). So χ = 1 − H(1−f/2). Set χ = 0.5 ⇒ H(1−f/2) = 0.5. Solving H(q) = 0.5 gives q ≈ 0.8900 (or 0.1100). Take q = 1−f/2 = 0.8900 ⇒ f/2 = 0.1100 ⇒ f ≈ 0.220. So at about 22% depolarization the accessible-information ceiling halves to 0.5 bits.