physics2u
Tier
⌕ Search ⌘K
physics2u.com/vault/von-neumann-entropy-properties.html
Derivation

Von Neumann Entropy and Subadditivity

D-416 Home PU-404 Threads chance · energy Depends on shannon-entropy-from-multiplicity, Reduced States and the Partial Trace, spectral-theorem-hermitian-operators
Statement

For a density operator ρ on a finite-dimensional Hilbert space, the von Neumann entropy S(ρ) = −Tr(ρ ln ρ) equals the Shannon entropy −∑k pk ln pk of its eigenvalue distribution, is invariant under unitaries, is concave in ρ, and for a bipartite state ρAB obeys subadditivity S(ρAB) ≤ S(ρA) + S(ρB) together with the Araki–Lieb lower bound S(ρAB) ≥ |S(ρA) − S(ρB)|.

Why it matters

The von Neumann entropy is the quantum generalisation of the Boltzmann–Gibbs–Shannon entropy: it is the unique measure (up to normalisation) that quantifies the missing information in a quantum state, sets the thermodynamic entropy of a system described by ρ, and gives the asymptotic compression limit of a quantum source. Its structural inequalities — concavity, subadditivity, the Araki–Lieb bound — are the load-bearing facts behind the second law for quantum systems and behind the theory of entanglement.

Subadditivity in particular is what makes entropy an extensive-but-correlated quantity: correlations between subsystems can only reduce the total entropy below the sum of the parts. The gap S(ρA) + S(ρB) − S(ρAB) = I(A:B) is the quantum mutual information, the single most useful measure of total correlation.

Assumptions
The state is a valid density operator: ρ = ρ, ρ ≥ 0, Tr ρ = 1. Without positivity the eigenvalues need not lie in [0,1] and ln pk becomes complex or undefined; the entropy loses its interpretation as missing information and can turn negative or non-real. Finite dimension (or trace-class with summable spectrum): the spectrum is discrete and k pk = 1 converges. In infinite dimensions with continuous spectrum the sum becomes an integral that may diverge; S can be +∞ and several inequalities require extra regularity. The function f(x) = −x ln x is applied through the spectral calculus: S(ρ) = Tr f(ρ) with f(0) ≡ 0 by continuity. If the operator function is not defined via the spectral theorem (e.g. applied entrywise in a non-eigenbasis) the trace no longer equals k f(pk) and every result below fails. Concavity uses that −x ln x is operator-concave and that relative entropy is jointly convex: the subadditivity proof rests on S(ρAB ‖ ρA ⊗ ρB) ≥ 0 (Klein’s inequality). Drop non-negativity of relative entropy and subadditivity is no longer guaranteed.
Derivation
1
ρ = ∑k pk |k⟩⟨k| ,   pk ≥ 0 ,   ∑k pk = 1
Spectral theorem for the Hermitian, positive, unit-trace operator ρ: it has an orthonormal eigenbasis with real non-negative eigenvalues summing to one. A
2
f(ρ) = ∑k f(pk) |k⟩⟨k| ,   f(x) = −x ln x
Definition of a function of an operator via the spectral calculus: f acts on the eigenvalues while eigenprojectors are preserved. Here f(0)=0 by the limit x ln x → 0. A
3
S(ρ) = Tr f(ρ) = −Tr(ρ ln ρ) = −∑k pk ln pk
Take the trace of step 2 in the eigenbasis; the trace of a diagonal operator is the sum of its diagonal entries. This is exactly the Shannon entropy of the distribution {pk}. A
4
S(UρU) = −Tr(UρU ln(UρU)) = −Tr(U ρ ln ρ U) = S(ρ)
A unitary conjugation permutes nothing in the spectrum: ln(UρU) = U (ln ρ) U, and the trace is cyclic with UU = 1. Entropy depends only on eigenvalues. A
5
S(λρ1 + (1−λ)ρ2) ≥ λS(ρ1) + (1−λ)S(ρ2) ,   0 ≤ λ ≤ 1
Concavity. The map x ↦ −x ln x is operator-concave, and ρ ↦ Tr f(ρ) inherits concavity from operator-concavity of f (Peierls–Bogoliubov / Klein). Mixing states cannot decrease uncertainty. C
6
S(ρ ‖ σ) ≡ Tr(ρ ln ρ) − Tr(ρ ln σ) ≥ 0 ,   equality iff ρ = σ
Klein’s inequality (non-negativity of quantum relative entropy), proved from the strict convexity of x ln x applied to the eigenvalues of ρ and σ with the overlap matrix |⟨i|j⟩|2 as a doubly stochastic weight. C
7
choose ρ = ρAB ,   σ = ρA ⊗ ρB ,   ρA = TrB ρAB ,   ρB = TrA ρAB
Insert the reduced density matrices (partial traces) into Klein’s inequality. Both are legitimate density operators, so the inequality applies. B
8
ln(ρA ⊗ ρB) = ln ρA ⊗ 1B + 1A ⊗ ln ρB
Logarithm of a tensor product of commuting factors splits into a sum; ρA⊗1 and 1⊗ρB commute so ln is additive. B
9
Tr(ρAB ln(ρA⊗ρB)) = Tr(ρA ln ρA) + Tr(ρB ln ρB) = −S(ρA) − S(ρB)
Substitute step 8; the marginal TrB ρAB = ρA collapses each cross term to a single-subsystem trace. This is the defining property of the partial trace. B
10
0 ≤ S(ρAB ‖ ρA⊗ρB) = −S(ρAB) + S(ρA) + S(ρB)
Assemble steps 6, 7 and 9: Tr(ρAB ln ρAB) = −S(ρAB) and the mixed term is −S(ρA)−S(ρB). Rearranging gives subadditivity. B
11
ρAB ↦ |ψ⟩ABR ,   ρAB = TrR |ψ⟩⟨ψ| ,   S(ρAB) = S(ρR)
Purify ρAB on a reference R. A pure global state has equal-spectrum marginals across any bipartition, so S(AB) = S(R) and S(B) = S(AR). C
12
S(ρAR) ≤ S(ρA) + S(ρR)  ⇒  S(ρB) ≤ S(ρA) + S(ρAB)
Apply subadditivity (step 10) to the bipartition A : R and substitute S(AR)=S(B), S(R)=S(AB) from step 11. This yields S(AB) ≥ S(B) − S(A); symmetry in A↔B gives the absolute value. C
Result
S(ρ) = −∑k pk ln pk   |   |S(ρA) − S(ρB)| ≤ S(ρAB) ≤ S(ρA) + S(ρB)

Reading. The quantum entropy is nothing but the classical Shannon entropy of the probabilities {pk} read off the density matrix in its own eigenbasis; measurement in any other basis can only increase it. For a joint system, the total ignorance is at most the sum of the ignorance about each part (subadditivity, with the deficit equal to the mutual information I(A:B)) and at least the difference of the two (Araki–Lieb). The lower bound is what forces a highly entangled pure whole to have low total entropy while its parts are individually very mixed.

Units check. Each pk is dimensionless and ln pk is dimensionless, so S is a pure number (measured in nats with the natural log, or bits if log2 is used; multiply by Boltzmann’s constant kB, units J K−1, to obtain the thermodynamic entropy). Every term in the inequalities carries the same units, so the bounds are dimensionally consistent.

Limiting cases
  • Pure state: ρ = |ψ⟩⟨ψ| has one eigenvalue 1 and the rest 0, so S = 0 — maximal knowledge, zero missing information.
  • Maximally mixed: ρ = 1/d gives pk = 1/d and S = ln d, the largest possible value in dimension d.
  • Diagonal (classical) state: if ρ is already diagonal in the reference basis, S is literally the Shannon entropy of a classical distribution — the quantum formula contains the classical one.
  • Product state: ρAB = ρA⊗ρB saturates subadditivity: S(AB) = S(A)+S(B), so I(A:B)=0 (no correlations).
  • Pure entangled bipartite state: S(AB)=0 forces S(A)=S(B) (the entanglement entropy) and saturates Araki–Lieb from below.
Breaks when
  • Non-positive or non-normalised “states”. If ρ has a negative eigenvalue (e.g. an unphysical output of an approximate reconstruction, or a partial transpose ρTB), ln pk is undefined and S is no longer real; the interpretation and every inequality collapse.
  • Infinite-dimensional systems with unbounded entropy. For a harmonic oscillator or a field mode with a heavy-tailed occupation, −∑ pk ln pk can diverge to +∞; differences like the mutual information may be ill-defined and the Araki–Lieb bound requires at least one finite marginal entropy.
  • Strong subadditivity is not what was proved here. The three-party inequality S(ABC)+S(B) ≤ S(AB)+S(BC) does not follow from ordinary subadditivity; assuming it does (Lieb–Ruskai is a separate, much deeper result) is a genuine failure of this derivation’s scope.
  • Rényi/Tsallis substitutes. Replacing −∑ p ln p by a Rényi entropy Sα with α ≠ 1 breaks subadditivity in general — the proof used the specific additivity ln(ρA⊗ρB) = ln ρA ⊕ ln ρB, which only the logarithm supplies.
Failure modes
  • Diagonalising in the wrong basis. Computing −∑i ⟨i|ρ|i⟩ ln⟨i|ρ|i⟩ in a fixed basis instead of the eigenbasis gives the (larger) Shannon entropy of the measured distribution, not S(ρ). The two agree only when ρ is already diagonal in that basis.
  • Forgetting 0 ln 0 = 0. Treating a zero eigenvalue as contributing −0·(−∞) and either dropping it wrongly or reporting NaN; the correct limit is a clean zero.
  • Assuming S(AB) ≥ S(A). This classical intuition (“the whole knows at least as much as a part”) is false quantum-mechanically: for a pure entangled state S(AB)=0 < S(A). Only the Araki–Lieb combination survives.
  • Using log2 and ln inconsistently. Mixing bits and nats within one calculation scales results by ln 2 ≈ 0.693 and corrupts every comparison.
  • Concavity vs. convexity sign slip. Writing S(λρ1+(1−λ)ρ2) ≤ λS1+(1−λ)S2; entropy is concave, so mixing raises entropy, not lowers it.
Discussion

The central conceptual point is that S(ρ) is a spectral invariant: it is a function of the eigenvalues of ρ alone, indifferent to the eigenvectors. This is why unitary evolution — which merely rotates the eigenbasis — leaves it fixed, and why the entropy of an isolated system governed by a Hamiltonian is a constant of the motion. Entropy increase in quantum thermodynamics is therefore never a property of unitary dynamics on the full state; it arises only when we discard information, either by tracing out an environment (the reduced state ρA generically has higher entropy than the pure global state) or by coarse-graining our description.

Subadditivity is the quantum expression of the fact that correlations store information non-locally. The mutual information I(A:B) = S(A)+S(B)−S(AB) = S(ρAB‖ρA⊗ρB) ≥ 0 measures the total — classical plus quantum — correlation between the parts, and vanishes exactly for product states. Because it is a relative entropy against the uncorrelated reference, it can never be negative, which is precisely the content of subadditivity. The Araki–Lieb bound is its mirror image obtained by purification: it encodes that a globally pure but internally entangled system can look maximally disordered locally while carrying zero total entropy — the hallmark of entanglement that has no classical analogue.

These inequalities are the scaffolding of quantum information theory. The Holevo bound, the quantum data-processing inequality, the monogamy of entanglement, and Landauer’s principle for the thermodynamic cost of erasure all descend from entropy’s concavity and (strong) subadditivity. In many-body physics the same quantities produce the celebrated area law: for a gapped ground state the entanglement entropy of a region scales with the area of its boundary rather than its volume, a statement about how subadditivity constrains correlations in low-energy states.

The deepest of the family, strong subadditivity S(ABC)+S(B) ≤ S(AB)+S(BC), is equivalent to the joint convexity of quantum relative entropy and to the monotonicity of relative entropy under completely positive trace-preserving maps (the data-processing inequality). Lieb and Ruskai’s 1973 proof used Lieb’s concavity theorem — the operator concavity of (A,B) ↦ Tr(K At K B1−t) — a genuinely operator-algebraic input beyond the elementary spectral argument used for ordinary subadditivity above. Equality in strong subadditivity characterises quantum Markov chains, states reconstructible from their marginals, which underpins the modern recovery-map refinements of Fawzi–Renner.

Common misconceptions. Von Neumann entropy is not “the entropy of the wavefunction” — a pure state always has S=0 no matter how complicated |ψ⟩ looks; the entropy lives in the mixedness of ρ. It is also not the same as the measurement (Shannon) entropy in a chosen basis, which is always at least as large and equals S(ρ) only for a measurement in the eigenbasis. And subadditivity does not say the whole is more disordered than its parts — quantum-mechanically it can be far less.

Worked examples
1
Bell state, reduced entropy.   |Φ+⟩ = (|00⟩ + |11⟩)/√2
Take the maximally entangled two-qubit state; compute S(A), S(B), S(AB) and check both bounds. A
2
ρAB = |Φ+⟩⟨Φ+| ⇒ S(ρAB) = 0  (pure)
A pure state has spectrum {1,0,0,0}, so S=−1·ln1 = 0. A
3
ρA = TrB ρAB = ½(|0⟩⟨0| + |1⟩⟨1|) = ½ 12
Partial trace over B: the off-diagonal terms |0⟩⟨1|⊗⟨0|1⟩ vanish, leaving the maximally mixed qubit. Same for ρB. A
4
S(ρA) = S(ρB) = −2·(½ ln½) = ln 2 = 0.693 nats = 1 bit
Shannon entropy of {½,½}. A
5
check:  |SA−SB| = 0 ≤ SAB = 0 ≤ SA+SB = 2 ln 2
Both bounds hold; Araki–Lieb is saturated (lower), subadditivity is strict with gap I(A:B)=2 ln 2. A
S(A)=S(B)=1 bit,  S(AB)=0,  I(A:B)=2 bits

Reading. Each qubit alone is maximally random (1 bit of ignorance), yet the pair is perfectly known (0 bits). The 2 bits of mutual information is the entanglement made visible: it is twice the entanglement entropy because it counts correlations symmetrically. Units: pure numbers (bits).

1
Classical correlated bits vs. a general qubit.   ρAB = ½(|00⟩⟨00| + |11⟩⟨11|)
A separable, classically correlated mixed state; compute the entropies and mutual information, then separately entropy of a thermal-like single qubit. A
2
spectrum of ρAB = {½, ½, 0, 0} ⇒ S(ρAB) = ln 2 = 1 bit
The state is already diagonal; Shannon entropy of {½,½}. A
3
ρA = ½(|0⟩⟨0|+|1⟩⟨1|) ⇒ S(ρA) = S(ρB) = ln 2 = 1 bit
Partial trace gives the maximally mixed qubit; no coherences to lose since none were present. A
4
I(A:B) = SA+SB−SAB = 1+1−1 = 1 bit
Ordinary (classical) correlation; contrast with the Bell state’s 2 bits. Subadditivity holds with gap 1 bit; Araki–Lieb gives 0 ≤ 1. B
5
single qubit  ρ = diag(p, 1−p),  p = 0.9 ⇒ S = −0.9 ln0.9 − 0.1 ln0.1
Binary entropy at a biased population, e.g. a qubit in near-ground thermal occupation. Symbols first, then numbers. A
6
S = −0.9(−0.1054) − 0.1(−2.3026) = 0.0948 + 0.2303 = 0.325 nats = 0.469 bits
Evaluate; convert nats→bits by dividing by ln 2. A
S(AB)=1 bit,  I(A:B)=1 bit (classical);  H2(0.9)=0.469 bits

Reading. Classical correlation gives at most 1 bit of mutual information for two bits, exactly half the Bell-state value — a quantitative signature that the maximally entangled state is “more correlated” than any classical mixture. The biased single qubit shows entropy falling well below its 1-bit maximum as the population concentrates. Units: bits (dimensionless).

Problems
  1. (A) Depolarised qubit. A qubit is in ρ = (1−ε)|0⟩⟨0| + ε 1/2 with ε=0.2. Find its eigenvalues and S(ρ) in bits.
    Solution ρ = diag(1−ε+ε/2,  ε/2) = diag(1−ε/2,  ε/2) = diag(0.9, 0.1). Then S = −0.9 ln0.9 − 0.1 ln0.1 = 0.325 nats = 0.325/0.693 = 0.469 bits — the binary entropy H2(0.1).
  2. (A) Three-level maximally mixed. A qutrit is maximally mixed, ρ = 1/3. Give S in nats and in bits, and state why no qutrit state exceeds it.
    Solution pk=1/3 for k=1,2,3, so S = −3·(1/3)ln(1/3) = ln 3 = 1.099 nats = 1.585 bits. It is the maximum because for fixed dimension d the Shannon entropy of a distribution is maximised by the uniform distribution, giving ln d; any other spectrum is more peaked and has lower entropy (by concavity / Jensen).
  3. (B) Verify subadditivity for a Werner-like diagonal state. Let ρAB = diag(0.4, 0.1, 0.1, 0.4) in the basis {|00⟩,|01⟩,|10⟩,|11⟩}. Compute S(AB), the marginals, S(A), S(B), and confirm both bounds.
    Solution S(AB) = −2(0.4 ln0.4) − 2(0.1 ln0.1) = −2(−0.3665) − 2(−0.2303) = 0.7330+0.4605 = 1.194 nats = 1.721 bits. Marginals: pA(0)=0.4+0.1=0.5, pA(1)=0.1+0.4=0.5, so ρA=diag(0.5,0.5) and S(A)=ln2=1 bit; identically S(B)=1 bit. Bounds: |SA−SB|=0 ≤ 1.721 ≤ SA+SB=2 bits. Both hold; mutual information I(A:B)=2−1.721=0.279 bits of (classical) correlation.
  4. (B) Mutual information of a partially entangled pure state. |ψ⟩ = cosθ|00⟩ + sinθ|11⟩ with θ=π/6. Find S(A), S(AB), and I(A:B).
    Solution Schmidt coefficients cos2θ=0.75, sin2θ=0.25. Pure state: S(AB)=0. ρA=diag(0.75,0.25), so S(A)=−0.75 ln0.75 − 0.25 ln0.25 = 0.2158+0.3466 = 0.562 nats = 0.811 bits = H2(0.25); equal to S(B). Mutual information I(A:B)=S(A)+S(B)−S(AB)=2×0.811=1.623 bits. Araki–Lieb |SA−SB|=0=S(AB) is saturated.
  5. (C) Araki–Lieb sharpness via purification. Show that for any bipartite ρAB one has S(AB) ≥ S(A)−S(B), and give a state that makes it an equality.
    Solution Purify: introduce R with |ψ⟩ABR pure and ρAB=TrR|ψ⟩⟨ψ|. For a global pure state, marginals across a cut have equal spectra, so S(B)=S(AR) and S(AB)=S(R). Apply subadditivity to the cut A:R: S(AR) ≤ S(A)+S(R), i.e. S(B) ≤ S(A)+S(AB), giving S(AB) ≥ S(B)−S(A); swapping A↔B gives S(AB) ≥ |S(A)−S(B)|. Equality example: take ρABA⊗|0⟩⟨0|B with ρA pure — then S(AB)=S(A)=S(B)=0 trivially; more informatively, any pure entangled state gives S(AB)=0=|S(A)−S(B)| since S(A)=S(B), saturating the bound.