Simultaneous Diagonalization of Commuting Operators
Statement
Let \(\hat{A}\) and \(\hat{B}\) be self-adjoint operators on a finite-dimensional complex inner-product space \(V\), \(\dim V = n\), and suppose they commute, \([\hat{A},\hat{B}] = \hat{A}\hat{B} - \hat{B}\hat{A} = 0\). Then \(V\) possesses a single orthonormal basis \(\{\lvert e_1\rangle,\dots,\lvert e_n\rangle\}\) whose every vector is simultaneously an eigenvector of \(\hat{A}\) and of \(\hat{B}\), so that one unitary \(U\) brings both to diagonal form: \(U^{\dagger}\hat{A}U = \operatorname{diag}(a_1,\dots,a_n)\) and \(U^{\dagger}\hat{B}U = \operatorname{diag}(b_1,\dots,b_n)\). The converse also holds: two self-adjoint operators sharing an orthonormal eigenbasis necessarily commute.
Why it matters
This is the mathematical definition of compatible observables. In quantum mechanics a measurement projects a state onto an eigenvector; if two observables share an eigenbasis, a single state can have a definite value of both at once, and the two measurements can be performed in either order without disturbing each other. The theorem tells us precisely when that is possible: exactly when the operators commute.
It underwrites the entire bookkeeping of good quantum numbers and complete sets of commuting observables (CSCO). The reason the hydrogen atom's stationary states are labelled by \((n,\ell,m)\) is that \(\hat{H}\), \(\hat{L}^2\) and \(\hat{L}_z\) mutually commute and can be simultaneously diagonalized, so each eigenvector carries one label from each operator. Where operators fail to commute, no such joint labelling exists, and an uncertainty relation appears in its place.
Assumptions
Derivation
Result
Reading. Commuting self-adjoint operators are two lists of real numbers attached to one shared set of orthogonal directions: the direction \(\lvert e_k\rangle\) is stretched by \(a_k\) under \(\hat{A}\) and by \(b_k\) under \(\hat{B}\). Each common eigenvector carries a joint eigenvalue label \((a_k,b_k)\) — the mathematical origin of a state's simultaneous quantum numbers. When one operator's spectrum is degenerate, the other operator's eigenvalue distinguishes the vectors inside that degenerate block.
Units check. \(U\) is built from normalized, dimensionless states and is itself dimensionless, so \(U^{\dagger}\hat{A}U\) carries the units of \(\hat{A}\) and \(U^{\dagger}\hat{B}U\) the units of \(\hat{B}\). Each eigenvalue \(a_k\) inherits the physical dimension of the observable \(\hat{A}\) (e.g. energy) and each \(b_k\) that of \(\hat{B}\) (e.g. \(\hbar^2\) for \(\hat{L}^2\)); the two lists need not share units. The commutator \([\hat{A},\hat{B}]\) carries units \([\hat{A}][\hat{B}]\), and its vanishing is dimensionally consistent on both sides.
Limiting cases
- \(\hat{A}\) non-degenerate (all \(a_k\) distinct): every \(E_a\) is one-dimensional, so step 5 is trivial and each \(\hat{A}\)-eigenvector is automatically a \(\hat{B}\)-eigenvector. \(\hat{A}\) alone forms a CSCO and the common basis is unique up to phases.
- \(\hat{A}=c\hat{I}\) (a multiple of the identity): every vector lies in one eigenspace, \(\hat{A}\) imposes no constraint, and the common basis is whatever diagonalizes \(\hat{B}\) alone.
- \(\hat{B}=f(\hat{A})\) (a function of \(\hat{A}\)): the two automatically commute and share \(\hat{A}\)'s eigenbasis, with \(b_k=f(a_k)\). This is the tightest possible dependence.
- Both already diagonal in the working basis: \(U=I\), and the joint eigenvalues are read straight off the two diagonals.
Breaks when
- The operators do not commute, \([\hat{A},\hat{B}]\neq 0\). Step 2 collapses: \(\hat{B}\) maps some \(\hat{A}\)-eigenvector out of its eigenspace, so no complete common eigenbasis exists. For \(\hat{\sigma}_x\) and \(\hat{\sigma}_z\) (which anticommute) there is not a single shared eigenvector, and the position–momentum pair \([\hat{x},\hat{p}]=i\hbar\) has none at all — the Robertson uncertainty relation replaces simultaneous diagonalization.
- Non-self-adjoint (non-normal) operators. Commuting matrices always share at least one eigenvector over \(\mathbb{C}\), but if either is defective there is no eigenbasis to share. Two commuting nilpotent Jordan blocks are simultaneously triangularizable but not diagonalizable, and any shared eigenvectors are not orthonormal.
- Infinite dimensions with continuous spectrum. Commuting operators such as the free-particle \(\hat{H}=\hat{p}^2/2m\) and \(\hat{p}\) have no normalizable eigenvectors; the plane waves \(e^{ikx}\notin L^2(\mathbb{R})\). "Simultaneous diagonalization" must be reinterpreted as a joint projection-valued measure over the continuous joint spectrum.
Failure modes
- Believing degeneracy needs no extra work. When \(\hat{A}\) has a repeated eigenvalue, an arbitrary orthonormal basis of \(E_a\) is generally not a \(\hat{B}\)-eigenbasis; you must still diagonalize \(\hat{B}\) inside each degenerate block (step 5). Skipping this leaves \(\hat{B}\) block-diagonal but not diagonal.
- Confusing "share one eigenvector" with "simultaneously diagonalizable". Non-commuting operators can accidentally share a single eigenvector; that does not give a full common basis. The theorem asserts a complete shared basis, which requires commutation.
- Thinking \(\hat{B}\) may act freely on a degenerate \(\hat{A}\)-eigenspace. Commutation forces \(\hat{B}\) to be block-diagonal — it must map \(E_a\) into itself. Writing down a \(\hat{B}\) with cross-block entries silently violates \([\hat A,\hat B]=0\).
- Assuming commuting operators are equal or proportional. Commutation is far weaker than equality; \(\hat{L}^2\) and \(\hat{L}_z\) commute yet are entirely different operators with different spectra.
- Using the real dot product on complex eigenvectors. Normalizing and checking orthogonality with \(u^{T}v\) instead of the Hermitian \(u^{\dagger}v\) breaks \(U^\dagger U = I\), exactly as in the single-operator spectral theorem.
Discussion
The engine of the proof is a two-stage decoupling. The first spectral theorem breaks \(V\) into \(\hat{A}\)-eigenspaces; commutation guarantees \(\hat{B}\) respects that block structure; the second spectral theorem then diagonalizes \(\hat{B}\) inside each block. Physically, \(\hat{A}\) sorts states into bins by its own eigenvalue, and \(\hat{B}\) — because it cannot move a state between bins — can only rearrange states within a bin, where it too can be diagonalized. The result is a grid of joint eigenvalues \((a,b)\) labelling an orthonormal basis.
This is why quantum states carry several quantum numbers at once. A maximal set of mutually commuting observables — a CSCO — pins down each basis vector with a unique tuple of eigenvalues, with no remaining degeneracy. For the hydrogen atom \(\{\hat{H},\hat{L}^2,\hat{L}_z\}\) is such a set; each successive operator lifts the degeneracy left by the previous ones, exactly as \(\hat{B}\) splits a degenerate \(\hat{A}\)-block in the proof. When a symmetry produces degeneracy, one looks for an operator that commutes with the Hamiltonian to label the degenerate multiplet — the direct link between this theorem and the symmetry thread.
The converse (step 10) elevates the result to an equivalence: for self-adjoint (more generally normal) operators, commuting and being simultaneously diagonalizable are the same statement. This is the finite-dimensional heart of a deeper structural fact — a family of mutually commuting self-adjoint operators generates a commutative algebra whose common eigenbasis realizes its joint spectrum, the finite-dimensional shadow of the spectral theorem for commuting operators and of the Gelfand picture in which a commutative operator algebra is functions on a joint spectrum. It is also the precise statement dual to the uncertainty principle: the failure of two observables to be jointly diagonalizable is quantified by the size of their commutator through the Robertson relation.
Common misconceptions. Simultaneous diagonalizability does not require that each operator's eigenvectors are individually forced — only that some common orthonormal basis exists; when both are degenerate, that basis is highly non-unique. And commuting is not a loose statement that "measuring one never disturbs the other" — it is the sharp condition that a joint eigenbasis exists, which is what makes the order of the two measurements immaterial on those joint eigenstates.
Worked examples
Example 1 — a non-degenerate \(\hat{A}\) fixes the basis for \(\hat{B}\).
Reading. The joint eigenvalue labels are \((a,b)=(+1,3)\) and \((-1,1)\); one rotation into the \((1,\pm1)\) directions diagonalizes both operators at once.
Example 2 — a degenerate \(\hat{A}\) whose degeneracy is lifted by \(\hat{B}\).
Reading. The joint labels \((1,1),(1,-1),(2,5)\) are all distinct even though \(\hat A\) alone repeats the value \(1\): \(\hat B\) supplies the second quantum number that resolves the degeneracy, so \(\{\hat A,\hat B\}\) is a CSCO.
Problems
- Prove the converse directly: if self-adjoint \(\hat{A},\hat{B}\) share an orthonormal eigenbasis \(\{\lvert k\rangle\}\), then \([\hat{A},\hat{B}]=0\).
Solution
Let \(\hat{A}\lvert k\rangle=a_k\lvert k\rangle\) and \(\hat{B}\lvert k\rangle=b_k\lvert k\rangle\) for every basis vector. Then \(\hat{A}\hat{B}\lvert k\rangle=\hat{A}(b_k\lvert k\rangle)=a_k b_k\lvert k\rangle\), and \(\hat{B}\hat{A}\lvert k\rangle=b_k a_k\lvert k\rangle\). The operators \(\hat{A}\hat{B}\) and \(\hat{B}\hat{A}\) agree on every vector of a basis, hence are equal as linear maps, so \([\hat{A},\hat{B}]=\hat{A}\hat{B}-\hat{B}\hat{A}=0\). - The matrices \(\hat{A}=\left(\begin{smallmatrix}3&1\\1&3\end{smallmatrix}\right)\) and \(\hat{B}=\left(\begin{smallmatrix}1&2\\2&1\end{smallmatrix}\right)\) commute. Find their common orthonormal eigenbasis and both diagonal forms, and list the joint eigenvalues.
Solution
Check: \(\hat{A}\hat{B}=\left(\begin{smallmatrix}5&7\\7&5\end{smallmatrix}\right)=\hat{B}\hat{A}\), so they commute. Both are of the form \(cI+d\,\hat\sigma_x\), so both are diagonalized by the \(\hat\sigma_x\) eigenvectors \(\lvert e_1\rangle=\tfrac1{\sqrt2}(1,1)^T\), \(\lvert e_2\rangle=\tfrac1{\sqrt2}(1,-1)^T\). Then \(\hat{A}\): \(a=3\pm1=\{4,2\}\); \(\hat{B}\): \(b=1\pm2=\{3,-1\}\). So \(U=\tfrac1{\sqrt2}\left(\begin{smallmatrix}1&1\\1&-1\end{smallmatrix}\right)\), \(U^{\dagger}\hat{A}U=\operatorname{diag}(4,2)\), \(U^{\dagger}\hat{B}U=\operatorname{diag}(3,-1)\). Joint eigenvalues \((4,3)\) and \((2,-1)\). - For \(\hat{A}=\operatorname{diag}(2,2,5)\) and \(\hat{B}=\left(\begin{smallmatrix}3&1&0\\1&3&0\\0&0&7\end{smallmatrix}\right)\), verify they commute and construct a common orthonormal eigenbasis, showing that \(\hat{B}\) lifts the degeneracy of \(\hat{A}\).
Solution
\(\hat{A}\) is scalar (\(=2I\)) on the top-left \(2\times2\) block and \(5\) on the last slot, both scalars, so it commutes with the block-diagonal \(\hat{B}\); explicitly \(\hat{A}\hat{B}=\hat{B}\hat{A}\) block by block. The degenerate space is \(E_2=\operatorname{span}\{\hat x,\hat y\}\) (\(a=2\)). Restrict \(\hat{B}\) there: \(\left(\begin{smallmatrix}3&1\\1&3\end{smallmatrix}\right)\) has eigenvalues \(4,2\) with eigenvectors \(\tfrac1{\sqrt2}(1,1,0)^T\) and \(\tfrac1{\sqrt2}(1,-1,0)^T\). The slot \(a=5\) gives \((0,0,1)^T\) with \(b=7\). Common basis \(\{\tfrac1{\sqrt2}(1,1,0),\tfrac1{\sqrt2}(1,-1,0),(0,0,1)\}\); joint eigenvalues \((2,4),(2,2),(5,7)\), all distinct — so \(\{\hat A,\hat B\}\) is a CSCO. - Show that if \([\hat{A},\hat{B}]=0\) and \(\hat{A}\) has a non-degenerate spectrum, then every eigenvector of \(\hat{A}\) is automatically an eigenvector of \(\hat{B}\).
Solution
Let \(\hat{A}\lvert v\rangle=a\lvert v\rangle\). By commutation, \(\hat{A}(\hat{B}\lvert v\rangle)=\hat{B}\hat{A}\lvert v\rangle=a(\hat{B}\lvert v\rangle)\), so \(\hat{B}\lvert v\rangle\in E_a\). Non-degeneracy means \(\dim E_a=1\), i.e. \(E_a=\operatorname{span}\{\lvert v\rangle\}\). Any vector of \(E_a\) is a scalar multiple of \(\lvert v\rangle\), so \(\hat{B}\lvert v\rangle=b\lvert v\rangle\) for some scalar \(b\): \(\lvert v\rangle\) is a \(\hat{B}\)-eigenvector. This is why a single non-degenerate operator already forms a CSCO. - Demonstrate the failure of the theorem when the operators do not commute: show that the Pauli matrices \(\hat{\sigma}_x=\left(\begin{smallmatrix}0&1\\1&0\end{smallmatrix}\right)\) and \(\hat{\sigma}_z=\left(\begin{smallmatrix}1&0\\0&-1\end{smallmatrix}\right)\) share no common eigenvector, and connect this to their commutator.
Solution
The eigenvectors of \(\hat{\sigma}_z\) are \((1,0)^T\) and \((0,1)^T\) (eigenvalues \(+1,-1\)). The eigenvectors of \(\hat{\sigma}_x\) are \(\tfrac1{\sqrt2}(1,1)^T\) and \(\tfrac1{\sqrt2}(1,-1)^T\) (eigenvalues \(+1,-1\)). No vector appears in both lists (up to scalar), so there is no common eigenvector, let alone a common basis. Consistently, \([\hat{\sigma}_x,\hat{\sigma}_z]=\hat{\sigma}_x\hat{\sigma}_z-\hat{\sigma}_z\hat{\sigma}_x=\left(\begin{smallmatrix}-1&0\\0&1\end{smallmatrix}\right)-\left(\begin{smallmatrix}1&0\\0&-1\end{smallmatrix}\right)=\left(\begin{smallmatrix}-2&0\\0&2\end{smallmatrix}\right)=-2i\,\hat{\sigma}_y\neq0\). A non-vanishing commutator forbids simultaneous diagonalization and forces the spin-\(x\)/spin-\(z\) uncertainty relation.