physics2u
Tier
⌕ Search ⌘K
Derivation

Variational Characterization of Eigenvalues

D-086 Home PU-104 Threads symmetry · energy Depends on Spectral Theorem for Self-Adjoint Operators
Statement

Let \(\hat{A}\) be a self-adjoint operator on a Hilbert space \(\mathcal{H}\) with a discrete, bounded-below spectrum and eigenvalues ordered \(\lambda_1 \le \lambda_2 \le \lambda_3 \le \cdots\). Define the Rayleigh quotient \(R(\psi) = \dfrac{\langle \psi, \hat{A}\psi\rangle}{\langle \psi,\psi\rangle}\) for \(\psi \neq 0\). Then the stationary points of \(R\) are exactly the eigenvectors of \(\hat{A}\), the stationary values are the eigenvalues, and the ordered eigenvalues obey the Courant–Fischer min–max principle \(\lambda_n = \min_{\dim S = n}\ \max_{\psi \in S \setminus \{0\}} R(\psi)\).

Why it matters

The variational principle turns a spectral (linear-algebra) problem into a minimisation problem, which is the entire foundation of the Rayleigh–Ritz method, the ground-state energy bounds of quantum mechanics, and finite-element eigenvalue solvers. Any normalised trial state gives a rigorous upper bound on the ground-state energy, so you can estimate \(\lambda_1\) without ever diagonalising \(\hat{A}\).

The min–max form is the workhorse for comparison theorems: it lets you bound how eigenvalues move when the operator or the domain changes (interlacing, monotonicity under perturbation, Weyl's law), because it characterises \(\lambda_n\) purely through subspaces rather than through the eigenvectors themselves.

Assumptions
\(\hat{A}\) is self-adjoint, \(\hat{A}=\hat{A}^\dagger\) on a dense domain.If merely symmetric but not self-adjoint the spectral theorem fails; \(R\) may be real yet its stationary values need not exhaust the spectrum, and residual or continuous spectrum can appear.
The spectrum is discrete and bounded below, with a complete orthonormal eigenbasis \(\{\varphi_n\}\).If dropped, the minimum need not be attained (only an infimum); a continuous spectrum yields stationary "generalised" eigenvectors that lie outside \(\mathcal{H}\).
The relevant trial vectors lie in the form domain \(Q(\hat{A})\), where \(\langle\psi,\hat{A}\psi\rangle\) is finite.If a trial state has divergent \(\langle\psi,\hat{A}\psi\rangle\) (for example a cusp giving infinite kinetic energy), the quotient is ill-defined and the bound is vacuous.
Derivation
1
\[ R(\psi) = \frac{\langle \psi, \hat{A}\psi\rangle}{\langle \psi,\psi\rangle}, \qquad \psi \in \mathcal{H}\setminus\{0\}. \]
Definition of the Rayleigh quotient; \(R\) is invariant under \(\psi \mapsto c\psi\), so we may work on the unit sphere \(\|\psi\|=1\). A
2
\[ \psi = \sum_n c_n \varphi_n, \qquad \hat{A}\varphi_n = \lambda_n \varphi_n, \qquad \langle \varphi_m,\varphi_n\rangle = \delta_{mn}. \]
The spectral theorem for self-adjoint operators supplies a complete orthonormal eigenbasis; every \(\psi\) expands in it. A
3
\[ \langle \psi,\hat{A}\psi\rangle = \sum_n \lambda_n |c_n|^2, \qquad \langle \psi,\psi\rangle = \sum_n |c_n|^2. \]
Apply \(\hat{A}\) termwise and use orthonormality; cross terms \(\langle\varphi_m,\varphi_n\rangle\) vanish for \(m\neq n\). Reality of \(\lambda_n\) follows from self-adjointness. A
4
\[ R(\psi) = \frac{\sum_n \lambda_n |c_n|^2}{\sum_n |c_n|^2} = \sum_n \lambda_n\, w_n, \qquad w_n = \frac{|c_n|^2}{\sum_k |c_k|^2}\ge 0,\ \ \sum_n w_n = 1. \]
Divide; the quotient is a convex (weighted) average of the eigenvalues with weights \(w_n\). A
5
\[ \lambda_1 = \sum_n \lambda_1 w_n \le \sum_n \lambda_n w_n = R(\psi) \le \sum_n \lambda_{\max} w_n = \lambda_{\max}. \]
A weighted average lies between the smallest and largest values averaged; equality on the left needs all weight on the \(\lambda_1\)-eigenspace. B
6
\[ \min_{\psi\neq 0} R(\psi) = \lambda_1, \quad\text{attained iff } \psi \in \ker(\hat{A}-\lambda_1). \]
The lower bound is achieved by choosing \(\psi=\varphi_1\) (so \(w_1=1\)); hence the global minimum of \(R\) is the lowest eigenvalue. B
7
\[ \delta\!\left[\langle\psi,\hat{A}\psi\rangle - \mu\big(\langle\psi,\psi\rangle-1\big)\right] = 0 \ \Longrightarrow\ \langle\delta\psi,\ \hat{A}\psi - \mu\psi\rangle + \text{c.c.} = 0. \]
Impose stationarity of \(\langle\psi,\hat{A}\psi\rangle\) at fixed norm using a Lagrange multiplier \(\mu\); vary \(\psi\) and its conjugate independently. Self-adjointness lets both variations act through \(\hat{A}\). C
8
\[ \hat{A}\psi = \mu\,\psi, \qquad \mu = \frac{\langle\psi,\hat{A}\psi\rangle}{\langle\psi,\psi\rangle} = R(\psi). \]
Because \(\delta\psi\) is arbitrary, the bracket vanishes: every stationary point is an eigenvector, and its stationary value is the eigenvalue \(\mu=R(\psi)\). C
9
\[ \lambda_n = \min_{\substack{\psi\perp \varphi_1,\dots,\varphi_{n-1}\\ \psi\neq 0}} R(\psi), \qquad \lambda_n = \min_{\dim S = n}\ \max_{\psi\in S\setminus\{0\}} R(\psi). \]
Repeating Step 6 on \(\{\varphi_1,\dots,\varphi_{n-1}\}^\perp\) gives \(\lambda_n\). Removing the dependence on the (unknown) eigenvectors yields the basis-free Courant–Fischer min–max, proved by comparing any \(n\)-dimensional \(S\) with \(\mathrm{span}\{\varphi_1,\dots,\varphi_n\}\) via a dimension count. C
Result
\[ \lambda_1 = \min_{\psi\neq 0} R(\psi), \qquad \lambda_n = \min_{\dim S = n}\ \max_{\psi\in S\setminus\{0\}} \frac{\langle\psi,\hat{A}\psi\rangle}{\langle\psi,\psi\rangle} = \max_{\dim T = n-1}\ \min_{\psi\perp T} R(\psi). \]

Reading. The lowest eigenvalue is the global minimum of the Rayleigh quotient, and the \(n\)-th eigenvalue is the smallest value the quotient can be forced to reach on some \(n\)-dimensional subspace. Every eigenvector is a stationary point of \(R\); every stationary value is an eigenvalue. Consequently any normalised trial state gives \(R(\psi)\ge \lambda_1\): a one-sided, always-valid upper bound on the ground state.

Units check. \(R\) inherits the units of \(\hat{A}\): the ratio \(\langle\psi,\hat{A}\psi\rangle/\langle\psi,\psi\rangle\) cancels the units of \(|\psi|^2\), leaving \([\hat{A}]\). For \(\hat{A}=\hat{H}\) (energy, joules), \(R\) is in joules, matching \([\lambda_n]=\mathrm{J}\). The weights \(w_n\) are dimensionless with \(\sum_n w_n = 1\), consistent with \(R\) being an average of eigenvalues.

Limiting cases
  • Trial state equal to an eigenvector \(\varphi_k\): \(R=\lambda_k\) exactly, and \(R\) is stationary there.
  • Trial state close to \(\varphi_1\), \(\psi=\varphi_1+\epsilon\,\varphi_2\): \(R=\lambda_1+\epsilon^2(\lambda_2-\lambda_1)+O(\epsilon^4)\); the error is second order in the state error (the variational payoff).
  • Degenerate ground state (\(\lambda_1=\lambda_2\)): the minimiser is any vector in the degenerate eigenspace; the minimum value is unchanged.
  • Finite-dimensional \(N\times N\) Hermitian matrix: \(\lambda_{\max}=\max R\), recovering the Courant–Fischer theorem of matrix analysis.
  • Two-sided extremes: for \(\hat{A}\) bounded above and below, \(\lambda_{\max}=\max_\psi R(\psi)\) by applying the argument to \(-\hat{A}\).
Breaks when
  • Non-self-adjoint operators. For a merely symmetric operator with nontrivial deficiency indices, or a non-normal operator, eigenvalues can be complex and \(R\) is not real; stationary points of \(\mathrm{Re}\,R\) need not be eigenvectors, so the principle collapses.
  • Spectrum not bounded below. For the Dirac operator or a free particle, \(\inf R=-\infty\); no minimiser exists and "\(\lambda_1\)" is undefined, so the variational bound is empty. One must restrict to a spectral gap or use projection.
  • Continuous spectrum below the discrete levels. If \(\lambda_n\) sits above the bottom of the essential spectrum \(\Sigma_{\mathrm{ess}}\), the min–max returns \(\inf\Sigma_{\mathrm{ess}}\) rather than a discrete eigenvalue; the bound is valid but does not isolate a bound state.
  • Trial states outside the form domain. If \(\langle\psi,\hat{A}\psi\rangle\) diverges (infinite kinetic energy from a discontinuous or cusped \(\psi\)), \(R\) is not defined and the estimate is meaningless.
Failure modes
  • Forgetting to normalise (or dividing by the wrong norm). Computing \(\langle\psi,\hat{H}\psi\rangle\) but omitting the denominator \(\langle\psi,\psi\rangle\) gives a number with the wrong scale and no bound property.
  • Believing the bound is two-sided. \(R(\psi)\ge\lambda_1\) is an upper bound on the ground state, never a lower bound; a "good-looking" trial energy could still be far above \(\lambda_1\).
  • Using a symmetry-mismatched trial state and calling it the ground state. A node-forced (for example antisymmetric) trial function bounds the lowest state of that symmetry sector, which may be an excited state overall.
  • Treating stationarity in a restricted family as exactness. Stationary \(R\) within a limited trial family is not a true eigenvalue unless the family contains the eigenvector.
  • Confusing the min–max subspace with the eigenspace. The optimal \(S\) in Courant–Fischer is \(\mathrm{span}\{\varphi_1,\dots,\varphi_n\}\); students often maximise over the wrong (for example \((n{-}1)\)-dimensional) space.
Discussion

The physical content is that eigenvalues are the "flat spots" of the expectation value under normalisation. In quantum mechanics \(R(\psi)=\langle\hat{H}\rangle_\psi\) is the mean energy, and the statement \(\delta\langle\hat H\rangle=0 \Rightarrow \hat H\psi=E\psi\) is exactly the time-independent Schrödinger equation reappearing as a stationarity condition. This is why the Rayleigh–Ritz method works: restrict \(\psi\) to a finite trial subspace, minimise \(R\), and the resulting values are guaranteed upper bounds to the true levels, converging from above as the subspace grows.

The min–max form connects to the symmetry thread through symmetry-adapted subspaces. If \(\hat{A}\) commutes with a symmetry group, \(\mathcal{H}\) splits into irreducible sectors; the variational principle applied within each sector gives the lowest level of each symmetry type. This is the rigorous basis for statements like "the ground state of a symmetric potential is nodeless and symmetric" (the Perron–Frobenius property for Schrödinger operators).

Through the energy thread, the principle underlies stability: a system relaxes toward the state minimising \(\langle\hat{H}\rangle\), so \(\lambda_1\) is the ground-state energy and \(\varphi_1\) the equilibrium configuration. The second-order error \(R-\lambda_1=O(\epsilon^2)\) explains why crude wavefunctions still give surprisingly accurate energies — the celebrated robustness of variational calculations in chemistry and condensed matter.

At the level of functional analysis, Courant–Fischer is the bridge to comparison theorems. Because \(\lambda_n\) is a minimum over subspaces of a maximum, monotonicity is immediate: enlarging \(\hat{A}\) (that is \(\hat{A}\le\hat{B}\) as quadratic forms) raises every \(\lambda_n\), and shrinking the domain (Dirichlet bracketing) raises them while enlarging it lowers them. Combined with a counting estimate this yields Weyl's asymptotic law \(N(\lambda)\sim C_d\,\mathrm{vol}(\Omega)\,\lambda^{d/2}\) and Cauchy interlacing for bordered matrices — all consequences of the single variational characterisation. Common misconceptions: that a lower trial energy always means a "better" wavefunction pointwise (it means better in the energy norm only); that stationarity implies a minimum (saddle points of \(R\) are excited eigenvectors); and that the principle gives lower bounds (it does not, without extra input such as Temple's inequality).

Worked examples
1
\[ \hat{A} = \begin{pmatrix} 2 & 1 \\ 1 & 2 \end{pmatrix}, \qquad \psi(\theta) = \begin{pmatrix}\cos\theta\\ \sin\theta\end{pmatrix}. \]
A \(2\times2\) real symmetric (Hermitian) matrix; parametrise the unit sphere by one angle \(\theta\). A
2
\[ R(\theta) = \psi^\dagger \hat{A}\psi = 2\cos^2\theta + 2\sin^2\theta + 2\cos\theta\sin\theta = 2 + \sin 2\theta. \]
Since \(\|\psi\|=1\) the denominator is 1; expand the quadratic form and use \(2\cos\theta\sin\theta=\sin2\theta\). A
3
\[ \frac{dR}{d\theta} = 2\cos 2\theta = 0 \ \Rightarrow\ \theta = -\tfrac{\pi}{4},\ \tfrac{\pi}{4}. \]
Stationary points of \(R\); by the theorem these should coincide with the eigenvectors. B
4
\[ R\!\left(-\tfrac{\pi}{4}\right) = 2 + \sin\!\left(-\tfrac{\pi}{2}\right) = 1, \qquad R\!\left(\tfrac{\pi}{4}\right) = 2 + \sin\!\left(\tfrac{\pi}{2}\right) = 3. \]
Evaluate; \(\theta=-\pi/4\) gives the eigenvector \(\tfrac{1}{\sqrt2}(1,-1)^{\top}\), and \(\theta=\pi/4\) gives \(\tfrac{1}{\sqrt2}(1,1)^{\top}\). B
\[ \lambda_1 = \min_\theta R = 1, \qquad \lambda_2 = \max_\theta R = 3. \]

Reading. The extrema of the Rayleigh quotient reproduce the eigenvalues \(\{1,3\}\) and their eigenvectors, exactly as the theorem predicts. A generic trial angle, say \(\theta=0\) giving \(\psi=(1,0)^\top\), yields \(R=2\ge\lambda_1=1\): a valid upper bound. Units. The matrix is dimensionless, so \(R\) and the eigenvalues are dimensionless numbers.

1
\[ \hat{H} = -\frac{\hbar^2}{2m}\frac{d^2}{dx^2}\ \text{on }[0,L],\quad \psi(x)=x(L-x),\quad \psi(0)=\psi(L)=0. \]
Infinite square well; a simple polynomial trial state satisfying the boundary conditions (no free parameter, a one-shot Rayleigh estimate). A
2
\[ \langle\psi,\psi\rangle = \int_0^L x^2(L-x)^2\,dx = \frac{L^5}{30}, \qquad \psi'(x) = L - 2x. \]
Normalisation integral and the derivative needed for the kinetic energy. A
3
\[ \langle\psi,\hat H\psi\rangle = \frac{\hbar^2}{2m}\int_0^L |\psi'(x)|^2\,dx = \frac{\hbar^2}{2m}\int_0^L (L-2x)^2\,dx = \frac{\hbar^2}{2m}\cdot\frac{L^3}{3}. \]
Integrate by parts once (the boundary terms vanish since \(\psi\) vanishes at the ends) to write the kinetic energy as \(\tfrac{\hbar^2}{2m}\!\int|\psi'|^2\). B
4
\[ R = \frac{\langle\psi,\hat H\psi\rangle}{\langle\psi,\psi\rangle} = \frac{\hbar^2}{2m}\cdot\frac{L^3/3}{L^5/30} = \frac{\hbar^2}{2m}\cdot\frac{10}{L^2} = \frac{5\hbar^2}{mL^2}. \]
Form the quotient; the \(L\) powers combine to \(1/L^2\), the correct scaling for a confined particle. B
\[ E_1 \le R = \frac{5\hbar^2}{mL^2} = \frac{10\,\hbar^2}{2mL^2}, \qquad E_1^{\mathrm{exact}} = \frac{\pi^2\hbar^2}{2mL^2} = \frac{9.8696\,\hbar^2}{2mL^2}. \]

Reading. The trial estimate \(10\) (in units of \(\hbar^2/2mL^2\)) sits just above the exact \(\pi^2\approx 9.87\), an error of only \(1.3\%\), and lies above it as guaranteed. Units. \([\hbar^2/mL^2]=(\mathrm{J\,s})^2/(\mathrm{kg\,m^2})=\mathrm{J}\), an energy, matching \([E_1]\).

Problems
  1. For \(\hat{A}=\begin{pmatrix}3&0\\0&5\end{pmatrix}\), use a trial vector \((\cos\theta,\sin\theta)\) to write \(R(\theta)\) and find its extrema.
    Solution \(R(\theta)=3\cos^2\theta+5\sin^2\theta = 3+2\sin^2\theta\). Then \(dR/d\theta = 4\sin\theta\cos\theta=2\sin2\theta=0\) at \(\theta=0,\pi/2\), giving \(R=3\) and \(R=5\). So \(\lambda_1=3\) with eigenvector \((1,0)\) and \(\lambda_2=5\) with eigenvector \((0,1)\), matching the diagonal entries.
  2. A normalised trial state is \(\psi = \sqrt{0.9}\,\varphi_1 + \sqrt{0.1}\,\varphi_2\) with \(\lambda_1=1\,\mathrm{eV}\) and \(\lambda_2=4\,\mathrm{eV}\). Compute \(R\) and the error above \(\lambda_1\).
    Solution Weights \(w_1=0.9,\ w_2=0.1\). \(R=0.9(1)+0.1(4)=0.9+0.4=1.3\,\mathrm{eV}\). Error \(R-\lambda_1=0.3\,\mathrm{eV}\), i.e. \(30\%\) above the ground state. A \(10\%\) admixture in probability produced a \(30\%\) energy error because \(\lambda_2-\lambda_1\) is large; the second-order smallness applies to the amplitude \(\sqrt{0.1}\), not to the probability.
  3. Prove that for any \(\psi\), \(R(\psi)\le\lambda_{\max}\) using the eigenexpansion, and state when equality holds.
    Solution Write \(R=\sum_n\lambda_n w_n\) with \(w_n\ge0\) and \(\sum_n w_n=1\). Since \(\lambda_n\le\lambda_{\max}\), we get \(R=\sum_n\lambda_n w_n\le\lambda_{\max}\sum_n w_n=\lambda_{\max}\). Equality requires \(w_n=0\) whenever \(\lambda_n<\lambda_{\max}\), i.e. \(\psi\) lies entirely in the top eigenspace \(\ker(\hat A-\lambda_{\max})\).
  4. For the harmonic oscillator \(\hat H = -\frac{\hbar^2}{2m}\frac{d^2}{dx^2}+\frac{1}{2} m\omega^2 x^2\), use the Gaussian trial \(\psi=e^{-\alpha x^2}\) and minimise \(R(\alpha)\).
    Solution Standard Gaussian integrals give \(\langle T\rangle=\frac{\hbar^2\alpha}{2m}\) and \(\langle V\rangle=\frac{m\omega^2}{8\alpha}\), so \(R(\alpha)=\frac{\hbar^2\alpha}{2m}+\frac{m\omega^2}{8\alpha}\). Set \(dR/d\alpha=\frac{\hbar^2}{2m}-\frac{m\omega^2}{8\alpha^2}=0\Rightarrow\alpha^2=\frac{m^2\omega^2}{4\hbar^2}\Rightarrow\alpha=\frac{m\omega}{2\hbar}\). Substituting, \(R=\frac{\hbar^2}{2m}\cdot\frac{m\omega}{2\hbar}+\frac{m\omega^2}{8}\cdot\frac{2\hbar}{m\omega}=\frac{\hbar\omega}{4}+\frac{\hbar\omega}{4}=\frac{\hbar\omega}{2}\). This equals the exact ground state \(E_0=\tfrac12\hbar\omega\) because the Gaussian is the true eigenstate, so the bound is saturated.
  5. Using Courant–Fischer, show that if \(\hat A\le\hat B\) as quadratic forms (\(\langle\psi,\hat A\psi\rangle\le\langle\psi,\hat B\psi\rangle\) for all \(\psi\)), then \(\lambda_n(\hat A)\le\lambda_n(\hat B)\) for every \(n\).
    Solution For any subspace \(S\) with \(\dim S=n\) and any \(\psi\in S\), \(R_A(\psi)=\frac{\langle\psi,\hat A\psi\rangle}{\|\psi\|^2}\le\frac{\langle\psi,\hat B\psi\rangle}{\|\psi\|^2}=R_B(\psi)\). Taking the max over \(\psi\in S\): \(\max_{\psi\in S}R_A(\psi)\le\max_{\psi\in S}R_B(\psi)\). Now take the min over all \(n\)-dimensional \(S\): \(\lambda_n(\hat A)=\min_S\max_\psi R_A\le\min_S\max_\psi R_B=\lambda_n(\hat B)\). This is the monotonicity theorem underlying Dirichlet–Neumann bracketing and perturbation bounds.