physics2u
Tier
⌕ Search ⌘K
Derivation

Rayleigh-Ritz Variational Principle

D-224 Home PU-301 Threads energy · chance Depends on Spectral Theorem for Hermitian Observables
Statement

For any normalisable trial state \( |\psi\rangle \) in the domain of a Hamiltonian \( \hat{H} \) whose spectrum is bounded below with ground-state energy \( E_0 \), the Rayleigh quotient satisfies \( \displaystyle E[\psi] \equiv \frac{\langle \psi | \hat{H} | \psi \rangle}{\langle \psi | \psi \rangle} \ge E_0 \), with equality if and only if \( |\psi\rangle \) is a ground state.

Why it matters

The variational principle turns an intractable eigenvalue problem into a minimisation. Rather than solving \( \hat{H}|\psi\rangle = E|\psi\rangle \) exactly, we choose a parametrised family of trial states and minimise the energy over the parameters; the minimum is a rigorous ceiling on the true ground-state energy that only ever improves as the family is enlarged.

It underlies almost every practical method of quantum chemistry and condensed-matter physics, from the Hartree–Fock approximation to the ground-state energy of helium, and it gives the tightest possible one-sided error control: the answer is guaranteed to be an over-estimate, never a wild guess in an unknown direction.

Assumptions
\( \hat{H} \) is Hermitian (self-adjoint) with a complete orthonormal eigenbasis.Without self-adjointness the eigenvalues may be complex and the ordering "\( \ge E_0 \)" is meaningless; the spectral decomposition used in the proof does not exist.
The spectrum is bounded below, \( E_0 = \inf \operatorname{spec}(\hat{H}) > -\infty \).If the energy can go to \( -\infty \) (e.g. an attractive \( -1/r^2 \) potential that "falls to the centre") there is no ground state to bound and the infimum of \( E[\psi] \) is \( -\infty \).
\( |\psi\rangle \) is normalisable and lies in the domain of \( \hat{H} \).A non-normalisable or too-singular trial function makes \( \langle\psi|\hat{H}|\psi\rangle \) or \( \langle\psi|\psi\rangle \) diverge, so the Rayleigh quotient is undefined.
The eigenstates form a genuine basis (discrete plus continuous spectrum resolved by the spectral theorem).If one drops the completeness guaranteed by the spectral theorem for self-adjoint operators, the expansion \( |\psi\rangle = \sum_n c_n |n\rangle \) may miss part of the state and the inequality can be violated.
Derivation
1
\[ \hat{H}|n\rangle = E_n |n\rangle, \qquad \langle m|n\rangle = \delta_{mn}, \qquad E_0 \le E_1 \le E_2 \le \cdots \]
By the spectral theorem for self-adjoint observables (prior result), \( \hat{H} \) has a complete orthonormal eigenbasis with real eigenvalues; order them so \( E_0 \) is the smallest. A
2
\[ |\psi\rangle = \sum_n c_n |n\rangle, \qquad c_n = \langle n | \psi \rangle \]
Completeness of the eigenbasis lets any state in the domain be expanded; the coefficients are the projections onto each eigenvector. A
3
\[ \langle \psi | \psi \rangle = \sum_{m,n} c_m^{*} c_n \langle m | n \rangle = \sum_n |c_n|^2 \]
Insert the expansion into the norm and use orthonormality \( \langle m|n\rangle=\delta_{mn} \) to collapse the double sum. A
4
\[ \langle \psi | \hat{H} | \psi \rangle = \sum_{m,n} c_m^{*} c_n \, E_n \, \langle m | n \rangle = \sum_n E_n |c_n|^2 \]
Act with \( \hat{H} \) on each \( |n\rangle \) to pull out \( E_n \), then apply orthonormality again. A
5
\[ \langle \psi | \hat{H} | \psi \rangle - E_0 \langle \psi | \psi \rangle = \sum_n (E_n - E_0)\,|c_n|^2 \]
Subtract \( E_0 \) times the norm (step 3) from the energy (step 4), combining the two sums term by term. B
6
\[ E_n - E_0 \ge 0 \ \ \text{and}\ \ |c_n|^2 \ge 0 \quad \Longrightarrow \quad \sum_n (E_n - E_0)\,|c_n|^2 \ge 0 \]
Every factor is non-negative because \( E_0 \) is the smallest eigenvalue and \( |c_n|^2 \) is a modulus squared; a sum of non-negative terms is non-negative. B
7
\[ \langle \psi | \hat{H} | \psi \rangle \ge E_0 \, \langle \psi | \psi \rangle \quad \Longrightarrow \quad E[\psi] = \frac{\langle \psi | \hat{H} | \psi \rangle}{\langle \psi | \psi \rangle} \ge E_0 \]
Rearrange the inequality from step 6 and divide by the strictly positive norm \( \langle\psi|\psi\rangle>0 \) (allowed because dividing by a positive number preserves direction). B
8
\[ E[\psi] = E_0 \iff (E_n-E_0)|c_n|^2 = 0 \ \forall n \iff c_n = 0 \ \text{whenever } E_n > E_0 \]
Equality in a sum of non-negative terms requires each term to vanish, so \( |\psi\rangle \) has support only in the \( E_0 \) eigenspace — it is a ground state. C
Result
\[ E[\psi] = \frac{\langle \psi | \hat{H} | \psi \rangle}{\langle \psi | \psi \rangle} \ge E_0, \qquad \text{equality} \iff \hat{H}|\psi\rangle = E_0 |\psi\rangle \]

Reading. The average energy measured in any state you can write down is at least the true ground-state energy. A trial wavefunction can never fool the Hamiltonian into reporting an energy below the floor; the best you can do is hit the floor exactly, which happens precisely when the trial state is the ground state. Minimising \( E[\psi] \) over a family therefore squeezes the estimate down towards \( E_0 \) from above.

Units check. \( \langle\psi|\hat{H}|\psi\rangle \) carries units of (state)\(^2\times\)energy and \( \langle\psi|\psi\rangle \) carries units of (state)\(^2\); their ratio is an energy, in joules (SI) or electronvolts, matching \( E_0 \). The inequality compares two quantities of identical dimension, as required.

Limiting cases
  • Trial state equals the exact ground state: \( c_0=1 \), all other \( c_n=0 \), so \( E[\psi]=E_0 \) — the bound is saturated.
  • Trial state is a single excited eigenstate \( |k\rangle \): \( E[\psi]=E_k \ge E_0 \) — the bound holds but is loose by \( E_k-E_0 \).
  • Small admixture of excited states, \( |\psi\rangle=|0\rangle+\epsilon|1\rangle \): \( E[\psi]-E_0 = \dfrac{\epsilon^2 (E_1-E_0)}{1+\epsilon^2} = O(\epsilon^2) \) — a first-order error in the wavefunction gives only a second-order error in the energy.
  • Degenerate ground state (\( E_0=E_1 \)): equality holds for any superposition within the degenerate subspace, not just one vector.
Breaks when
  • Unbounded-below spectrum. For potentials singular enough that the energy has no floor (e.g. \( V=-\lambda/r^2 \) with \( \lambda \) beyond the critical coupling, or an unregularised \( -1/r^4 \)), \( \inf E[\psi]=-\infty \) and there is no \( E_0 \) to bound — the principle is vacuous.
  • Non-self-adjoint or open (non-Hermitian) Hamiltonians. With gain/loss or complex potentials the eigenvalues are complex, "\( \ge \)" is undefined, and the real part of \( E[\psi] \) can dip below \( \operatorname{Re}(E_0) \); the spectral-theorem step fails.
  • Trial state outside the operator domain. A wavefunction with a discontinuous slope at a potential step, or one whose kinetic energy integral \( \int|\nabla\psi|^2 \) diverges, makes \( \langle\hat{H}\rangle \) ill-defined and can spuriously return a value below \( E_0 \).
  • Excited-state targeting without a symmetry constraint. The bound is only for the ground state; naively minimising for an excited level, without projecting out lower states or fixing a symmetry sector, collapses back to \( E_0 \).
Failure modes
  • Forgetting the normalisation denominator. Minimising \( \langle\psi|\hat{H}|\psi\rangle \) alone rather than the Rayleigh quotient; the true stationary condition is \( \langle\psi|\hat{H}|\psi\rangle - E\langle\psi|\psi\rangle \) stationary, giving \( \hat{H}|\psi\rangle=E|\psi\rangle \) only with the denominator present.
  • Believing a lower number is always better. Students take a variational estimate below a published value as "more accurate"; if it is below \( E_0 \) it signals a bug (a domain violation or a sign error), because the bound is one-sided.
  • Using a complex parameter as if real. Dropping the modulus in \( |c_n|^2 \) or in a complex variational parameter and getting a "lower" but non-physical energy.
  • Confusing \( E[\psi]\ge E_0 \) with \( \langle\hat{H}\rangle=E_0 \). Reporting the trial energy as the ground-state energy rather than as an upper bound.
  • Not enlarging the basis monotonically. Removing a basis function and expecting the energy to keep falling; adding functions can only lower (or hold) the minimum, never raise it — removing them does the opposite.
Discussion

The variational principle is really a statement about a Rayleigh quotient of a bounded-below self-adjoint operator: its infimum over the whole Hilbert space equals the smallest eigenvalue. What makes it powerful in practice is the quadratic insensitivity shown in the limiting cases — a wavefunction that is wrong at first order in some small parameter gives an energy that is wrong only at second order. This is why crude trial functions (a single Gaussian for the hydrogen atom, a product of one-electron orbitals for helium) already capture most of the binding energy.

The stationarity form of the principle is the true engine: setting \( \delta E[\psi]=0 \) subject to normalisation yields, via a Lagrange multiplier, the Schrödinger equation itself, with the multiplier identified as the energy. Thus every eigenstate is a stationary point of the Rayleigh quotient — the ground state is the global minimum, and higher eigenstates are saddle points. Restricting the variation to a finite subspace and demanding stationarity gives the Rayleigh–Ritz matrix eigenvalue problem \( \mathbf{H}\mathbf{c} = E\,\mathbf{S}\mathbf{c} \), whose lowest root bounds \( E_0 \) — and, by the Hylleraas–Undheim / MacDonald theorem, whose \( k \)-th root bounds the \( k \)-th exact eigenvalue from above.

On the "chance" thread, the proof is a probabilistic statement in disguise: \( |c_n|^2 \) is the probability of measuring energy \( E_n \) in the trial state, so \( E[\psi]=\sum_n E_n |c_n|^2 \) is literally the expectation value of the energy random variable. The inequality \( \sum_n E_n|c_n|^2 \ge E_0\sum_n|c_n|^2 \) is then nothing more than the elementary fact that the mean of a distribution supported on \( \{E_n\} \) cannot lie below the smallest value in its support. The variational principle is a quantum expectation bounded by the minimum of its outcome set.

Common misconceptions. The bound is not that "the ground state minimises energy" in a classical sense — the trial state need not evolve or relax; the inequality holds instantaneously for any state you write down. Nor does a lower variational number automatically mean a better wavefunction in every respect: energy converges quadratically while other observables (dipole moments, densities) converge only linearly, so a good energy can coexist with a mediocre wavefunction.

Worked examples
1
\[ \hat{H} = -\frac{\hbar^2}{2m}\frac{d^2}{dx^2} + \tfrac{1}{2}m\omega^2 x^2, \qquad \psi(x)=e^{-\alpha x^2} \]
Take a Gaussian trial function for the 1D harmonic oscillator with one variational width parameter \( \alpha>0 \). A
2
\[ E(\alpha)=\frac{\langle\psi|\hat{H}|\psi\rangle}{\langle\psi|\psi\rangle}=\frac{\hbar^2\alpha}{2m}+\frac{m\omega^2}{8\alpha} \]
Evaluate the Gaussian integrals: \( \langle T\rangle=\hbar^2\alpha/2m \) and \( \langle V\rangle=m\omega^2/8\alpha \), then form the quotient. B
3
\[ \frac{dE}{d\alpha}=\frac{\hbar^2}{2m}-\frac{m\omega^2}{8\alpha^2}=0 \;\Rightarrow\; \alpha_\star=\frac{m\omega}{2\hbar} \]
Minimise over the parameter; the stationary point is the width that balances kinetic and potential contributions. B
4
\[ E(\alpha_\star)=\frac{\hbar^2}{2m}\cdot\frac{m\omega}{2\hbar}+\frac{m\omega^2}{8}\cdot\frac{2\hbar}{m\omega}=\tfrac{1}{4}\hbar\omega+\tfrac{1}{4}\hbar\omega=\tfrac{1}{2}\hbar\omega \]
Substitute \( \alpha_\star \) back into \( E(\alpha) \). A
\[ E_{\min}=\tfrac{1}{2}\hbar\omega = E_0 \]

Reading. Because the Gaussian family contains the exact ground state, the bound is saturated. Numerically, with \( \hbar\omega = 1\ \text{eV} \), \( E_{\min}=0.500\ \text{eV} \) exactly — a check that the method returns the known answer when the trial family is rich enough.

1
\[ \hat{H}=-\frac{\hbar^2}{2m}\nabla^2-\frac{e^2}{4\pi\varepsilon_0 r}, \qquad \psi(r)=e^{-r/a} \]
Hydrogen atom with a single-exponential trial function of variational length scale \( a \) (a "wrong-guess" width, not fixed to the Bohr radius). A
2
\[ E(a)=\frac{\hbar^2}{2m a^2}-\frac{e^2}{4\pi\varepsilon_0\, a} \]
Standard hydrogenic integrals give kinetic \( \hbar^2/2ma^2 \) and potential \( -e^2/(4\pi\varepsilon_0 a) \) for the normalised \( 1s \)-type state. B
3
\[ \frac{dE}{da}=-\frac{\hbar^2}{m a^3}+\frac{e^2}{4\pi\varepsilon_0\, a^2}=0 \;\Rightarrow\; a_\star=\frac{4\pi\varepsilon_0\hbar^2}{m e^2}=a_0 \]
Minimise; the optimal length scale is exactly the Bohr radius \( a_0\approx 5.29\times10^{-11}\ \text{m} \). B
4
\[ E(a_0)=-\frac{m e^4}{2(4\pi\varepsilon_0)^2\hbar^2}=-\frac{1}{2}\,\frac{e^2}{4\pi\varepsilon_0\, a_0}=-13.6\ \text{eV} \]
Insert \( a_\star=a_0 \); the two terms combine to the Rydberg energy. A
\[ E_{\min}=-13.6\ \text{eV} = E_0 \]

Reading. The exponential family again contains the true \( 1s \) ground state, so the variational estimate equals the exact \( -13.6\ \text{eV} \). Had we frozen \( a \) at, say, \( 2a_0 \), we would get \( E(2a_0)=-10.2\ \text{eV} > E_0 \): still a valid upper bound, just looser — a concrete illustration of "never below the floor".

Problems
  1. A trial state is an equal-amplitude mix of the two lowest eigenstates, \( |\psi\rangle=\tfrac{1}{\sqrt2}(|0\rangle+|1\rangle) \), with \( E_0=1.0\ \text{eV} \) and \( E_1=4.0\ \text{eV} \). Compute \( E[\psi] \) and confirm the bound.
    Solution \( E[\psi]=\tfrac12 E_0+\tfrac12 E_1=\tfrac12(1.0)+\tfrac12(4.0)=2.5\ \text{eV} \). Since \( 2.5\ \text{eV} \ge E_0=1.0\ \text{eV} \), the bound holds; the excess \( 1.5\ \text{eV} \) is the contribution \( \tfrac12(E_1-E_0) \) from the unwanted excited-state weight.
  2. For the harmonic-oscillator Gaussian trial function, suppose the width is fixed too broad at \( \alpha=\tfrac{m\omega}{4\hbar} \) (half the optimum). Using \( E(\alpha)=\dfrac{\hbar^2\alpha}{2m}+\dfrac{m\omega^2}{8\alpha} \), find \( E \) in units of \( \hbar\omega \) and the fractional error.
    Solution First term: \( \dfrac{\hbar^2}{2m}\cdot\dfrac{m\omega}{4\hbar}=\dfrac{\hbar\omega}{8} \). Second: \( \dfrac{m\omega^2}{8}\cdot\dfrac{4\hbar}{m\omega}=\dfrac{\hbar\omega}{2} \). Total \( E=\dfrac{\hbar\omega}{8}+\dfrac{\hbar\omega}{2}=\dfrac{5}{8}\hbar\omega=0.625\,\hbar\omega \). Error over \( E_0=0.5\,\hbar\omega \) is \( 0.125/0.5=25\% \). Still \( \ge E_0 \), as required.
  3. A trial function has coefficients \( c_0=0.8,\ c_1=0.6 \) (already normalised, \( |c_0|^2+|c_1|^2=1 \)) over levels \( E_0=-13.6\ \text{eV},\ E_1=-3.4\ \text{eV} \). Compute \( E[\psi] \) and the gap to \( E_0 \).
    Solution \( |c_0|^2=0.64,\ |c_1|^2=0.36 \). \( E[\psi]=0.64(-13.6)+0.36(-3.4)=-8.704-1.224=-9.928\ \text{eV} \). Gap: \( E[\psi]-E_0=-9.928-(-13.6)=3.672\ \text{eV}\ge 0 \). Bound satisfied.
  4. Show, for a real one-parameter family \( E(\lambda) \), that the variational minimum makes the energy error second order in the wavefunction error. Take \( |\psi\rangle=|0\rangle+\lambda|1\rangle \) (unnormalised) with orthonormal \( |0\rangle,|1\rangle \) and eigenvalues \( E_0,E_1 \); expand \( E(\lambda) \) for small \( \lambda \).
    Solution \( \langle\psi|\psi\rangle=1+\lambda^2 \); \( \langle\psi|\hat H|\psi\rangle=E_0+\lambda^2 E_1 \). So \( E(\lambda)=\dfrac{E_0+\lambda^2 E_1}{1+\lambda^2}=(E_0+\lambda^2 E_1)(1-\lambda^2+\cdots)=E_0+\lambda^2(E_1-E_0)+O(\lambda^4) \). The wavefunction error is \( O(\lambda) \) but the energy error \( E-E_0=\lambda^2(E_1-E_0) \) is \( O(\lambda^2) \): quadratic, and non-negative since \( E_1>E_0 \).
  5. A student obtains a variational energy of \( -14.2\ \text{eV} \) for hydrogen (true \( E_0=-13.6\ \text{eV} \)) and claims a better ground-state estimate. Explain what must have gone wrong and give one specific culprit.
    Solution The variational principle forbids \( E[\psi]<E_0 \) for a valid trial state, so \( -14.2\ \text{eV}<-13.6\ \text{eV} \) signals an error, not an improvement. Likely culprits: (i) a trial function outside the domain — e.g. a cusp or discontinuous derivative making the kinetic integral \( \int|\nabla\psi|^2 \) under-counted or mis-integrated by parts; (ii) an arithmetic/sign error in \( \langle T\rangle \) or \( \langle V\rangle \); or (iii) forgetting the normalisation denominator. The correct interpretation is that the true answer is at most \( -13.6\ \text{eV} \), so \( -14.2\ \text{eV} \) is unphysical.