Equivalence of Ensembles in the Thermodynamic Limit
Statement
The canonical partition function is the Laplace transform of the microcanonical density of states, Z(β) = ∫ dE Ω(E) e−βE. Evaluating this integral by steepest descent shows that, for any system whose entropy is extensive and concave (heat capacity C > 0), the integrand is a sharp Gaussian of relative width ∼ 1/√N centred on the energy E* where the microcanonical temperature equals the bath temperature. Consequently the canonical free energy equals the Legendre transform of the microcanonical entropy, F(T) = E* − T S(E*), up to corrections of relative order (ln N)/N → 0.
Why it matters
Statistical mechanics offers several ways to describe the same equilibrium system — fixed energy (microcanonical), fixed temperature (canonical), fixed chemical potential (grand canonical). If these disagreed about measurable thermodynamics, the whole framework would be ambiguous. This derivation is the concrete demonstration that, in the thermodynamic limit, they do not disagree: choosing an ensemble is a matter of calculational convenience, not physics.
The same saddle-point machinery reappears everywhere — large-deviation theory, the method of steepest descent in field theory, and the Legendre-transform structure that connects every pair of thermodynamic potentials. Understanding it once here explains why fluctuations of extensive quantities are always suppressed by 1/√N.
Assumptions
Derivation
Result
Reading. The canonical free energy is the Legendre transform of the microcanonical entropy: pass from the variable E to its conjugate 1/T = S′(E), and F = E* − TS emerges automatically. The neglected term is a single logarithm — smaller than F by a factor (ln N)/N. Physically, a system held at fixed T has an energy that is not fixed but is pinned to E*(T) with fractional scatter only 1/√N; on macroscopic scales it is indistinguishable from a system held at fixed energy E*. That is the equivalence of ensembles.
Units check. The exponent S/k − βE is dimensionless (both terms pure numbers), as an argument of exp must be. The variance kT2C has units [J K−1][K2][J K−1] = J2, so σE is an energy. In F, each of E*, TS, and kT ln(·) carries units of J. Consistent.
Limiting cases
- Large N: the log correction and the 1/√N spread both vanish relative to the extensive terms — the ensembles coincide exactly per particle.
- Finite N (small system): the Gaussian broadens, F picks up the measurable −(kT/2) ln(2πkT2C) term, and the two ensembles differ at 𝒪(1/N) — the regime of nanothermodynamics.
- High temperature (ideal gas): C = (3/2)Nk is constant, S(E) ∝ ln E is smooth and concave, and equivalence is textbook-clean.
- Approaching a critical point: C diverges, so σE ∼ √C grows; the peak is still sharp relative to E* ∼ N as long as C grows slower than N2, but the approach to equivalence slows.
Breaks when
- Long-range interactions (gravitating stars, plasmas, small nuclei): the energy is non-additive, S(E) can be non-concave, and the microcanonical ensemble supports negative-heat-capacity states that the canonical ensemble simply cannot reach. The two ensembles predict different equilibria — the archetypal failure.
- At a first-order phase transition: S(E) has a straight coexistence segment (S″ = 0, C → ∞), the quadratic saddle expansion collapses, and the canonical ensemble discontinuously hops between the two phase energies while the microcanonical ensemble occupies the forbidden intermediate energies.
- Genuinely small or mesoscopic systems: when N is not large the 1/√N fluctuations and the log correction are experimentally resolvable, and ensemble choice becomes a physical statement about the actual constraints (isolated vs. thermostatted), not a convenience.
Failure modes
- Treating the saddle point as a maximum without checking C > 0. If S″ > 0 the stationary point is a minimum of the integrand and the naive Gaussian formula gives an imaginary width — a signal that steepest descent must be deformed off the real axis, not that the answer is complex.
- Confusing E* with the microcanonical energy of the isolated system. E* is the most probable canonical energy; it coincides with a fixed E only in the limit, and equating them at finite N discards the fluctuation term.
- Dropping the √(2πkT2C) prefactor and then claiming F = E* − TS is exact. It is exact only per particle in the limit; the prefactor is the whole content of the kT ln N correction.
- Using CV vs CP inconsistently. The energy fluctuation formula uses the heat capacity at the fixed variables of the ensemble (here constant V, N); substituting CP gives the wrong width.
- Assuming higher Taylor terms are always negligible. Near a critical point the third and fourth derivatives of φ matter (non-Gaussian fluctuations); the quadratic truncation is a large-N, away-from-criticality statement.
Discussion
The heart of the argument is that an integral of eNf(x) with N large is completely controlled by the single point where f is stationary. Statistical mechanics manufactures exactly this structure because entropy and energy are both extensive: the competition S/k − βE between maximising entropy and minimising energy is a battle between two 𝒪(N) quantities, and its outcome is pinned to microscopic sharpness. The Legendre transform F = E* − TS is not an extra postulate; it is what steepest descent produces.
The consistency in Step 9 — that the saddle width equals the independently derived fluctuation kT2C — is worth dwelling on. The canonical distribution of energy, P(E) ∝ Ω(E) e−βE, is a Gaussian centred on E* with this variance. The Laplace-transform relation and the Boltzmann-weight relation are two views of one object, and they must and do give the same fluctuation. This is the microscopic origin of the fluctuation–dissipation link between the width of energy fluctuations and the response coefficient C.
The equivalence generalises to every conjugate pair: (particle number, chemical potential), (volume, pressure), (magnetisation, field). Each grand or generalised ensemble is a further Laplace transform, each is evaluated at its own saddle, and each agrees with the others in the thermodynamic limit. The whole edifice of thermodynamic potentials related by Legendre transforms is the large-N shadow of these saddle-point identities.
Modern large-deviation theory makes the statement sharp: ensemble equivalence holds if and only if the microcanonical entropy density s(e) = lim S/N is concave. When s(e) has a non-concave dip, its canonical counterpart φ(β) = lim (ln Z)/N is the concave envelope (Legendre–Fenchel transform) of s, and the transform is not invertible: a whole interval of energies collapses to one temperature. The microcanonical description then strictly contains more information than the canonical one, and negative-heat-capacity states — forbidden canonically because σE2 = kT2C ≥ 0 — live perfectly well microcanonically. Self-gravitating systems, which grow hotter as they lose energy, are the standard physical realisation.
Common misconceptions. "The ensembles are always equivalent" — false; equivalence is a theorem about short-range, concave-entropy systems in the limit, and it fails for gravity and across first-order transitions. "Fixed temperature means fixed energy" — no; it means energy fluctuates with relative size 1/√N, which merely happens to be unmeasurable for a mole. "Saddle point means the integrand is small there" — on the contrary, for a concave exponent the saddle is the maximum and carries essentially all the integral.
Worked examples
Reading. The canonical energy of a mole of gas at room temperature is pinned to one part in 1012. No thermometer can resolve this, so "fixed T" and "fixed E" describe operationally identical macrostates — the equivalence in numbers.
Reading. Even a bounded-energy system with a finite C shows the universal 1/√N = 10−10 suppression (the sech factor shifts it only by a number of order one). The canonical and microcanonical descriptions of this paramagnet agree to ten significant figures.
Problems
- Starting from Z = ∫ dE eS(E)/k − βE, show directly that the saddle-point condition reproduces the thermodynamic definition of temperature, and interpret the result physically.
Solution
Set the derivative of the exponent to zero: d/dE[S(E)/k − βE] = S′(E*)/k − β = 0, so S′(E*) = kβ = 1/T. Since the microcanonical temperature is defined by 1/Tmc = ∂S/∂E, the dominant energy E* is precisely the one whose microcanonical temperature matches the bath temperature T. The Boltzmann weight e−βE pushes toward low energy, the phase-space factor eS/k pushes toward high energy, and they balance where the two "forces" −β and +S′/k are equal. - For a monatomic ideal gas of N = 1.0 mol at T = 300 K, compute the absolute energy fluctuation σE in joules and compare with U.
Solution
N = 6.02×1023. U = (3/2)NkT = 1.5 × 6.02×1023 × 1.38×10−23 × 300 = 3.74×103 J. CV = (3/2)Nk = 12.5 J/K. σE = √(kT2CV) = √(1.38×10−23 × 3002 × 12.5) = √(1.55×10−17) = 3.9×10−9 J. Ratio σE/U = 3.9×10−9/3.74×103 = 1.05×10−12, matching √(2/3N). The energy is fixed to a nanojoule out of kilojoules. - Estimate the neglected logarithmic correction −(kT/2) ln(2πkT2C) for the gas of Problem 2 and express it as a fraction of F ∼ −NkT.
Solution
With kT2C = σE2 = 1.55×10−17 J2, the argument 2πkT2C = 9.7×10−17, ln(·) = ln(9.7×10−17) ≈ −36.9. So the correction is −(kT/2)(−36.9) = +18.5 kT = 18.5 × 4.14×10−21 = 7.7×10−20 J. Compared with |F| ∼ NkT = 2.5×103 J, the fraction is 3×10−23 — utterly negligible, of order (ln N)/N. (The units mismatch inside the log is an artefact of not making the argument dimensionless; the physically meaningful statement is that the correction scales as kT ln N, hence 𝒪(ln N/N) relative to F.) - A hypothetical isolated system has microcanonical heat capacity C < 0 over some energy range (it heats up as it loses energy). Explain, using Step 6 and Step 7, why the canonical ensemble cannot reproduce these states, and name a physical example.
Solution
Step 6 gives the saddle curvature φ″ = −1/(kT2C). If C < 0 then φ″ > 0, so the stationary energy is a minimum of the integrand, not a maximum — the Gaussian in Step 7 would have variance σ2 = kT2C < 0, i.e. imaginary width, which is unphysical. Equivalently, the canonical fluctuation identity σE2 = kT2C forces C ≥ 0, so no canonical (fixed-T) system can ever exhibit negative heat capacity. Such states exist only microcanonically (fixed E), where energy cannot flow to a bath. The standard example is a self-gravitating system — a star cluster or gas cloud — whose long-range, non-additive gravity gives a non-concave S(E) and negative C; the ensembles are genuinely inequivalent there. - For the two-level paramagnet of Worked Example 2, at what temperature is the heat capacity (and hence σE at fixed N) maximal? Compute σE/|U| there for N = 1.0×1020.
Solution
The per-spin heat capacity c(x) = k x2 sech2x with x = βε peaks (the Schottky maximum) at x ≈ 1.20, i.e. kT = ε/1.20, T = ε/(1.20k) = 10−21/(1.20 × 1.38×10−23) = 60.4 K. There sech(1.20) = 0.552 and tanh(1.20) = 0.834. Using the closed forms from the example, σE = √N ε sechx = 1010 × 10−21 × 0.552 = 5.5×10−12 J, and |U| = Nε tanhx = 1020 × 10−21 × 0.834 = 8.34×10−2 J. Ratio = 5.5×10−12/8.34×10−2 = 6.6×10−11. Even at the Schottky peak, where fluctuations are largest, the relative width stays at the 1/√N = 10−10 scale, so the ensembles remain equivalent.