The Boltzmann distribution from maximum entropy
D-013 boltzmann-distribution Home PU-203 Threads chance · energy Depends on — verified
Statement
The probability distribution over states that maximises entropy at fixed mean energy is exponential in the energy, and the exponent defines temperature.
Why it matters
The Boltzmann factor is the single most-used result in physical science. Presented as an assertion it is a magic trick; derived, it is a statement about counting and nothing more. This route also makes clear that temperature is not a primitive concept — it is a Lagrange multiplier.
Assumptions
Derivation
Result
Reading. States of higher energy are exponentially less likely, with kBT setting the scale over which "less likely" becomes "essentially never". Every thermodynamic quantity follows from Z: U = −∂ln Z/∂β, F = −kBT ln Z, and so on.
Units check. The exponent is dimensionless: [J] / ([J K−1][K]). ✓
Limiting cases
- Uniform distribution over Ω equally-likely states recovers Boltzmann's S = kB ln Ω.
- High temperature (kBT ≫ level spacing): all accessible states become nearly equiprobable.
- Low temperature (kBT ≪ level spacing): the system collapses onto the ground state.
Breaks when
- The system is small enough that its energy fluctuations are comparable to its mean — the ensembles are then inequivalent.
- Long-range interactions make energy non-additive, so the reservoir decomposition fails. Self-gravitating systems are the standard pathology, and they can exhibit negative heat capacity.
- The system is not in equilibrium. Nothing here applies to a driven or relaxing system.
Failure modes
- Confusing states with energy levels. A level with degeneracy g contributes g e−βE, and forgetting the degeneracy is the most common numerical error in the topic.
- Treating Z as a mere normalisation. It encodes the complete thermodynamics; the normalisation role is incidental.
- Assuming the highest-probability state is where the system is found. With degeneracy the most probable energy can be far above the ground state — see D-014, where the peak of the speed distribution is not at zero.
Worked number
The ratio of populations of two levels is p2/p1 = e−ΔE/kBT, with Z cancelling. At 300 K, kBT = 0.0259 eV. For the CO vibrational spacing of 0.27 eV from D-010, the ratio is e−10.4 ≈ 3 × 10−5 — three molecules in a hundred thousand are vibrationally excited. This one line explains the frozen-out heat capacity of D-020.
Discussion
The single most important thing to see is that the exponential form is not chosen — it is forced. When a large isolated composite (system + reservoir) is free to share a fixed total energy, the fundamental postulate weights every microstate equally. Overwhelmingly the most microstates correspond to the reservoir absorbing energy Ei whenever the small system sits in state i, and the number of reservoir microstates grows as eSR(U−Ei)/kB. Taylor-expanding the reservoir entropy to first order in the small quantity Ei produces exactly e−Ei/kBT. The max-entropy route in the derivation and this counting route are the same statement dressed differently: the exponential is what you get when the only thing you know is the mean energy.
The reason a single scalar function Z carries the entire thermodynamics is that the energy constraint entered the exponent linearly. That makes Z(β) a generating function: each derivative with respect to β pulls down another power of E. Hence U = −∂lnZ/∂β is the mean, and the second derivative gives the variance, ∂²lnZ/∂β² = ⟨E²⟩ − ⟨E⟩² = kBT²CV. That last equality is a fluctuation–response relation of exactly the type that recurs across physics: the equilibrium fluctuations of a quantity and the system's response to a conjugate perturbation are one and the same object. The free energy F = −kBT lnZ then makes Z the bridge between microscopic spectrum and macroscopic thermodynamics.
Temperature appearing as a Lagrange multiplier is not a curiosity of this derivation — it is the deepest content. In Jaynes' information-theoretic reading, the Boltzmann distribution is simply the least-committal probability assignment (maximum Shannon/Gibbs entropy) consistent with a known mean, and β is the multiplier enforcing that mean. The pair (S as a function of U) and (lnZ as a function of β) are Legendre transforms of one another, the same duality that links Lagrangian and Hamiltonian, or internal energy and free energy. This viewpoint also predicts its own generalisations without new ideas: add a number constraint and you get a second multiplier μ/kBT and the grand canonical distribution; constrain volume-conjugate or magnetisation-conjugate quantities and you get the corresponding ensembles. It even permits T < 0: if the energy spectrum is bounded above (as for a fixed set of spins), a distribution that favours high-energy states is still a legitimate maximiser, and the multiplier β simply goes negative — negative temperatures are hotter than infinite temperature, not colder than zero.
Common misconceptions. (i) That the Boltzmann factor gives the probability of an energy. It gives the probability of a state; the probability of an energy level of degeneracy g carries an extra factor g, which is why the most-populated level need not be the ground level (see the rotational example below). (ii) That Z is "just normalisation". Normalisation is incidental; Z encodes the full thermodynamics. (iii) That temperature is a primitive that causes the distribution. Here it is the other way round: fixing the mean energy defines a multiplier, and that multiplier is what we have always called 1/kBT.
Worked examples
Example 1 — A single spin in a magnetic field (two-level system). An electron magnetic moment of magnitude μ = μB = 9.274×10−24 J T−1 sits in a field B = 2 T at T = 4 K. Find the population ratio, the fractional populations, and the mean energy per spin.
Answer. About 66% of spins align with the field and 34% oppose it, giving a mean energy of −6.0×10−24 J per spin (a polarisation of only 0.32). Even at 4 K and 2 T the sample is far from fully aligned, because μB is still a fraction of kBT — the origin of Curie's law, in which magnetisation scales as B/T in the weak-field regime.
Example 2 — The most-populated rotational level of CO (degeneracy matters). A linear rotor has levels EJ = BrotJ(J+1) with degeneracy gJ = 2J+1. For CO the rotational constant is B̃ = 1.93 cm−1. Which level is most heavily populated at 300 K?
Answer. The most populated rotational level of CO at 300 K is J = 7, with P7 ≈ 8.9 P0 — nearly nine times the ground level, even though it lies 56Brot higher. The peak sits well above the ground state purely because of the 2J+1 degeneracy, the same mechanism that shifts the peak of the Maxwell speed distribution away from zero (D-014).
Problems
- (Warm-up) Two non-degenerate levels are split by ΔE = 0.10 eV. What fraction of systems occupies the upper level at 300 K, and at 1000 K? (kBT = 0.0259 eV at 300 K.)
Solution
Ratio r = pu/pℓ = e−ΔE/kBT. At 300 K: r = e−0.10/0.0259 = e−3.86 = 0.0210, so upper fraction = r/(1+r) = 0.0206 (about 2%). At 1000 K: kBT = 0.0259×(1000/300) = 0.0863 eV, r = e−1.16 = 0.314, fraction = 0.239 (about 24%). Raising T flattens the distribution toward equal occupation. - (Two-level thermodynamics) For a two-level system with energies 0 and ε (both non-degenerate), write Z and ⟨E⟩, evaluate ⟨E⟩/ε when βε = 1, and give the limits T→0 and T→∞.
Solution
Z = 1 + e−βε. Then ⟨E⟩ = ε e−βε/(1+e−βε) = ε/(eβε+1). At βε = 1: ⟨E⟩ = ε/(e+1) = ε/3.718 = 0.269 ε. Limits: as T→0 (βε→∞), ⟨E⟩→0 (system frozen in ground state); as T→∞ (βε→0), ⟨E⟩→ε/2 (both states equally likely). A two-level system therefore stores at most ε/2 — it saturates, unlike an unbounded spectrum. - (Degeneracy) A system has three levels at energies 0, ε, 2ε with degeneracies 1, 2, 1. Take ε = 0.025 eV and a temperature such that kBT = 0.025 eV (i.e. 290 K). Find Z and the probability of each level. Which single microstate is most probable?
Solution
Here βε = 1. Z = 1·e0 + 2·e−1 + 1·e−2 = 1 + 0.736 + 0.135 = 1.871. Level probabilities: P(0) = 1/1.871 = 0.534; P(ε) = 0.736/1.871 = 0.393; P(2ε) = 0.135/1.871 = 0.072. The middle level is boosted by its degeneracy, but per state the two states at ε each carry e−1/Z = 0.197, less than the ground state's 0.534. So the single most probable microstate is the non-degenerate ground state — the distinction between "most probable level" and "most probable state" is the whole point. - (Schottky heat capacity) For the two-level system of problem 2, show that the heat capacity is C = kB(βε)² eβε/(eβε+1)². It peaks at βε ≈ 2.40. For ε = 1.0 meV, at what temperature does the peak occur, and what is Cmax?
Solution
With ⟨E⟩ = ε/(eβε+1) and C = d⟨E⟩/dT, put x = βε = ε/kBT so dx/dT = −x/T. Then d⟨E⟩/dx = −εex/(ex+1)², giving C = (−εex/(ex+1)²)(−x/T) = kBx² ex/(ex+1)² (using ε/T = xkB). This is the Schottky anomaly: C→0 at both low and high T, with a bump in between. Setting dC/dx = 0 gives numerically x ≈ 2.40. At ε = 1.0 meV: T = ε/(kBx) = 1.0×10−3 eV/(8.617×10−5 eV K−1 × 2.40) = 4.8 K. And Cmax = kB(2.40)²e2.40/(e2.40+1)² = kB(5.76×11.02/144.5) = 0.44 kB per system. - (Capstone — quantum oscillator) A quantum harmonic oscillator has levels En = nℏω, n = 0, 1, 2, … (measure from the ground state; non-degenerate). (a) Sum the geometric series for Z. (b) Obtain ⟨E⟩ = −∂lnZ/∂β. (c) Evaluate ⟨E⟩ for the CO stretch, ℏω = 0.27 eV, at 300 K, and explain the frozen-out heat capacity of D-020.
Solution
(a) Z = Σn≥0 e−nβℏω = 1/(1 − e−βℏω) (a geometric series with ratio e−βℏω < 1). (b) lnZ = −ln(1 − e−βℏω), so ⟨E⟩ = −∂lnZ/∂β = ℏω/(eβℏω − 1) — the Planck–Einstein result. (Its high-T limit is kBT, recovering equipartition for a 1D oscillator; its low-T limit is ℏωe−βℏω → 0.) (c) βℏω = 0.27/0.0259 = 10.4, so e10.4 = 3.35×104 and ⟨E⟩ = 0.27/(3.35×104 − 1) = 8.1×10−6 eV per molecule — utterly negligible next to kBT = 0.026 eV. The mode is "frozen": it holds essentially no energy and its contribution to CV is exponentially suppressed (≈ kB(βℏω)²e−βℏω), which is precisely why the vibrational heat capacity of CO does not switch on until temperatures of order ℏω/kB ≈ 3100 K (D-020).