physics2u
Tier
⌕ Search ⌘K
Derivation · the atom

The Boltzmann distribution from maximum entropy

D-013 boltzmann-distribution Home PU-203 Threads chance · energy Depends on verified

Statement

The probability distribution over states that maximises entropy at fixed mean energy is exponential in the energy, and the exponent defines temperature.

Why it matters

The Boltzmann factor is the single most-used result in physical science. Presented as an assertion it is a magic trick; derived, it is a statement about counting and nothing more. This route also makes clear that temperature is not a primitive concept — it is a Lagrange multiplier.

Assumptions
All microstates of the isolated total system are equally likely.The fundamental postulate. It is a postulate, not a theorem; attempts to derive it from dynamics (ergodicity) succeed only for restricted classes of system.
The system is weakly coupled to a much larger reservoir.Interaction energy negligible compared with either subsystem's energy. Fails for small systems and for anything with long-range forces, where surface terms are not negligible.
Only the mean energy is constrained.Adding a constraint on particle number produces the grand canonical distribution and a second multiplier, the chemical potential.
Derivation
1
S = −kB Σi pi ln pi
The Gibbs entropy. For a uniform distribution over Ω states it reduces to kB ln Ω, recovering Boltzmann's expression, so this is a generalisation rather than a new assumption. A
2
Σi pi = 1 ,    Σi pi Ei = U
The two constraints. Normalisation, and a fixed mean energy — the latter is what "in contact with a reservoir at fixed temperature" amounts to. A
3
δ [ S − α kB Σ pi − β kB Σ pi Ei ] = 0
Maximise S subject to both constraints by the method of Lagrange multipliers. The multipliers α and β are undetermined for now; their meaning emerges rather than being imposed. B
4
−ln pi − 1 − α − β Ei = 0
Differentiate with respect to each pi independently and set to zero. Because the pi are now unconstrained (the constraints having been absorbed into the multipliers), each derivative vanishes separately. B
5
pi = e−1−α e−βEi ≡ (1/Z) e−βEi
Exponentiate. The first factor is a constant, fixed by normalisation, and is renamed 1/Z. The partition function is therefore not a definition imposed from outside — it is whatever makes the probabilities sum to one. B
6
∂S/∂U = kB β ≡ 1/T   ⟹   β = 1/(kBT)
Substitute the solution back into S and differentiate with respect to U. Comparing with the thermodynamic definition of temperature identifies the multiplier. Temperature has arrived as a derived quantity. C
Result
pi = e−Ei/kBT / Z ,    Z = Σi e−Ei/kBT

Reading. States of higher energy are exponentially less likely, with kBT setting the scale over which "less likely" becomes "essentially never". Every thermodynamic quantity follows from Z: U = −∂ln Z/∂β, F = −kBT ln Z, and so on.

Units check. The exponent is dimensionless: [J] / ([J K−1][K]). ✓

Limiting cases
  • Uniform distribution over Ω equally-likely states recovers Boltzmann's S = kB ln Ω.
  • High temperature (kBT ≫ level spacing): all accessible states become nearly equiprobable.
  • Low temperature (kBT ≪ level spacing): the system collapses onto the ground state.
Breaks when
  • The system is small enough that its energy fluctuations are comparable to its mean — the ensembles are then inequivalent.
  • Long-range interactions make energy non-additive, so the reservoir decomposition fails. Self-gravitating systems are the standard pathology, and they can exhibit negative heat capacity.
  • The system is not in equilibrium. Nothing here applies to a driven or relaxing system.
Failure modes
  • Confusing states with energy levels. A level with degeneracy g contributes g e−βE, and forgetting the degeneracy is the most common numerical error in the topic.
  • Treating Z as a mere normalisation. It encodes the complete thermodynamics; the normalisation role is incidental.
  • Assuming the highest-probability state is where the system is found. With degeneracy the most probable energy can be far above the ground state — see D-014, where the peak of the speed distribution is not at zero.
Worked number

The ratio of populations of two levels is p2/p1 = e−ΔE/kBT, with Z cancelling. At 300 K, kBT = 0.0259 eV. For the CO vibrational spacing of 0.27 eV from D-010, the ratio is e−10.4 ≈ 3 × 10−5 — three molecules in a hundred thousand are vibrationally excited. This one line explains the frozen-out heat capacity of D-020.

Run the check →

Discussion

The single most important thing to see is that the exponential form is not chosen — it is forced. When a large isolated composite (system + reservoir) is free to share a fixed total energy, the fundamental postulate weights every microstate equally. Overwhelmingly the most microstates correspond to the reservoir absorbing energy Ei whenever the small system sits in state i, and the number of reservoir microstates grows as eSR(UEi)/kB. Taylor-expanding the reservoir entropy to first order in the small quantity Ei produces exactly eEi/kBT. The max-entropy route in the derivation and this counting route are the same statement dressed differently: the exponential is what you get when the only thing you know is the mean energy.

The reason a single scalar function Z carries the entire thermodynamics is that the energy constraint entered the exponent linearly. That makes Z(β) a generating function: each derivative with respect to β pulls down another power of E. Hence U = −∂lnZ/∂β is the mean, and the second derivative gives the variance, ∂²lnZ/∂β² = ⟨E²⟩ − ⟨E⟩² = kBT²CV. That last equality is a fluctuation–response relation of exactly the type that recurs across physics: the equilibrium fluctuations of a quantity and the system's response to a conjugate perturbation are one and the same object. The free energy F = −kBT lnZ then makes Z the bridge between microscopic spectrum and macroscopic thermodynamics.

Temperature appearing as a Lagrange multiplier is not a curiosity of this derivation — it is the deepest content. In Jaynes' information-theoretic reading, the Boltzmann distribution is simply the least-committal probability assignment (maximum Shannon/Gibbs entropy) consistent with a known mean, and β is the multiplier enforcing that mean. The pair (S as a function of U) and (lnZ as a function of β) are Legendre transforms of one another, the same duality that links Lagrangian and Hamiltonian, or internal energy and free energy. This viewpoint also predicts its own generalisations without new ideas: add a number constraint and you get a second multiplier μ/kBT and the grand canonical distribution; constrain volume-conjugate or magnetisation-conjugate quantities and you get the corresponding ensembles. It even permits T < 0: if the energy spectrum is bounded above (as for a fixed set of spins), a distribution that favours high-energy states is still a legitimate maximiser, and the multiplier β simply goes negative — negative temperatures are hotter than infinite temperature, not colder than zero.

Common misconceptions. (i) That the Boltzmann factor gives the probability of an energy. It gives the probability of a state; the probability of an energy level of degeneracy g carries an extra factor g, which is why the most-populated level need not be the ground level (see the rotational example below). (ii) That Z is "just normalisation". Normalisation is incidental; Z encodes the full thermodynamics. (iii) That temperature is a primitive that causes the distribution. Here it is the other way round: fixing the mean energy defines a multiplier, and that multiplier is what we have always called 1/kBT.

Worked examples

Example 1 — A single spin in a magnetic field (two-level system). An electron magnetic moment of magnitude μ = μB = 9.274×10−24 J T−1 sits in a field B = 2 T at T = 4 K. Find the population ratio, the fractional populations, and the mean energy per spin.

1
Z = e+βμB + eβμB = 2 cosh(βμB)
Two states: aligned (E = −μB) and anti-aligned (E = +μB). The splitting is ΔE = 2μB. Keep everything symbolic first.
2
βμB = μB/kBT = (9.274×10−24 × 2)/(1.381×10−23 × 4) = 0.336
Evaluate the dimensionless exponent. Only now do numbers enter.
3
p+/p = e−2βμB = e−0.672 = 0.511
The ratio of anti-aligned to aligned populations; Z cancels, as always for a ratio.
4
p = e+0.336/2cosh(0.336) = 0.662,   p+ = 0.338
Divide each Boltzmann factor by Z = 2.114. The two fractions sum to 1, as they must.
5
E⟩ = μB(p+p) = −μB tanh(βμB) = −(1.855×10−23)(0.323)
The mean energy per spin. tanh(0.336) = 0.323 gives the polarisation directly.
E⟩ = −μB tanh(μB/kBT) = −6.0×10−24 J

Answer. About 66% of spins align with the field and 34% oppose it, giving a mean energy of −6.0×10−24 J per spin (a polarisation of only 0.32). Even at 4 K and 2 T the sample is far from fully aligned, because μB is still a fraction of kBT — the origin of Curie's law, in which magnetisation scales as B/T in the weak-field regime.

Example 2 — The most-populated rotational level of CO (degeneracy matters). A linear rotor has levels EJ = BrotJ(J+1) with degeneracy gJ = 2J+1. For CO the rotational constant is B̃ = 1.93 cm−1. Which level is most heavily populated at 300 K?

1
PJ ∝ (2J+1) eβBrotJ(J+1)
The population of a level is its degeneracy times the Boltzmann factor of a state in it. Dropping the 2J+1 is the classic error flagged in the failure modes.
2
d/dJ [ ln(2J+1) − βBrotJ(J+1) ] = 0  ⟹  Jmax = √(kBT/2Brot) − ½
Treat J as continuous and maximise the log-population. The competition is degeneracy (rising) against the Boltzmann factor (falling).
3
Brot = hc×(1.93 cm−1) = (1.986×10−23 J·cm)(1.93 cm−1) = 3.83×10−23 J
Convert the spectroscopic constant to energy with hc. Here hc = 1.986×10−23 J·cm.
4
kBT = 4.14×10−21 J;  Jmax = √(4.14×10−21/(2×3.83×10−23)) − ½ = 7.35 − 0.5 = 6.85
At T = 300 K. The nearest physical (integer) level is J = 7.
5
P7/P0 = 15 eBrot×56/kBT = 15 e−0.518 = 8.9
Compare the peak level (g = 15, E = 56Brot) with the ground level (g = 1, E = 0).
Jmax = √(kBT/2Brot) − ½  ≈  7

Answer. The most populated rotational level of CO at 300 K is J = 7, with P7 ≈ 8.9 P0 — nearly nine times the ground level, even though it lies 56Brot higher. The peak sits well above the ground state purely because of the 2J+1 degeneracy, the same mechanism that shifts the peak of the Maxwell speed distribution away from zero (D-014).

Problems
  1. (Warm-up) Two non-degenerate levels are split by ΔE = 0.10 eV. What fraction of systems occupies the upper level at 300 K, and at 1000 K? (kBT = 0.0259 eV at 300 K.)
    SolutionRatio r = pu/p = e−ΔE/kBT. At 300 K: r = e−0.10/0.0259 = e−3.86 = 0.0210, so upper fraction = r/(1+r) = 0.0206 (about 2%). At 1000 K: kBT = 0.0259×(1000/300) = 0.0863 eV, r = e−1.16 = 0.314, fraction = 0.239 (about 24%). Raising T flattens the distribution toward equal occupation.
  2. (Two-level thermodynamics) For a two-level system with energies 0 and ε (both non-degenerate), write Z and ⟨E⟩, evaluate ⟨E⟩/ε when βε = 1, and give the limits T→0 and T→∞.
    SolutionZ = 1 + eβε. Then ⟨E⟩ = ε eβε/(1+eβε) = ε/(eβε+1). At βε = 1: ⟨E⟩ = ε/(e+1) = ε/3.718 = 0.269 ε. Limits: as T→0 (βε→∞), ⟨E⟩→0 (system frozen in ground state); as T→∞ (βε→0), ⟨E⟩→ε/2 (both states equally likely). A two-level system therefore stores at most ε/2 — it saturates, unlike an unbounded spectrum.
  3. (Degeneracy) A system has three levels at energies 0, ε, 2ε with degeneracies 1, 2, 1. Take ε = 0.025 eV and a temperature such that kBT = 0.025 eV (i.e. 290 K). Find Z and the probability of each level. Which single microstate is most probable?
    SolutionHere βε = 1. Z = 1·e0 + 2·e−1 + 1·e−2 = 1 + 0.736 + 0.135 = 1.871. Level probabilities: P(0) = 1/1.871 = 0.534; P(ε) = 0.736/1.871 = 0.393; P(2ε) = 0.135/1.871 = 0.072. The middle level is boosted by its degeneracy, but per state the two states at ε each carry e−1/Z = 0.197, less than the ground state's 0.534. So the single most probable microstate is the non-degenerate ground state — the distinction between "most probable level" and "most probable state" is the whole point.
  4. (Schottky heat capacity) For the two-level system of problem 2, show that the heat capacity is C = kB(βε)² eβε/(eβε+1)². It peaks at βε ≈ 2.40. For ε = 1.0 meV, at what temperature does the peak occur, and what is Cmax?
    SolutionWith ⟨E⟩ = ε/(eβε+1) and C = d⟨E⟩/dT, put x = βε = ε/kBT so dx/dT = −x/T. Then d⟨E⟩/dx = −εex/(ex+1)², giving C = (−εex/(ex+1)²)(−x/T) = kBx² ex/(ex+1)² (using ε/T = xkB). This is the Schottky anomaly: C→0 at both low and high T, with a bump in between. Setting dC/dx = 0 gives numerically x ≈ 2.40. At ε = 1.0 meV: T = ε/(kBx) = 1.0×10−3 eV/(8.617×10−5 eV K−1 × 2.40) = 4.8 K. And Cmax = kB(2.40)²e2.40/(e2.40+1)² = kB(5.76×11.02/144.5) = 0.44 kB per system.
  5. (Capstone — quantum oscillator) A quantum harmonic oscillator has levels En = nω, n = 0, 1, 2, … (measure from the ground state; non-degenerate). (a) Sum the geometric series for Z. (b) Obtain ⟨E⟩ = −∂lnZ/∂β. (c) Evaluate ⟨E⟩ for the CO stretch, ℏω = 0.27 eV, at 300 K, and explain the frozen-out heat capacity of D-020.
    Solution(a) Z = Σn≥0 eω = 1/(1 − eβω) (a geometric series with ratio eβω < 1). (b) lnZ = −ln(1 − eβω), so ⟨E⟩ = −∂lnZ/∂β = ω/(eβω − 1) — the Planck–Einstein result. (Its high-T limit is kBT, recovering equipartition for a 1D oscillator; its low-T limit is ℏωeβω → 0.) (c) βω = 0.27/0.0259 = 10.4, so e10.4 = 3.35×104 and ⟨E⟩ = 0.27/(3.35×104 − 1) = 8.1×10−6 eV per molecule — utterly negligible next to kBT = 0.026 eV. The mode is "frozen": it holds essentially no energy and its contribution to CV is exponentially suppressed (≈ kB(βω)²eβω), which is precisely why the vibrational heat capacity of CO does not switch on until temperatures of order ℏω/kB ≈ 3100 K (D-020).