The Central Limit Theorem
Statement
Let X1, X2, …, XN be independent, identically distributed random variables with finite mean μ = E[Xi] and finite, nonzero variance σ2 = Var(Xi). Form the sum SN = ∑i=1N Xi and standardise it: ZN = (SN − Nμ)/(σ√N). Then as N → ∞, ZN converges in distribution to a standard Gaussian: ZN →d 𝒩(0,1), meaning P(ZN ≤ z) → Φ(z) = ∫−∞z (1/√(2π)) e−u2/2 du for every z. The limiting shape is Gaussian regardless of the shape of the parent distribution, provided only that its variance is finite.
Why it matters
This is the reason the bell curve is everywhere. Any measurable quantity that is the accumulation of many small, independent contributions — the total charge from countless electrons, the pressure from myriad molecular impacts, the error left after many independent instrument imperfections — is driven toward a Gaussian by this theorem, and the individual contributions’ own distributions are washed out. It is why a single functional form, with just two parameters (mean and variance), describes the fluctuations of astonishingly disparate systems.
It also underwrites the entire practice of error analysis. The standard-error result fixes the width of a sample mean’s distribution; the central limit theorem fixes its shape, and only that shape lets us convert a standard error into a confidence interval (“68% within 1σ, 95% within 2σ”). Every χ2 test, every Gaussian likelihood, every “nσ discovery” in physics inherits its legitimacy from here.
Assumptions
Derivation
Result
Reading. Add up many independent finite-variance contributions, subtract the mean, and rescale by σ√N; whatever the individual variables looked like, the result is a standard bell curve. All the parent distribution’s idiosyncratic features — skewness, kurtosis, discreteness — are relegated to higher-order corrections that die as powers of 1/√N. Only the first two moments, μ and σ2, survive; everything else is forgotten.
Units check. SN and Nμ both carry the units of X (say joules); their difference does too. The denominator σ√N carries the units of σ, also joules (√N is a dimensionless count). Hence ZN is dimensionless, as it must be to be a standard normal deviate. The characteristic-function argument t is therefore dimensionless, so e−t2/2 is well-formed.
Limiting cases
- N = 1: Z1 = (X−μ)/σ is just the standardised parent — no Gaussianity yet; the theorem is an N→∞ statement.
- Parent already Gaussian: ZN is exactly 𝒩(0,1) for every N — the Gaussian is the fixed point of the sum-and-rescale map (a stable law with index 2).
- Symmetric parent: odd moments vanish, so the leading correction is the O(1/N) kurtosis term, not the O(1/√N) skewness term — convergence is noticeably faster.
- σ → 0 (degenerate): the variables are constants; ZN is undefined (0/0). The theorem needs a genuinely fluctuating parent.
Breaks when
- Infinite variance / heavy tails. For a Cauchy (Lorentzian) parent, σ2 does not exist and Step 5 has no quadratic term; instead φY(s) = e−|s|, giving [φ(t/N)]N = e−|t| — the sample mean of Cauchy variables is again Cauchy, never Gaussian. More generally a parent with tail index α < 2 needs normalisation N1/α and converges to an α-stable law with power-law tails, not a bell curve.
- Strong dependence. If the Xi are correlated with slowly decaying (long-range) correlations, Step 4’s factorisation fails, the variance of the sum grows faster than N, and the correct rescaling and limit change — e.g. the increments of fractional Brownian motion converge to a non-standard Gaussian process with the wrong exponent, or to non-Gaussian Hermite processes at criticality.
- A dominant term (Lindeberg violated). If one summand carries a finite fraction of the total variance no matter how large N is (e.g. variances σi2 ∝ 2i), that term never becomes negligible; its own non-Gaussian shape survives in the limit and the sum does not Gaussianise. Averaging many terms only works when none of them individually matters.
- Finite N in the tails. Even with all hypotheses satisfied, the approximation is excellent near the centre but poor many standard deviations out: rare-event (large-deviation) probabilities are governed by the parent’s tail, not the Gaussian, and can be wrong by orders of magnitude for the same N that makes the bulk look perfectly normal.
Failure modes
- Thinking the sum itself becomes Gaussian. SN has mean Nμ and variance Nσ2 that both grow with N — it does not converge to anything. Only the standardised variable ZN has a limit. The theorem is about shape after centring and rescaling.
- Applying it to a single measurement or tiny N. The CLT is asymptotic; for N=1 the distribution is just the parent. “Assume it’s Gaussian because CLT” with N=3 and a skewed parent is unjustified.
- Ignoring the variance condition. Invoking the CLT for power-law / heavy-tailed data (financial returns, network degrees, city sizes) where σ2 is effectively infinite — the sample “mean” then fails to stabilise and confidence intervals are meaningless.
- Forgetting independence. Treating N correlated readings as if they were independent inflates the effective sample size and mis-applies the σ√N scaling; the sum can be far from Gaussian.
- Trusting the tails. Using the Gaussian to estimate a 6σ or 8σ probability for moderate N — the central limit theorem controls the bulk, not the extreme tail, which converges much more slowly (Cramér / large-deviation regime).
- Confusing the CLT with the law of large numbers. The LLN says the sample mean converges to μ (a statement about location); the CLT says the fluctuation about μ, magnified by √N, is Gaussian (a statement about shape). They answer different questions.
Discussion
The proof reveals precisely why the Gaussian is universal: the characteristic function turns summing variables into multiplying functions, and taking a logarithm turns that into adding functions. The theorem is then the statement that when you add N copies of a function that starts as 1 − (something)/2N, only the quadratic term survives the N→∞ limit — the linear term was removed by centring, and every higher term is suppressed by extra powers of 1/√N. The Gaussian is universal because it is the unique fixed point of “add two independent copies and rescale” among finite-variance laws; the sum-and-rescale operation is a flow on the space of distributions, and e−t2/2 is where everything with finite variance ends up.
This fixed-point language is not a metaphor — it is the seed of the renormalisation group. In statistical physics the CLT is the trivial (Gaussian) fixed point of block-spin summation; a critical system is exactly one where the CLT breaks because correlations are long-ranged and fluctuations do not add in quadrature, giving non-Gaussian scaling exponents. The stable laws (Cauchy, Lévy) are the other fixed points, reached when variance is infinite. So the theorem is the boring, universal centre of a much richer landscape: it tells you what happens when nothing interesting (no criticality, no heavy tails) is going on, which is most of the time, which is why the bell curve is everywhere.
There is a precise hierarchy of corrections beyond the leading Gaussian, made explicit by the Edgeworth expansion. Keeping the cubic term in Step 5 (the skewness γ1 = E[Y3]) gives lnφZN(t) = −t2/2 + (γ1/6)(it)3/√N + …, so the leading correction to the density is O(1/√N) and proportional to the parent’s skewness; the next, kurtosis, correction is O(1/N). This is why symmetric parents converge faster (no 1/√N term) and why the Berry–Esseen theorem bounds the distance to Gaussian by Cρ/(σ3√N) with ρ=E|X−μ|3 — a uniform, non-asymptotic control on the rate that the bare theorem does not give.
Common misconceptions. (i) “The CLT says everything is Gaussian.” No — it says sums of many independent finite-variance pieces are, after standardising; a single skewed or heavy-tailed quantity need not be. (ii) “Larger N makes the data Gaussian.” The raw data keep their own distribution forever; it is the standardised sum that Gaussianises. (iii) “CLT and the law of large numbers are the same.” One is about where the mean goes, the other about the shape of its fluctuations. (iv) “It works for any distribution.” Finite variance is essential; Cauchy is the standard counterexample where it fails outright.
Worked examples
Example 1 — Sum of uniforms as a Gaussian generator. A classic quick-and-dirty way to make an approximately standard-normal random number is to add twelve independent U(0,1) variables and subtract 6. Justify this and estimate P(S12 > 7), where S12 = ∑i=112 Ui.
Reading. Twelve flat distributions, each utterly un-bell-shaped, sum to something almost indistinguishable from a Gaussian — the CLT in action at modest N. The −6 centres it and the variance happens to be exactly 1, which is why this recipe was popular before better generators existed. (Its flaw: the sum is bounded to [0,12], so it cannot produce deviates beyond ±6σ — a tail failure exactly as the “Breaks when” box warns.)
Example 2 — Coin flips and the normal approximation to the binomial (de Moivre–Laplace). A fair coin is flipped N = 100 times. Let H be the number of heads. Estimate P(H ≥ 60) using the CLT.
Reading. The exact binomial answer is 0.0284, so the CLT (with continuity correction) is right to better than 1% relative — excellent for N=100 and symmetric p=1/2. Getting 60 or more heads is a modest 1.9σ upward fluctuation. Without the continuity correction one would use Z=2.0 and get 0.0228, noticeably worse — the correction matters at these N.
Problems
- A machine fills bags with a mean mass μ = 500 g and standard deviation σ = 20 g per bag, bags independent. A pallet holds N = 100 bags. Using the CLT, estimate the probability that the total pallet mass exceeds 50.4 kg.
Solution
Total S100 has mean Nμ = 50000 g and variance Nσ2 = 100×400 = 40000 g², so SD = √40000 = 200 g. Standardise: Z = (50400 − 50000)/200 = 2.0. So P(S>50400) = P(Z>2) = 1 − Φ(2) = 1 − 0.9772 = 0.0228 (about 2.3%). Note the SD of the total grows only as √N (200 g) while the mean grows as N (50 kg) — the sum is relatively more predictable for larger N. - Explain, using characteristic functions, why the sum of two independent Cauchy variables (each with φ(t)=e−|t|) rescaled by 1/2 is again Cauchy, and why this makes the CLT fail. Which step of the derivation is invalid for the Cauchy?
Solution
For independent X1,X2, φX1+X2(t) = φ(t)² = e−2|t|. The sample mean is M = (X1+X2)/2, whose characteristic function is φX1+X2(t/2) = e−2|t/2| = e−|t| — identical to a single Cauchy. So averaging does not narrow the distribution at all; the “standardised” mean stays Cauchy for every N and never approaches a Gaussian. The failure is at Step 5: the Taylor expansion φ(s)=1−s²/2+… requires a finite second moment; the Cauchy has E[X2]=∞, and indeed e−|s| = 1 − |s| + … is not differentiable at s=0 (a |s|, not s², cusp), so there is no quadratic-order term to inherit. - The lifetime of a component is exponentially distributed with mean τ = 5 hours (so μ=5, σ=5 h). A system replaces components one after another. Estimate the probability that the total time to get through N = 50 components exceeds 300 hours.
Solution
For an exponential, μ=τ=5 h and σ=τ=5 h. Total time T = ∑150 has mean Nμ = 250 h and variance Nσ2 = 50×25 = 1250 h², so SD = √1250 ≈ 35.36 h. Standardise: Z = (300 − 250)/35.36 = 50/35.36 ≈ 1.414. So P(T>300) ≈ P(Z>1.41) = 1 − Φ(1.41) ≈ 1 − 0.9207 = 0.079 (about 8%). (The exact answer via the Gamma/Erlang distribution is ≈ 0.083; the exponential is right-skewed, so at N=50 a small O(1/√N) skewness error remains.) - A parent distribution has skewness γ1 = 2. Roughly how large must N be for the leading Edgeworth skewness correction, of order γ1/√N, to fall below 0.1? What does this say about the rate of convergence compared with a symmetric parent?
Solution
Require γ1/√N < 0.1, i.e. √N > γ1/0.1 = 2/0.1 = 20, so N > 400. You need a few hundred samples for a moderately skewed parent, versus far fewer for a symmetric one where the 1/√N term vanishes identically and the leading correction is the smaller O(1/N) kurtosis term. Skewness is the slowest-dying defect, which is why asymmetric parents (exponential, Poisson at small mean) approach Gaussian noticeably more slowly than symmetric ones (uniform, binomial at p=1/2). - Show that if X is itself Gaussian, then ZN is exactly 𝒩(0,1) for every N, using characteristic functions. What special property of the Gaussian does this express?
Solution
A single standardised Gaussian summand has φY(s) = e−s²/2 exactly (no higher-order remainder). Then from Step 4, φZN(t) = [φY(t/√N)]N = [e−(t/√N)²/2]N = [e−t²/(2N)]N = e−t²/2 for every N, including N=1. So ZN is precisely standard normal, not merely asymptotically. This expresses that the Gaussian is a stable law (index 2): a normalised sum of independent Gaussians is Gaussian. It is the attracting fixed point of the sum-and-rescale flow, which is exactly why all finite-variance distributions converge to it rather than to something else.