Conditional expectation as projection
Statement
Let \( (\Omega,\mathcal{F},P) \) be a probability space, let \( \mathcal{G} \subseteq \mathcal{F} \) be a sub-\(\sigma\)-algebra, and let \( X \in L^2(\Omega,\mathcal{F},P) \). Equip \( L^2(\Omega,\mathcal{F},P) \) with the inner product \( \langle U,V\rangle = E[UV] \). Then \( L^2(\Omega,\mathcal{G},P) = \{Y \in L^2(\Omega,\mathcal{F},P) : Y \text{ is } \mathcal{G}\text{-measurable}\} \) is a closed linear subspace of \( L^2(\Omega,\mathcal{F},P) \), and the conditional expectation \( E[X\mid\mathcal{G}] \) coincides (as an a.s.-equivalence class) with the orthogonal projection of \( X \) onto \( L^2(\Omega,\mathcal{G},P) \): it is the unique \( Y_0 \in L^2(\Omega,\mathcal{G},P) \) satisfying \[ \|X - Y_0\| = \min_{Y \in L^2(\Omega,\mathcal{G},P)} \|X-Y\|, \qquad \|U\|:=\sqrt{E[U^2]}, \] equivalently characterised by the orthogonality relation \[ E\big[(X-Y_0)Z\big] = 0 \quad \forall\, Z \in L^2(\Omega,\mathcal{G},P). \]
Why it matters
This theorem converts conditioning — an operation defined measure-theoretically via the Radon–Nikodym theorem and only characterised up to almost-sure equality — into a piece of Euclidean geometry. \( E[X\mid\mathcal{G}] \) is exactly the best mean-square predictor of \( X \) using only the information encoded in \( \mathcal{G} \); "best" here means literally closest in the \( L^2 \) norm, and "using only the information in \( \mathcal{G} \)" means literally restricted to the subspace of \( \mathcal{G} \)-measurable square-integrable functions.
Once conditioning is recognised as a projection, an entire toolkit becomes available for free: linearity, idempotence, self-adjointness, contractivity, the Pythagorean identity (giving the law of total variance), and the tower property (composition of nested projections). This is also the rigorous foundation for treating regression functions, Kalman/Wiener filtering, and Bayesian posterior means as instances of a single geometric operation.
Hypotheses
Proof
Result
Reading. Conditioning on \(\mathcal{G}\) is, up to a.s. equality, the same operation as dropping a perpendicular from \(X\) onto the subspace of "\(\mathcal{G}\)-measurable" vectors inside the Hilbert space \(L^2\): it is the closest \(\mathcal{G}\)-measurable function to \(X\) in mean-square distance, and the "error" \(X - E[X\mid\mathcal{G}]\) is orthogonal (uncorrelated, in the \(E[UV]\) sense) to every \(\mathcal{G}\)-measurable square-integrable function.
Scope. Valid exactly under the hypotheses above: \(X \in L^2\) of a probability space, \(\mathcal{G}\) a genuine sub-\(\sigma\)-algebra. The projection/orthogonality picture is special to \(L^2\); for \(X \in L^1 \setminus L^2\) the conditional expectation still exists (via Radon–Nikodym) but is no longer literally a metric projection, only a norm-decreasing (contractive) operator obtained by continuous extension from \(L^2\) (or \(L^\infty\)) — the geometric "closest point" language does not transfer verbatim.
Corollaries & converses
- Best predictor property. For every \(\mathcal{G}\)-measurable \(Y\in L^2\), \(E[(X-E[X\mid\mathcal{G}])^2] \le E[(X-Y)^2]\), with equality iff \(Y=E[X\mid\mathcal{G}]\) a.s. — this is just restating "nearest point" in prediction language.
- Idempotence and self-adjointness. \(P_\mathcal{G}\) is linear, \(P_\mathcal{G}^2 = P_\mathcal{G}\) (projecting twice does nothing new: \(E[E[X\mid\mathcal{G}]\mid\mathcal{G}]=E[X\mid\mathcal{G}]\)), and \(\langle P_\mathcal{G}X,Y\rangle = \langle X,P_\mathcal{G}Y\rangle\) for all \(X,Y \in L^2\) — general facts about orthogonal projections in Hilbert space, instantiated here.
- Contractivity. \(\|P_\mathcal{G}X\| \le \|X\|\), i.e. \(E[E[X\mid\mathcal{G}]^2] \le E[X^2]\): conditioning never increases mean-square size. Equivalently, by Pythagoras, \(E[X^2] = E[E[X\mid\mathcal{G}]^2] + E[(X-E[X\mid\mathcal{G}])^2]\) — the law of total variance when \(E[X]=0\) is subtracted off appropriately.
- Tower property. If \(\mathcal{G}_1 \subseteq \mathcal{G}_2 \subseteq \mathcal{F}\) then \(\mathcal{H}_{\mathcal{G}_1} \subseteq \mathcal{H}_{\mathcal{G}_2}\), so composing the two nearest-point projections onto nested closed subspaces gives \(P_{\mathcal{G}_1}P_{\mathcal{G}_2}=P_{\mathcal{G}_1}=P_{\mathcal{G}_2}P_{\mathcal{G}_1}\), i.e. \(E[E[X\mid\mathcal{G}_2]\mid\mathcal{G}_1]=E[X\mid\mathcal{G}_1]\) — the tower property is a one-line fact about nested Hilbert subspaces.
- Converse fails in general. Not every orthogonal projection of \(L^2(\Omega,\mathcal{F},P)\) onto a closed subspace is a conditional expectation. Example: fix \(f \in L^2\) with \(\|f\|=1\) and \(f\) not a.s. constant, and let \(Q\) be projection onto the one-dimensional subspace \(\mathrm{span}(f)\), i.e. \(QX = \langle X,f\rangle f\). \(Q\) is a genuine orthogonal projection, but \(\mathrm{span}(f)\) is not of the form \(L^2(\mathcal{G},P)\) for a \(\sigma\)-algebra \(\mathcal{G}\) (it is not even an algebra of functions, e.g. it need not contain \(f^2\) or fix constants unless \(f\equiv1\)), so \(Q\) is not any \(E[\cdot\mid\mathcal{G}]\). The projections that are conditional expectations are exactly those onto subspaces closed under multiplication that contain the constants and separate no more than a \(\sigma\)-algebra's worth of information (equivalently, projections that are also positivity-preserving and unital, \(Q\mathbf{1}=\mathbf{1}\)) — a sharper operator-theoretic characterisation beyond the scope of this statement.
Fails without
- Drop \(X \in L^2\): let \(X\) be standard Cauchy on \((\Omega,\mathcal{F},P)\) with any nontrivial \(\mathcal{G}\). Then \(E[|X|]=\infty\), \(X \notin L^1\), so \(X\) is not even a vector of the Hilbert space \(L^2(\mathcal{F})\); the minimisation problem \(\min_{Y\in L^2(\mathcal{G})}E[(X-Y)^2]\) is not merely hard, it has no well-posed left-hand side, and no orthogonal-projection interpretation of any (even \(L^1\)-defined) conditional expectation of \(X\) is available.
- Drop "\(\mathcal{G}\) is a \(\sigma\)-algebra" (allow a mere algebra \(\mathcal{G}_0\)): on \(\Omega=[0,1]\) with Lebesgue measure, let \(\mathcal{G}_0\) be the algebra of finite unions of dyadic intervals \([k2^{-n},(k+1)2^{-n})\) over all \(n,k\) (an algebra, not a \(\sigma\)-algebra, since it omits limits such as arbitrary Borel sets built as countable unions of shrinking dyadic pieces). Let \(X=\mathbf{1}_{\{\omega \le 1/3\}}\). The natural candidates for a "best \(\mathcal{G}_0\)-measurable approximation" — simple functions constant on ever-finer dyadic partitions — approach \(X\) in \(L^2\)-distance with infimum error \(0\), but no single \(\mathcal{G}_0\)-measurable simple function attains it (attaining error \(0\) would require a set with \(P\)-boundary exactly \(1/3\) to lie in the algebra, which it does not for any finite dyadic level). The infimum is not attained: the minimiser \(Y_0\) required by the theorem does not exist in \(\mathcal{H}_{\mathcal{G}_0}\).
- Drop \(P(\Omega) \lt \infty\): take \((\mathbb{R},\mathcal{B}(\mathbb{R}),\mathrm{Leb})\) (infinite measure) and \(\mathcal{G}=\{\emptyset,\mathbb{R}\}\). Then \(\mathbf{1}_\mathbb{R}\notin L^2(\mathrm{Leb})\), so \(\mathcal{H}_\mathcal{G} \cap L^2 = \{0\}\) only, and the projection of a nonzero \(X\) onto \(\mathcal{H}_\mathcal{G}\) is forced to be \(0\), which need not equal any sensible notion of "the average value of \(X\)" (there generally is none, since Lebesgue measure assigns infinite mass). The correspondence with the classical, Radon–Nikodym-based conditional expectation breaks down.
Common errors
- Minimising \(E[(X-Y)^2]\) over all \(Y \in L^2(\mathcal{F})\) instead of only \(\mathcal{G}\)-measurable \(Y\) — this trivially gives \(Y=X\) itself and destroys the entire content of the theorem; the constraint \(Y \in \mathcal{H}_\mathcal{G}\) is not optional.
- Checking the orthogonality relation \(E[(X-Y_0)Z]=0\) for only one convenient \(Z\) (e.g. \(Z=1\)) and concluding \(Y_0=E[X\mid\mathcal{G}]\); the relation must hold for every \(\mathcal{G}\)-measurable \(Z\in L^2\), not a single test function.
- Conflating "\(X\) uncorrelated with every \(\mathcal{G}\)-measurable function" (orthogonality in \(L^2\)) with "\(X\) independent of \(\mathcal{G}\)" — orthogonality only forces \(E[X\mid\mathcal{G}]\) to be constant (equal to \(E[X]\)), not full independence of \(X\) from the sets in \(\mathcal{G}\).
- Treating \(E[X\mid\mathcal{G}]\) as a specific function of \(\omega\) rather than a \(P\)-a.s. equivalence class; asking "what is \(E[X\mid\mathcal{G}](\omega)\)" for one fixed \(\omega\) on a \(P\)-null set is meaningless — only integrals of \(E[X\mid\mathcal{G}]\) against \(\mathcal{G}\)-measurable test functions are pinned down.
- Forgetting that this Hilbert-space argument is special to \(p=2\): assuming an analogous "orthogonal projection" picture computes \(E[X\mid\mathcal{G}]\) for \(X \in L^1\) or \(L^p, p\ne2\), where there is no inner product and the correct statement is only that conditional expectation is a norm-\(1\) contraction, not a metric projection.
Discussion
Kolmogorov's 1933 axiomatisation defined \(E[X\mid\mathcal{G}]\) purely via the Radon–Nikodym theorem, with no reference to geometry — it is simply the density of the measure \(A\mapsto E[X\mathbf{1}_A]\) with respect to \(P\) restricted to \(\mathcal{G}\). The Hilbert-space reading proved here, tracing back to the projection theorem of the Hilbert-space theory developed by von Neumann and Riesz in the 1930s, is logically a corollary of that definition (via the uniqueness clause of Radon–Nikodym in Step 6), yet it supplies the intuition that makes the whole theory usable: \(E[X\mid\mathcal{G}]\) is "the shadow \(X\) casts" on the space of \(\mathcal{G}\)-measurable functions.
This is also the rigorous population-level statement of ordinary least squares: if \(\mathcal{G}=\sigma(X_1,\dots,X_p)\) is generated by observed covariates, \(E[Y\mid\mathcal{G}]\) is precisely the (generally nonlinear) function of \(X_1,\ldots,X_p\) minimising mean-square prediction error for \(Y\); linear regression is the further restriction of the same minimisation to the strictly smaller closed subspace spanned by \(1,X_1,\dots,X_p\) inside \(\mathcal{H}_\mathcal{G}\), and is itself an orthogonal projection by exactly the same Hilbert projection theorem. The Kalman filter and Wiener filter are dynamic versions of the same statement, projecting onto \(\sigma\)-algebras generated by observations up to time \(t\).
The projection viewpoint also explains, without further computation, why conditional expectation commutes with taking limits of nested \(\sigma\)-algebras in the \(L^2\)-martingale convergence theorem: if \(\mathcal{G}_n \uparrow \mathcal{G}_\infty = \sigma(\bigcup_n \mathcal{G}_n)\), then \(P_{\mathcal{G}_n}X \to P_{\mathcal{G}_\infty}X\) in \(L^2\) is the statement that projections onto an increasing sequence of closed subspaces converge (in norm, on each fixed vector) to the projection onto the closure of their union — a standard Hilbert-space fact that underlies the \(L^2\) martingale convergence theorem and, via truncation arguments, its \(L^1\) and a.s. counterparts (Doob).
Common misconception. It is tempting to think the theorem says conditional expectation "is" a projection in some metaphorical sense. It is not a metaphor: with \(L^2(\mathcal{F})\) literally realised as a Hilbert space via \(\langle U,V\rangle=E[UV]\), \(E[X\mid\mathcal{G}]\) is the literal image of \(X\) under the literal orthogonal-projection operator onto the literal closed subspace \(L^2(\mathcal{G})\) — every word in the Hilbert-space statement is doing exactly the work its analogue in \(\mathbb{R}^n\) would do.
Worked examples
Reading. Conditioning on the coarse partition replaces \(X\) by its average on each block — exactly the projection onto piecewise-constant functions.
Reading. Knowing \(\omega_1\) exactly, the best mean-square guess for \(X=\omega_1+\omega_2\) replaces the unknown \(\omega_2\) by its mean; the residual mean-square error is exactly \(\mathrm{Var}(\omega_2)\), matching the Pythagorean decomposition of the corollaries.
Problems
- Show directly from the orthogonality characterisation that when \(\mathcal{G}=\{\emptyset,\Omega\}\) is trivial, \(E[X\mid\mathcal{G}]=E[X]\) (the constant function).
Solution
\(\mathcal{H}_\mathcal{G}\) consists exactly of the a.s.-constant functions (the only \(\{\emptyset,\Omega\}\)-measurable functions). For constant \(Y_0=c\), orthogonality requires \(E[(X-c)Z]=0\) for every constant \(Z\), i.e. (taking \(Z=1\)) \(E[X]-c=0\), so \(c=E[X]\). This is consistent for every constant \(Z=z\) since \(E[(X-c)z]=z(E[X]-c)=0\). Hence \(E[X\mid\{\emptyset,\Omega\}]=E[X]\). - Prove \(E\big[E[X\mid\mathcal{G}]\,\big|\,\mathcal{G}\big] = E[X\mid\mathcal{G}]\) a.s. using idempotence of orthogonal projections.
Solution
Let \(Y_0=P_\mathcal{G}X=E[X\mid\mathcal{G}]\). Since \(Y_0\in\mathcal{H}_\mathcal{G}\) already, the closest point in \(\mathcal{H}_\mathcal{G}\) to \(Y_0\) is \(Y_0\) itself (distance \(0\) is achieved and is trivially minimal). By the uniqueness clause of the Hilbert projection theorem, \(P_\mathcal{G}Y_0=Y_0\), i.e. \(E[E[X\mid\mathcal{G}]\mid\mathcal{G}]=E[X\mid\mathcal{G}]\) a.s. - Using the Pythagorean identity for the orthogonal decomposition \(X = E[X\mid\mathcal{G}] + (X-E[X\mid\mathcal{G}])\), derive the law of total variance \(\mathrm{Var}(X) = E[\mathrm{Var}(X\mid\mathcal{G})] + \mathrm{Var}(E[X\mid\mathcal{G}])\), where \(\mathrm{Var}(X\mid\mathcal{G}):=E[(X-E[X\mid\mathcal{G}])^2\mid\mathcal{G}]\).
Solution
Apply the theorem to \(X-E[X]\) (still in \(L^2\)) with the same \(\mathcal{G}\): since \(E[X\mid\mathcal{G}]-E[X]=E[X-E[X]\mid\mathcal{G}]\) (linearity of conditional expectation, itself immediate from linearity of \(P_\mathcal{G}\)), the two summands \(E[X\mid\mathcal{G}]-E[X]\) and \(X-E[X\mid\mathcal{G}]\) are orthogonal in \(L^2\) (Step 4 orthogonality applied to \(X-E[X]\), since \(E[X\mid\mathcal{G}]-E[X]\in\mathcal{H}_\mathcal{G}\)). Pythagoras gives \(E[(X-E[X])^2] = E[(E[X\mid\mathcal{G}]-E[X])^2] + E[(X-E[X\mid\mathcal{G}])^2]\), i.e. \(\mathrm{Var}(X) = \mathrm{Var}(E[X\mid\mathcal{G}]) + E[\mathrm{Var}(X\mid\mathcal{G})]\), using that \(E[E[X\mid\mathcal G]]=E[X]\) (take \(A=\Omega\) in the defining property) so the first term is exactly \(\mathrm{Var}(E[X\mid\mathcal{G}])\), and \(E\big[E[(X-E[X\mid\mathcal G])^2\mid\mathcal G]\big]=E[(X-E[X\mid\mathcal G])^2]\) by the tower property (Problem 4). - For \(\mathcal{G}_1\subseteq\mathcal{G}_2\subseteq\mathcal{F}\), prove the tower property \(E[E[X\mid\mathcal{G}_2]\mid\mathcal{G}_1]=E[X\mid\mathcal{G}_1]\) a.s. by composing projections onto nested closed subspaces.
Solution
\(\mathcal{G}_1\subseteq\mathcal{G}_2\) gives \(\mathcal{H}_{\mathcal{G}_1}\subseteq\mathcal{H}_{\mathcal{G}_2}\) as closed subspaces of \(L^2(\mathcal{F})\) (any \(\mathcal{G}_1\)-measurable function is \(\mathcal{G}_2\)-measurable). Let \(Y_2=P_{\mathcal{G}_2}X\). We must show \(P_{\mathcal{G}_1}Y_2 = P_{\mathcal{G}_1}X\). By Step 4 orthogonality applied at level \(\mathcal{G}_2\), \(X-Y_2\) is orthogonal to all of \(\mathcal{H}_{\mathcal{G}_2}\), and in particular to all of \(\mathcal{H}_{\mathcal{G}_1}\subseteq\mathcal{H}_{\mathcal{G}_2}\). So for any \(Z\in\mathcal{H}_{\mathcal{G}_1}\), \(E[(X-Y_2)Z]=0\), i.e. \(E[XZ]=E[Y_2Z]\). This says exactly that \(Y_2\) and \(X\) have the same inner product against every test vector in \(\mathcal{H}_{\mathcal{G}_1}\), so their projections onto \(\mathcal{H}_{\mathcal{G}_1}\) agree: \(P_{\mathcal{G}_1}Y_2=P_{\mathcal{G}_1}X\), i.e. \(E[E[X\mid\mathcal{G}_2]\mid\mathcal{G}_1]=E[X\mid\mathcal{G}_1]\) a.s. - Let \(\Omega=[0,1]\) with Lebesgue measure, \(X(\omega)=\omega^2\), and \(\mathcal{G}=\sigma(Y)\) where \(Y(\omega)=\min(\omega,1-\omega)\) (the \(\mathcal{G}\)-measurable functions are exactly those symmetric under \(\omega\mapsto1-\omega\)). Compute \(E[X\mid\mathcal{G}]\) and verify it via the orthogonality characterisation.
Solution
Guess. By the symmetry pairing \(\omega\leftrightarrow1-\omega\), guess \(Y_0(\omega) = \tfrac12\big[\omega^2+(1-\omega)^2\big] = \omega^2-\omega+\tfrac12\); note \(Y_0\) is invariant under \(\omega\mapsto1-\omega\) (direct check: replacing \(\omega\) by \(1-\omega\) swaps the two squared terms), hence \(\mathcal{G}\)-measurable, being a function of \(\min(\omega,1-\omega)\) alone (any symmetric function of \(\omega\) is a function of \(Y=\min(\omega,1-\omega)\), since \(\{\omega,1-\omega\}\) is recovered from \(Y\)). Verification. \(\mathcal{H}_\mathcal{G}\) consists of \(g(\omega)\) with \(g(\omega)=g(1-\omega)\) a.e. For such \(g\), split \(\int_0^1(X-Y_0)g\,d\omega = \int_0^1\big(\omega^2-\omega^2+\omega-\tfrac12\big)g(\omega)\,d\omega=\int_0^1(\omega-\tfrac12)g(\omega)\,d\omega\). Split the integral at \(\tfrac12\) and substitute \(u=1-\omega\) in the second half: \(\int_0^{1/2}(\omega-\tfrac12)g(\omega)\,d\omega + \int_0^{1/2}\big((1-u)-\tfrac12\big)g(1-u)\,du = \int_0^{1/2}(\omega-\tfrac12)g(\omega)\,d\omega+\int_0^{1/2}(\tfrac12-u)g(u)\,du\) (using \(g(1-u)=g(u)\)) \(= \int_0^{1/2}\big[(\omega-\tfrac12)+(\tfrac12-\omega)\big]g(\omega)\,d\omega = 0\). This holds for every \(g\in\mathcal H_{\mathcal G}\), so by the orthogonality characterisation (Step 4/6 of the proof), \(Y_0=E[X\mid\mathcal{G}]\) a.s.: \(E[\omega^2\mid\mathcal{G}](\omega)=\omega^2-\omega+\tfrac12\).