The Hahn–Banach theorem
Statement
Let \(X\) be a real vector space and let \(p:X\to\mathbb{R}\) be a sublinear functional, i.e. \(p(x+y)\leq p(x)+p(y)\) for all \(x,y\in X\) (subadditivity) and \(p(tx)=t\,p(x)\) for all \(t\geq 0\), \(x\in X\) (positive homogeneity). Let \(M\subseteq X\) be a linear subspace and let \(f:M\to\mathbb{R}\) be linear with \(f(x)\leq p(x)\) for all \(x\in M\). Then there exists a linear functional \(F:X\to\mathbb{R}\) such that \(F|_M=f\) and \(F(x)\leq p(x)\) for all \(x\in X\). Consequently, if \(X\) is a normed space, \(M\) a subspace, and \(f\in M^{*}\) a bounded linear functional, taking \(p(x)=\lVert f\rVert_{M^{*}}\lVert x\rVert\) yields \(F\in X^{*}\) with \(F|_M=f\) and \(\lVert F\rVert_{X^{*}}=\lVert f\rVert_{M^{*}}\): a bounded functional extends to the whole space without increasing its norm.
Why it matters
Hahn–Banach is the existence theorem of functional analysis: it is what guarantees that a normed space, however abstract or infinite-dimensional, has "enough" continuous linear functionals to see every vector, separate disjoint convex sets, and identify a space with a natural subspace of its bidual. Without it there is no guarantee that \(X^{*}\) is not embarrassingly small — indeed there are exotic topological vector spaces where the dual is trivial, and the theorem's hypotheses (sublinearity, local convexity in the topological version) are exactly what rules that pathology out for normed spaces.
Its extension-without-loss-of-control mechanism recurs everywhere: separating hyperplane theorems in convex analysis, existence of Banach limits and invariant means, weak-* compactness arguments, the identification of duals of quotient and subspace, and the very definition of the second dual embedding \(X\hookrightarrow X^{**}\) all rest on it.
Hypotheses
Proof
Result
Reading. A linear functional that is "small" (dominated by a sublinear gauge \(p\)) on a subspace can always be pushed out to the whole space while remaining small everywhere. In the normed-space case, "small" means bounded by a multiple of the norm, and the theorem says the operator norm of the extension can be taken exactly equal to the operator norm on the subspace — no growth is forced.
Scope. Proved here for real scalars; the complex version (Bohnenblust–Sobczyk) follows by extending \(\operatorname{Re} f\) as a real-linear functional dominated by the seminorm \(p\) and recovering \(F(x) = g(x) - i\,g(ix)\). The theorem needs no completeness, separability, or local structure beyond sublinearity of \(p\) — it holds in an arbitrary real (or complex) vector space, with no topology at all; continuity/boundedness only enters when one specialises \(p\) to a norm multiple to get the normed-space corollary used in practice.
Corollaries & converses
- Norm-attaining functional at a point. For any \(x_0\neq0\) in a normed space \(X\), there exists \(F\in X^{*}\) with \(\lVert F\rVert=1\) and \(F(x_0)=\lVert x_0\rVert\): apply the theorem to \(f(tx_0)=t\lVert x_0\rVert\) on \(M=\mathbb{R}x_0\), dominated by \(p(x)=\lVert x\rVert\).
- The dual separates points. If \(x\neq y\) in \(X\) then some \(F\in X^{*}\) has \(F(x)\neq F(y)\) — immediate from the previous corollary applied to \(x-y\).
- Canonical embedding \(X\hookrightarrow X^{**}\) is isometric. \(\hat{x}(F):=F(x)\) satisfies \(\lVert\hat x\rVert_{X^{**}}=\lVert x\rVert_X\); \(\leq\) is Cauchy–Schwarz-type and always true, \(\geq\) needs the norm-attaining corollary above.
- Distance to a subspace via functionals. If \(M\) is a closed subspace and \(x_0\notin M\), there is \(F\in X^{*}\) with \(F|_M=0\), \(\lVert F\rVert=1\), \(F(x_0)=\operatorname{dist}(x_0,M)\) (geometric Hahn–Banach / separation form).
- Converse fails to be meaningful as stated — the theorem is an existence statement, not an equivalence, so there is no converse to test; what can be asked is uniqueness of the extension, and that is false in general (Worked Example 1 below exhibits a continuum of distinct extensions), becoming unique exactly when \(p\) is strictly convex on the relevant directions (e.g. Hilbert space norms via Riesz representation).
Fails without
- Without subadditivity of \(p\): take \(X=\mathbb{R}^2\), \(M=\{(t,t)\}\), \(f(t,t)=2t\), and \(p(x,y)=\max(x,y,0)\) restricted to be homogeneous but engineered to fail the triangle inequality on the relevant pair — concretely \(p(x,y)=(x+y)^2\) if \(x+y \geq 0\) and \(0\) otherwise fails subadditivity (\(p(1,0)+p(-1,1)=1+0=1 \lt p(0,1)=1\) is not violated, but \(p(1,-0.5)+p(-0.5,1)=0.25+0.25=0.5 \lt p(0.5,0.5)=1\) is), and the one-step lemma's key inequality in Proof Step 5 — which needed exactly \(p(y_1+y_2)\leq p(y_1)+p(y_2)\) — collapses, so \(\alpha\) need not exist and no consistent extension across even one extra dimension can be guaranteed.
- Without \(f\) already dominated on \(M\): \(X=\mathbb{R}\), \(M=\{0\}\subsetneq X\), \(f(0)=0\) trivially, but if one instead insists \(f\) is a functional on \(M=\mathbb{R}\) itself with \(f(t)=t\) and asks it to be dominated by \(p(t)=-|t|\) (not sublinear — homogeneous but not subadditive, and in fact \(f(1)=1 \gt p(1)=-1\) already), the extension problem is vacuous/impossible before any subspace argument begins — domination must hold at the start.
- Without the axiom of choice, in general \(X\): in the Feferman–Lévy-type or Pincus-type models of ZF where the Boolean Prime Ideal theorem fails, there exist normed spaces (typically \(\ell^\infty\)-type constructions) and bounded functionals on subspaces with no norm-preserving extension to the whole space; the finite- and countable-dimensional cases survive because they need only finitely/countably many applications of the one-step lemma, which is choice-free, but genuinely uncountable-dimensional \(X\) needs Zorn's Lemma as used in Steps 1–3.
Common errors
- Believing the extension \(F\) is unique — it generically is not (see Worked Example 1); students often report "the" Hahn–Banach extension as if the theorem specified one.
- Applying the theorem with \(p\) taken to be a norm when \(p\) is only required to be sublinear — this under-uses the theorem (many gauge functions of interest, e.g. support functions of convex sets, are sublinear but not norms, being neither symmetric nor zero only at the origin).
- Forgetting the domination direction: the requirement is \(f(x)\leq p(x)\), a one-sided bound, not \(|f(x)|\leq p(x)\); conflating the two is harmless when \(p\) is even (\(p(-x)=p(x)\), as for a norm) but wrong in general, and the proof genuinely only uses the one-sided version.
- Trying to prove the complex case by directly bounding \(F\) with \(F\leq p\) as in the real case — \(\leq\) is meaningless for complex values; the correct route is via \(\operatorname{Re}F\) and the identity \(F(x)=g(x)-ig(ix)\).
- Treating this as a constructive theorem — for infinite-dimensional \(X\) the proof is a pure existence argument via Zorn's Lemma; no formula for \(F\) is produced (compare Worked Example 2, where the extension — a Banach limit — provably cannot be written down by any explicit rule).
Discussion
Hahn–Banach was isolated independently by Hans Hahn (1927) and Stefan Banach (1929) as the abstract engine behind extension results that had appeared piecemeal — Riesz's work on moment problems, Helly's earlier finite/countable extension results. Its real achievement is turning a geometric idea (you can always find a supporting hyperplane bounding a convex "cap") into a purely algebraic one-step lemma that Zorn's Lemma can then iterate to infinity without ever needing to see the whole space at once.
The analytic form proved here (domination by a sublinear functional) and the geometric form (a convex set with nonempty interior, disjoint from a subspace or convex set, can be separated by a hyperplane) are formally equivalent: the gauge (Minkowski functional) of a convex absorbing set is sublinear, and conversely every sublinear functional is the gauge of the convex set \(\{x : p(x)\leq 1\}\). This is why the theorem sits simultaneously in functional analysis and convex geometry, and why it underlies both duality theory in Banach spaces and separating-hyperplane arguments in optimisation (Lagrangian duality, supporting hyperplane theorems for convex programs).
A recurring subtlety: the theorem gives no control over the extension beyond the bound \(F\leq p\); in particular it does not preserve any extra structure \(f\) might have (positivity, multiplicativity, continuity in a topology finer than the norm) unless that structure is separately built into \(p\) or re-derived. The Krein–Milman and Krein–Rutman theorems, which do produce positivity-preserving or extreme-point-respecting extensions, require substantially more (compactness, cones) precisely because plain Hahn–Banach does not supply it for free.
Common misconception: that Hahn–Banach requires completeness of \(X\) (i.e. that it is really a Banach-space theorem). It is not — the algebraic form needs only a vector space and a sublinear functional; completeness is irrelevant to the extension mechanism and only becomes relevant in adjacent theorems (open mapping, closed graph, uniform boundedness) that genuinely need Baire category.
Worked examples
Reading. There is an entire one-parameter family of norm-preserving extensions, not a single one: Hahn–Banach guarantees existence, not uniqueness, and the sup-norm's non-strict convexity is exactly what allows this multiplicity.
Reading. This is a genuinely non-constructive existence result — no explicit formula for \(F\) on a general bounded sequence like \((0,1,0,1,\dots)\) can be written down (it assigns such a sequence some value in \([0,1]\), but which one depends on the Zorn's Lemma choices made), yet Hahn–Banach guarantees it exists and is bounded by \(\limsup\).
Problems
- Show that \(p(x) = \lVert x\rVert\) for a norm \(\lVert\cdot\rVert\) on a vector space \(X\) is sublinear, and identify exactly which part of the sublinearity definition uses the triangle inequality versus absolute homogeneity of the norm.
Solution
Subadditivity: \(p(x+y)=\lVert x+y\rVert \leq \lVert x\rVert+\lVert y\rVert = p(x)+p(y)\) is precisely the triangle inequality axiom of the norm. Positive homogeneity: for \(t\geq0\), \(p(tx)=\lVert tx\rVert = |t|\,\lVert x\rVert = t\lVert x\rVert = t\,p(x)\) (using \(|t|=t\) since \(t\geq0\)), which uses absolute homogeneity of the norm, \(\lVert tx\rVert=|t|\lVert x\rVert\), restricted to \(t\geq0\) — note sublinearity does not require the full homogeneity \(p(tx)=t p(x)\) for negative \(t\), so a norm is a special (symmetric) sublinear functional but sublinear functionals need not be symmetric, e.g. \(p(x,y)=\max(x,y,0)\) on \(\mathbb{R}^2\) is sublinear but not a norm. - In \(X=\mathbb{R}^2\) with \(M=\{(t,0):t\in\mathbb{R}\}\), \(f(t,0)=3t\), and \(p(x,y)=|x|+2|y|\), find one linear extension \(F(x,y)=ax+by\) with \(F\leq p\) on all of \(X\), and verify domination explicitly.
Solution
Need \(F(t,0)=at=3t\) so \(a=3\). Need \(ax+by\leq |x|+2|y|\) for all \((x,y)\), i.e. (testing signs) \(|a|\leq 1\) is required by setting \(y=0\): but \(a=3\not\leq1\). So closer inspection: setting \(y=0\), domination requires \(3x\leq|x|\) for all \(x\), which fails for \(x\gt0\) already (\(3x \gt x = |x|\) once \(x\gt0\)). Hence \(f\) is not dominated by \(p\) on \(M\) to begin with (\(f(1,0)=3 \gt p(1,0)=1\)), so the Hypotheses of the theorem fail and no such extension is asserted to exist — indeed none does, since any extension restricts to \(f\) on \(M\) and would inherit the violation. This illustrates the necessity of checking domination on \(M\) before invoking the theorem. - Using the corollary "for \(x_0\neq0\) there is \(F\in X^*\) with \(\lVert F\rVert=1,\ F(x_0)=\lVert x_0\rVert\)", prove that the canonical map \(\hat{\,\cdot\,}:X\to X^{**}\), \(\hat x(F)=F(x)\), is isometric: \(\lVert \hat x\rVert_{X^{**}} = \lVert x\rVert_X\).
Solution
(\(\leq\)) For any \(F\in X^*\) with \(\lVert F\rVert\leq1\), \(|\hat x(F)| = |F(x)| \leq \lVert F\rVert\,\lVert x\rVert \leq \lVert x\rVert\), so \(\lVert \hat x\rVert_{X^{**}} = \sup_{\lVert F\rVert\leq1}|F(x)| \leq \lVert x\rVert\). (\(\geq\)) If \(x=0\) both sides are \(0\); if \(x\neq0\), the corollary gives \(F_0\in X^*\) with \(\lVert F_0\rVert=1\) and \(F_0(x)=\lVert x\rVert\), so \(\lVert \hat x\rVert_{X^{**}} \geq |\hat x(F_0)| = \lVert x\rVert\). Combining, \(\lVert\hat x\rVert_{X^{**}}=\lVert x\rVert_X\). - Construct an explicit sublinear \(p\) on \(\mathbb{R}^2\) that is not a norm (fails symmetry or fails to vanish only at \(0\)), and verify it is still sublinear.
Solution
Take \(p(x,y) = \max(x,y,0)\) — the support function of the triangle with vertices \((0,0),(1,0),(0,1)\), or equivalently the gauge of a convex set that is not symmetric about the origin. Homogeneity: \(p(t x,ty)=\max(tx,ty,0) = t\max(x,y,0)=t\,p(x,y)\) for \(t\geq0\), directly. Subadditivity: \(\max(x_1+x_2,y_1+y_2,0) \leq \max(x_1,y_1,0)+\max(x_2,y_2,0)\); check by cases on which coordinate/zero attains each outer max — e.g. if the LHS max is attained by \(x_1+x_2\), then \(x_1+x_2 \leq \max(x_1,y_1,0)+\max(x_2,y_2,0)\) since \(x_1\leq\max(x_1,y_1,0)\) and \(x_2\leq\max(x_2,y_2,0)\) always. Not a norm: \(p(-1,-1)=0\neq p(1,1)=1\), so \(p(-v)\neq p(v)\), violating the symmetry norms require; also \(p(-1,0)=0\) while \((-1,0)\neq(0,0)\), so \(p\) vanishes off the origin — definiteness fails too. - Let \(X\) be a normed space, \(M\) a closed proper subspace, and \(x_0\in X\setminus M\) with \(d:=\operatorname{dist}(x_0,M) \gt 0\). Using Hahn–Banach, construct \(F\in X^*\) with \(F|_M=0\), \(F(x_0)=d\), and \(\lVert F\rVert=1\).
Solution
Let \(N = M\oplus\mathbb{R}x_0\) and define \(g:N\to\mathbb{R}\) by \(g(m+tx_0) = td\) (well-defined since the sum is direct, as \(x_0\notin M\)). Check \(g\) is bounded on \(N\) with \(\lVert g\rVert_{N^*}\leq1\): for \(t\neq0\), \(|g(m+tx_0)| = |t|d = |t|\operatorname{dist}(x_0,M) \leq |t|\,\lVert x_0 - (-m/t)\rVert = \lVert tx_0+m\rVert\) for every \(m\in M\) (using \(-m/t\in M\) and the definition of \(\operatorname{dist}\) as an infimum over \(M\)); taking the infimum over the representation gives \(|g(m+tx_0)|\leq \lVert m+tx_0\rVert\), and for \(t=0\), \(g=0\) trivially. So \(\lVert g\rVert_{N^*}\leq1\); also \(\lVert g\rVert_{N^*}\geq1\) by approximating \(d=\inf_{m\in M}\lVert x_0-m\rVert\) with a near-minimising \(m_k\), giving \(|g(x_0-m_k)|/\lVert x_0-m_k\rVert = d/\lVert x_0-m_k\rVert \to 1\). So \(\lVert g\rVert_{N^*}=1\). Apply the normed-space corollary of Hahn–Banach (Proof Step 8) to extend \(g\) to \(F\in X^*\) with \(\lVert F\rVert_{X^*}=\lVert g\rVert_{N^*}=1\); then \(F|_M = g|_M = 0\) and \(F(x_0)=g(x_0)=d\), as required.