The Jordan normal form
Statement
Let \( V \) be a nonzero finite-dimensional vector space over \( \mathbb{C} \) (more generally, over any field \( F \) such that the characteristic polynomial of \( T \) splits into linear factors over \( F \)), and let \( T : V \to V \) be a linear operator. Then there exists a basis of \( V \) in which the matrix of \( T \) is a direct sum of Jordan blocks \[ J_k(\lambda) \;=\; \begin{pmatrix} \lambda & 1 & & \\ & \lambda & \ddots & \\ & & \ddots & 1 \\ & & & \lambda \end{pmatrix} \in M_k(\mathbb{C}), \] i.e. \( [T] = J_{k_1}(\lambda_1) \oplus J_{k_2}(\lambda_2) \oplus \cdots \oplus J_{k_r}(\lambda_r) \), where the \( \lambda_j \) are eigenvalues of \( T \) (repetitions allowed). Moreover the multiset of blocks \( \{ (k_j, \lambda_j) \} \) is uniquely determined by \( T \): two operators (equivalently, two matrices in \( M_n(\mathbb{C}) \)) are similar if and only if they have the same Jordan blocks up to reordering.
Why it matters
The Jordan normal form is the complete classification of linear operators on a finite-dimensional complex vector space up to similarity. Diagonalisation fails precisely when eigenvectors run out; the Jordan form quantifies that failure exactly, replacing "diagonalisable" with the universally valid "diagonalisable plus nilpotent, in commuting pieces". Every similarity-invariant question about a single operator — minimal polynomial, ranks of powers of \( T - \lambda \), matrix functions \( f(T) \), solutions of \( \dot{x} = Ax \) — can be answered by inspecting the block data.
It is also the first structure theorem most students meet: the statement "a finitely generated torsion module over a PID is a direct sum of cyclic modules" specialises, for the ring \( \mathbb{C}[x] \) acting on \( V \) via \( T \), to exactly this theorem. Understanding the Jordan form well is a rehearsal for module theory, representation theory, and the study of unipotent and nilpotent elements in Lie theory.
Hypotheses
Proof
Write \( n = \dim V \geq 1 \). The proof has three movements: (I) split \( V \) into generalized eigenspaces (steps 1–6); (II) prove the structure theorem for nilpotent operators by induction (steps 7–11); (III) assemble, and prove uniqueness (steps 12–13).
Result
Reading. Choose coordinates wisely and any complex square matrix becomes almost diagonal: eigenvalues down the diagonal, and a few unavoidable \( 1 \)’s just above it recording exactly how far the operator is from diagonalisable. The sizes and eigenvalues of the blocks are a complete fingerprint: two matrices are similar precisely when their fingerprints agree.
Scope. Applies to every linear operator on a nonzero finite-dimensional vector space over \( \mathbb{C} \), and over any field \( F \) provided the characteristic polynomial splits over \( F \) (e.g. always over an algebraically closed field; over \( \mathbb{R} \) only when all eigenvalues are real). It fails in infinite dimensions and gives no canonical form for families of non-commuting operators. Over a non-splitting field the correct substitute is the rational canonical form.
Corollaries & converses
- Minimal polynomial read off. \( \mu_T(x) = \prod_i (x - \lambda_i)^{s_i} \) where \( s_i \) is the size of the largest block with eigenvalue \( \lambda_i \). In particular \( T \) is diagonalisable \( \iff \) \( \mu_T \) has simple roots \( \iff \) all blocks have size \( 1 \).
- Jordan–Chevalley decomposition. \( T = D + N \) with \( D \) diagonalisable, \( N \) nilpotent, \( DN = ND \); moreover \( D \) and \( N \) are polynomials in \( T \) (take \( D = \sum_i \lambda_i E_i \) with the projections \( E_i = q_i(T)p_i(T) \) of step 4), and this decomposition is unique.
- Similarity classification. The similarity classes in \( M_n(\mathbb{C}) \) are in bijection with functions assigning to each eigenvalue a partition (its block sizes). E.g. nilpotent \( n \times n \) classes \( \leftrightarrow \) partitions of \( n \).
- Every complex matrix is similar to its transpose (conjugate each block by the reversal permutation; see Problem 3).
- Matrix functions and ODEs. \( e^{tJ_k(\lambda)} = e^{t\lambda} \sum_{a=0}^{k-1} \frac{t^a}{a!} J_k(0)^a \), giving the \( t^a e^{\lambda t} \) solution basis of \( \dot{x} = Ax \).
- Converse direction. "Same Jordan form \( \Rightarrow \) similar" holds trivially. But weaker invariants do not classify: \( J_2(0) \) and the zero \( 2 \times 2 \) matrix share characteristic polynomial \( x^2 \), trace, determinant and eigenvalues, yet are not similar (different ranks). Even characteristic and minimal polynomial together fail from size \( 4 \): \( J_2(0) \oplus J_2(0) \) versus \( J_2(0) \oplus J_1(0) \oplus J_1(0) \) share \( \chi = x^4, \mu = x^2 \) but differ in rank.
Fails without
- Drop algebraic closure (splitting). \( A = \begin{pmatrix} 0 & -1 \\ 1 & 0 \end{pmatrix} \in M_2(\mathbb{R}) \): \( \chi_A = x^2 + 1 \) has no real root, so \( A \) has no real eigenvector and is similar over \( \mathbb{R} \) to no triangular, hence no Jordan, matrix. (Over \( \mathbb{C} \) it becomes \( \mathrm{diag}(i, -i) \).) Similarly \( \begin{pmatrix} 0 & 2 \\ 1 & 0 \end{pmatrix} \) over \( \mathbb{Q} \) fails because \( x^2 - 2 \) does not split in \( \mathbb{Q}[x] \).
- Drop finite-dimensionality. On \( V = \mathbb{C}[x] \), the operator \( M : f \mapsto xf \) has no eigenvalues (degree count), so \( V \) has no generalized eigenspace decomposition whatsoever. Worse, the differentiation operator \( D : f \mapsto f' \) on \( \mathbb{C}[x] \) satisfies \( \bigcup_k \ker D^k = V \) yet admits no basis of Jordan chains for a direct-sum-of-blocks matrix: \( D \) is surjective, while a direct sum of finite nilpotent blocks \( J_k(0) \) is never surjective on a nonzero space unless some chain is infinite — the block-decomposition dichotomy of the finite theorem simply breaks.
- Drop "vector space" for "module". Over \( \mathbb{Z} \), multiplication by \( 2 \) on \( \mathbb{Z}^2 \) is "diagonal", but the classification-by-blocks machinery needs division: \( \begin{pmatrix} 0 & 2 \\ 0 & 0 \end{pmatrix} \) and \( \begin{pmatrix} 0 & 1 \\ 0 & 0 \end{pmatrix} \) are conjugate over \( \mathbb{Q} \) but not over \( \mathbb{Z} \) (any \( P \in GL_2(\mathbb{Z}) \) conjugating one to the other would force \( 2 \mid \det P \cdot \text{unit} \), impossible); canonical forms over rings require the Smith normal form theory instead.
Common errors
- Geometric multiplicity tells you everything. It only counts the number of blocks for \( \lambda \) (\( = \dim\ker(A - \lambda I) \)), not their sizes. For sizes you need the ranks \( r_p(\lambda) \) of powers (step 13). A \( 4\times 4 \) nilpotent with \( \dim\ker N = 2 \) could be \( J_3 \oplus J_1 \) or \( J_2 \oplus J_2 \).
- Solving the chain top-down with a random eigenvector. Picking any eigenvector \( v \) and then trying to solve \( (A - \lambda I)w = v \) may be inconsistent: \( v \) must be chosen inside \( \operatorname{im}(A - \lambda I) \). Build chains from the top (a vector of maximal height) downward, or solve the system for \( v \) and \( w \) together.
- Conflating exponents. Algebraic multiplicity \( m_i \) = sum of block sizes for \( \lambda_i \); the exponent \( s_i \) in the minimal polynomial = largest block size. These agree only when there is a single block.
- Believing the Jordan basis is unique. Only the block multiset is unique; the basis (and the change-of-basis matrix \( P \)) is far from unique — e.g. any nonzero scalar rescales a whole chain.
- "Same characteristic polynomial \( \Rightarrow \) similar." False: \( J_2(0) \) vs \( 0_{2\times 2} \).
- Using JNF in floating-point computation. The Jordan form is a discontinuous function of the matrix: \( \begin{pmatrix} 0 & 1 \\ \varepsilon & 0 \end{pmatrix} \) is diagonalisable for every \( \varepsilon \neq 0 \) but is a \( J_2(0) \) at \( \varepsilon = 0 \). Numerically one uses the Schur form instead.
Discussion
Historically the theorem crystallised in the 1868–1874 period. Weierstrass’s 1868 theory of elementary divisors of matrix pencils contains the classification in equivalent form; Camille Jordan published the block form in his 1870 Traité des substitutions et des équations algébriques — notably working over the finite prime fields \( \mathbb{Z}/p\mathbb{Z} \), in the service of group theory, rather than over \( \mathbb{C} \). A sharp priority dispute with Kronecker followed; the modern verdict is that Weierstrass had the invariants first and Jordan the normal form. The "generalized eigenvector" language and the projection-based proof via Bézout are twentieth-century repackagings.
The right conceptual frame is module theory: \( T \) makes \( V \) a finitely generated torsion module over the PID \( \mathbb{C}[x] \), with \( x \) acting as \( T \). The structure theorem for such modules decomposes \( V \) as \( \bigoplus_j \mathbb{C}[x] / \left( (x - \lambda_j)^{k_j} \right) \), and each cyclic summand, in the basis \( \overline{(x-\lambda_j)^{k_j - 1}}, \dots, \overline{(x - \lambda_j)}, \overline{1} \), is exactly one Jordan block. Uniqueness of elementary divisors is uniqueness of blocks. Over a non-algebraically-closed field the same machine outputs the rational canonical form, whose companion blocks need no eigenvalues; the Jordan form is the special case where every irreducible factor of \( \chi_T \) is linear.
Analytically, the Jordan form is the finite-dimensional germ of spectral theory: the decomposition \( V = \bigoplus V_i \) with projections \( E_i \) that are polynomials in \( T \) prefigures the holomorphic functional calculus, where \( E_i = \frac{1}{2\pi i} \oint_{\gamma_i} (z - T)^{-1} \, dz \) around each eigenvalue. The nilpotent parts measure the failure of \( T \) to be normal-like; for normal operators (\( T T^* = T^* T \)) the spectral theorem forces every block to be \( 1 \times 1 \), which is why Jordan blocks never appear in the Hermitian theory. In applications the form is a theoretical instrument — it proves that solutions of \( \dot{x} = Ax \) are sums of \( t^a e^{\lambda t} \) terms, that \( \rho(A) \lt 1 \) forces \( A^m \to 0 \) — while actual computation routes through the numerically stable Schur triangularisation (Golub–Wilkinson analysed the ill-posedness of numerical JNF in 1976).
Deeper still, the nilpotent classification by partitions is the entrance to a rich geometry: the set of nilpotent matrices in \( M_n(\mathbb{C}) \) is an algebraic variety (the nilpotent cone), stratified by the \( GL_n \)-orbits that this theorem enumerates. Orbit closures are described by dominance order on partitions (Gerstenhaber–Hesselink): \( \overline{\mathcal{O}_{\mu}} \supseteq \mathcal{O}_{\nu} \iff \mu \trianglerighteq \nu \), the same order governing degenerations \( J_2 \rightsquigarrow J_1 \oplus J_1 \) under perturbation — the source of the discontinuity noted above. In Lie theory the Jordan–Chevalley decomposition generalises to semisimple algebraic groups, and the partition combinatorics reappears in Springer theory and the representation theory of \( S_n \). Common misconceptions. The Jordan form is not "just diagonalisation with extra steps": the set of non-diagonalisable matrices has measure zero, but it is exactly where the interesting degeneration phenomena (resonance in ODEs, defective eigenproblems, sensitivity of eigenvalues of order \( \varepsilon^{1/k} \)) live. And "generalized eigenvector" does not mean any solution of \( (A - \lambda I)^p v = 0 \) for huge \( p \) gives new information: the kernels stabilise at \( p = s_i \), the largest block size, never later.
Worked examples
Example 1. Find the Jordan normal form and a Jordan basis for \( A = \begin{pmatrix} 2 & 1 & 1 \\ 0 & 2 & 0 \\ 0 & 0 & 2 \end{pmatrix} \in M_3(\mathbb{C}) \).
Reading. One genuine \( 2 \)-chain and one loose eigenvector: \( A \) is defective with a one-dimensional defect.
Scope. Any \( 3 \times 3 \) with a triple eigenvalue is settled the same way by the pair \( \left( \operatorname{rank}(A - \lambda I), \operatorname{rank}(A-\lambda I)^2 \right) \).
Example 2. Find the Jordan normal form and a Jordan basis for \( A = \begin{pmatrix} 1 & 1 & 0 \\ -1 & 3 & 0 \\ -1 & 1 & 1 \end{pmatrix} \in M_3(\mathbb{C}) \).
Reading. The repeated eigenvalue \( 2 \) is defective (one eigenvector short); the simple eigenvalue \( 1 \) contributes a trivial block. General solution of \( \dot{x} = Ax \): \( x(t) = (c_1 + c_2 t) e^{2t} v + c_2 e^{2t} w + c_3 e^{t} u \).
Scope. The method (char. poly \( \to \) ranks \( \to \) chains, always solving for the eigenvector first inside \( \operatorname{im}(A - \lambda I) \) when several choices exist) works verbatim for any matrix whose eigenvalues are known exactly.
Problems
- Find the Jordan normal form of \( A = \begin{pmatrix} 1 & 1 \\ -1 & 3 \end{pmatrix} \) and an explicit \( P \) with \( P^{-1} A P \) in Jordan form.
Solution
\( \chi_A(x) = x^2 - 4x + 4 = (x-2)^2 \). \( A - 2I = \begin{pmatrix} -1 & 1 \\ -1 & 1 \end{pmatrix} \neq 0 \) has rank \( 1 \), so \( A \) is not diagonalisable and \( \text{JNF}(A) = J_2(2) \). Eigenvector: \( (A-2I)v = 0 \Rightarrow v = (1,1)^{\mathsf T} \). Generalized: \( (A - 2I)w = v \Rightarrow -w_1 + w_2 = 1 \); take \( w = (0,1)^{\mathsf T} \). Then \( P = \begin{pmatrix} 1 & 0 \\ 1 & 1 \end{pmatrix} \), and one checks \( Av = 2v \), \( Aw = (1,3)^{\mathsf T} = v + 2w \), so \( P^{-1} A P = \begin{pmatrix} 2 & 1 \\ 0 & 2 \end{pmatrix} \). - Let \( N \in M_4(\mathbb{C}) \) be nilpotent with \( N^2 = 0 \) and \( \operatorname{rank} N = 2 \). Determine the Jordan normal form of \( N \), and prove there is no other possibility.
Solution
All eigenvalues are \( 0 \) (nilpotency), so the JNF is \( \bigoplus J_{k_j}(0) \) with \( \sum k_j = 4 \). \( N^2 = 0 \) forces every \( k_j \leq 2 \) (a block \( J_k(0) \) satisfies \( J_k(0)^2 \neq 0 \) when \( k \geq 3 \), since it shifts a chain by two places and \( k - 2 \geq 1 \)). The number of blocks is \( \dim \ker N = 4 - \operatorname{rank} N = 2 \). A partition of \( 4 \) into exactly \( 2 \) parts, each \( \leq 2 \), must be \( 2 + 2 \). Hence \( \text{JNF}(N) = J_2(0) \oplus J_2(0) \), uniquely: any other partition either has the wrong number of parts (\( 4 = 2+1+1 \) has \( 3 \), \( 4 = 4 \) has \( 1 \)) or a part exceeding \( 2 \) (\( 3 + 1 \)). - Prove that every \( A \in M_n(\mathbb{C}) \) is similar to its transpose \( A^{\mathsf T} \).
Solution
Step 1: it suffices to prove it for a single Jordan block. If \( A = P J P^{-1} \) with \( J = \bigoplus_j J_{k_j}(\lambda_j) \), then \( A^{\mathsf T} = (P^{-1})^{\mathsf T} J^{\mathsf T} P^{\mathsf T} \), so \( A^{\mathsf T} \sim J^{\mathsf T} = \bigoplus_j J_{k_j}(\lambda_j)^{\mathsf T} \). If each \( J_k(\lambda)^{\mathsf T} \sim J_k(\lambda) \), then (conjugating blockwise by a block-diagonal matrix) \( J^{\mathsf T} \sim J \), whence \( A^{\mathsf T} \sim J \sim A \). Step 2: single block. Let \( R \in M_k \) be the reversal matrix, \( R_{ab} = 1 \) if \( a + b = k+1 \), else \( 0 \); note \( R^{-1} = R \). For any \( B \), \( (R B R)_{ab} = B_{k+1-a, \, k+1-b} \), i.e. conjugation by \( R \) rotates the matrix by half a turn. Applying this to \( B = J_k(\lambda) \): the diagonal \( (a = b) \) maps to itself, and the superdiagonal entry at \( (a, a+1) \) maps to position \( (k+1-a, k-a) \), which is the subdiagonal — exactly \( J_k(\lambda)^{\mathsf T} \). Hence \( R J_k(\lambda) R^{-1} = J_k(\lambda)^{\mathsf T} \). (Remark: the statement is true over every field via the rational canonical form; the JNF proof needs \( \mathbb{C} \) only because it invokes the theorem above.) - How many similarity classes of nilpotent matrices are there in \( M_5(\mathbb{C}) \)? List them.
Solution
A nilpotent matrix has all eigenvalues \( 0 \), so by existence and uniqueness of the JNF its similarity class is exactly the multiset of block sizes: a partition of \( 5 \). Conversely every partition occurs (write down the block-diagonal matrix). The partitions of \( 5 \) are: \( 5;\; 4+1;\; 3+2;\; 3+1+1;\; 2+2+1;\; 2+1+1+1;\; 1+1+1+1+1 \) — seven in total. So there are exactly \( 7 \) classes, from the regular nilpotent \( J_5(0) \) down to the zero matrix \( J_1(0)^{\oplus 5} \). (The count of classes of nilpotents in \( M_n \) is the partition function \( p(n) \).) - Let \( A \in M_n(\mathbb{C}) \) satisfy \( A^m = I \) for some integer \( m \geq 1 \). Prove that \( A \) is diagonalisable, and show by counterexample that the conclusion fails over a field of positive characteristic.
Solution
By the theorem, \( A \sim J = \bigoplus_j J_{k_j}(\lambda_j) \), and \( A^m = I \iff J^m = I \). Since \( A \) is invertible (\( A \cdot A^{m-1} = I \)), every \( \lambda_j \neq 0 \). Suppose some block has \( k := k_j \geq 2 \). Write \( J_k(\lambda) = \lambda I + N \) with \( N = J_k(0) \), \( N^k = 0 \), and \( \lambda I \) commuting with \( N \); the binomial theorem (valid for commuting matrices) gives \[ J_k(\lambda)^m = \sum_{a=0}^{m} \binom{m}{a} \lambda^{m-a} N^a = \lambda^m I + m \lambda^{m-1} N + \binom{m}{2}\lambda^{m-2} N^2 + \cdots \] The \( (1,2) \) entry of this matrix is \( m \lambda^{m-1} \) (only the \( N^1 \) term contributes to the first superdiagonal). In characteristic zero \( m \lambda^{m-1} \neq 0 \) since \( m \geq 1 \) and \( \lambda \neq 0 \) — contradicting \( J_k(\lambda)^m = I_k \), whose \( (1,2) \) entry is \( 0 \). Hence every block has size \( 1 \) and \( J \) is diagonal, i.e. \( A \) is diagonalisable. (Equivalently: the minimal polynomial divides \( x^m - 1 \), which has \( m \) distinct roots over \( \mathbb{C} \), so all blocks have size \( 1 \).) Counterexample in characteristic \( p \): over \( \mathbb{F}_p \) take \( A = J_2(1) = \begin{pmatrix} 1 & 1 \\ 0 & 1 \end{pmatrix} \). Then \( A^p = (I + N)^p = I + pN + \binom{p}{2} N^2 = I \) since \( p \equiv 0 \) and \( N^2 = 0 \); but \( A \neq I \) is a nontrivial Jordan block, hence not diagonalisable (its only eigenvalue is \( 1 \), and \( \ker(A - I) \) is one-dimensional). Here \( x^p - 1 = (x-1)^p \) has a repeated root, which is exactly where the characteristic-zero argument breaks.