physics2u
Tier
⌕ Search ⌘K
Derivation

Diagonalizability and the Eigenbasis Criterion

Statement

Let \(T\) be a linear operator on a finite-dimensional vector space \(V\) with \(\dim V = n\), represented in some basis by a matrix \(A\). Then \(T\) is diagonalizable (similar to a diagonal matrix, \(P^{-1}AP = D\)) if and only if its eigenvectors span \(V\) — equivalently, if and only if the characteristic polynomial splits over the field and, for every eigenvalue \(\lambda_i\), the geometric multiplicity equals the algebraic multiplicity, \(g_i = a_i\). In that case \(V = E_{\lambda_1}\oplus\cdots\oplus E_{\lambda_k}\) is the direct sum of the eigenspaces.

Why it matters

Diagonalization is the single most useful structural fact about a linear operator: in an eigenbasis the operator becomes independent scalar multiplication on each axis, so powers, exponentials, and functions of the operator reduce to functions applied entrywise. Normal-mode analysis, the time-evolution operator \(e^{-i\hat H t/\hbar}\), decoupling of coupled oscillators, and the whole apparatus of the spectral theorem all rest on knowing when this reduction is possible.

The criterion is also a sharp diagnostic. It isolates the two distinct obstructions to diagonalization — the field failing to contain the eigenvalues, and an eigenvalue whose eigenspace is too small (a defective operator) — and shows they are the only two. This is the symmetry thread's cleanest example of a global geometric property (an eigenbasis) being certified by a purely local count of multiplicities.

Assumptions
Finite dimension, \(n < \infty\).In infinite dimensions an operator may have no eigenvectors at all (the shift operator on \(\ell^2\) has empty point spectrum), and the multiplicity bookkeeping that drives the proof — a finite sum reaching \(n\) — has no meaning. The characteristic polynomial splits into linear factors over the field.Over a non-closed field such as \(\mathbb{R}\), a rotation has no real eigenvalues; no real eigenbasis exists even though the operator is perfectly well-behaved. Diagonalizability is then a statement about \(\mathbb{C}\), recovered by complexifying. Eigenvalues and their multiplicities come from the characteristic polynomial \(p(\lambda)=\det(A-\lambda I)\).This imports the eigenvalue–root correspondence and the similarity-invariance of \(p\); without it "algebraic multiplicity" is undefined and the count \(\sum a_i \le n\) is unavailable. Dimension is basis-independent, so \(\dim E_{\lambda_i}\) and "eigenvectors span \(V\)" are well-posed.If dimension were not an invariant, the geometric multiplicity \(g_i=\dim\ker(T-\lambda_i I)\) would depend on coordinates and could not be compared to \(a_i\).
Derivation
1
\[ T \text{ diagonalizable} \iff \exists\, P \text{ invertible},\ P^{-1}AP = D = \operatorname{diag}(d_1,\dots,d_n) \]
Definition: \(T\) is diagonalizable when some basis makes its matrix diagonal, and a change of basis acts by similarity \(A\mapsto P^{-1}AP\). A
2
\[ P^{-1}AP = D \iff AP = PD \iff A\,\mathbf{p}_j = d_j\,\mathbf{p}_j \quad (j=1,\dots,n) \]
Read the matrix identity column by column: the \(j\)-th column of \(PD\) is \(d_j\mathbf{p}_j\). So each column \(\mathbf{p}_j\) of \(P\) is an eigenvector with eigenvalue \(d_j\). A
3
\[ P \text{ invertible} \iff \{\mathbf{p}_1,\dots,\mathbf{p}_n\} \text{ linearly independent} \iff \text{a basis of eigenvectors exists} \]
An \(n\times n\) matrix is invertible iff its columns are independent, and \(n\) independent vectors in an \(n\)-dimensional space form a basis. Hence diagonalizable \(\iff\) eigenvectors span \(V\). B
4
\[ \text{eigenvectors of distinct eigenvalues are independent:}\quad \textstyle\sum_i E_{\lambda_i}=\bigoplus_i E_{\lambda_i} \]
Suppose \(\sum_i \mathbf{v}_i=\mathbf{0}\) with \(\mathbf{v}_i\in E_{\lambda_i}\). Applying \(\prod_{i\ne m}(T-\lambda_i I)\) annihilates every term but the \(m\)-th and multiplies it by \(\prod_{i\ne m}(\lambda_m-\lambda_i)\ne 0\), forcing \(\mathbf{v}_m=\mathbf{0}\). So the eigenspace sum is direct. C
5
\[ 1 \le g_i \equiv \dim E_{\lambda_i} \le a_i \equiv \text{(multiplicity of }\lambda_i\text{ in }p) \]
Extend a basis of \(E_{\lambda_i}\) to a basis of \(V\); in it \(A\) is block upper-triangular with a leading \(\lambda_i I_{g_i}\) block, so \(p(\lambda)=(\lambda_i-\lambda)^{g_i}\,q(\lambda)\) and \(a_i\ge g_i\). C
6
\[ \dim\!\Big(\bigoplus_i E_{\lambda_i}\Big) = \sum_{i=1}^{k} g_i, \qquad \sum_{i=1}^{k} a_i \le n \]
Dimensions add over a direct sum (step 4); the algebraic multiplicities sum to at most \(n=\deg p\), with equality precisely when \(p\) splits into linear factors. B
7
\[ \text{eigenvectors span } V \iff \sum_i g_i = n \]
The eigenvectors span iff the direct sum of eigenspaces fills \(V\), i.e. its dimension reaches \(n\) (using dimension-invariance). Combined with step 3, this is exactly diagonalizability. A
8
\[ \sum_i g_i = n,\ \ g_i\le a_i,\ \ \sum_i a_i\le n \ \Longrightarrow\ \sum_i a_i = n \ \text{and}\ g_i=a_i\ \forall i \]
A sum of terms each \(\le\) its counterpart can only reach \(n\) if the enclosing bound \(\sum a_i\) also equals \(n\) (so \(p\) splits) and no slack remains, forcing termwise equality \(g_i=a_i\). The converse is immediate: \(g_i=a_i\) and \(\sum a_i=n\) give \(\sum g_i=n\). C
Result
\[ T \text{ diagonalizable} \iff \text{eigenvectors span } V \iff \Big(p \text{ splits and } g_i=a_i\ \forall i\Big) \iff V=\bigoplus_{i} E_{\lambda_i} \]

Reading. An operator can be reduced to independent scalings on \(n\) axes exactly when it supplies enough eigenvectors to build a basis. Two things can go wrong and only two: the eigenvalues may live outside the field (\(p\) does not split), or a repeated eigenvalue may be "defective," its eigenspace \(E_{\lambda_i}\) smaller than the root multiplicity \(a_i\). If neither happens, the eigenspaces tile the whole space as a direct sum. Repetition of an eigenvalue is permitted — the identity is diagonal with a single eigenvalue of multiplicity \(n\) — it is only a shortfall \(g_i<a_i\) that obstructs.

Units check. Diagonalizability is a purely structural (dimensionless) statement: \(P\) is a change of basis and carries no units. Each eigenvalue \(\lambda_i=d_i\) inherits the units of the operator's diagonal — energies [J] for a Hamiltonian, inverse times [s\({}^{-1}\)] for a rate operator — and the relation \(A\mathbf{p}_j=d_j\mathbf{p}_j\) is homogeneous, both sides carrying [\(A\)]·[\(\mathbf{p}\)]. The multiplicities \(g_i,a_i\) are pure integers.

Limiting cases
  • Distinct eigenvalues: if \(p\) has \(n\) distinct roots then every \(a_i=1\), forcing \(g_i=1=a_i\); the operator is automatically diagonalizable. Distinctness is sufficient but not necessary.
  • Already diagonal or triangular with distinct diagonal: the standard basis vectors are eigenvectors and span trivially.
  • Identity / any scalar operator \(cI\): one eigenvalue of algebraic multiplicity \(n\), and \(E_c = V\) so \(g=n=a\) — diagonalizable despite maximal repetition.
  • Hermitian / real-symmetric / normal operators (symmetry thread): always diagonalizable, by the spectral theorem — a special case where \(g_i=a_i\) is guaranteed and the eigenbasis is even orthonormal.
  • Single Jordan block \(J_n(\lambda)\): the extreme opposite — \(a=n\) but \(g=1\), maximally defective.
Breaks when
  • The field is not algebraically closed. Over \(\mathbb{R}\) the rotation \(R(\theta)=\begin{pmatrix}\cos\theta&-\sin\theta\\\sin\theta&\cos\theta\end{pmatrix}\) with \(0<\theta<\pi\) has characteristic polynomial \(\lambda^2-2\cos\theta\,\lambda+1\), no real roots, hence no real eigenbasis. It becomes diagonalizable only after complexifying.
  • Defective eigenvalues (Jordan blocks). A repeated eigenvalue with \(g_i<a_i\), e.g. \(\begin{pmatrix}\lambda&1\\0&\lambda\end{pmatrix}\), has only a one-dimensional eigenspace; no basis of eigenvectors exists and the best reduction is Jordan form, not diagonal.
  • Infinite-dimensional spaces. Operators on function spaces can have continuous spectra or no eigenvectors; "eigenvectors span" must be replaced by spectral measures, and the finite multiplicity count is meaningless.
Failure modes
  • "Repeated eigenvalue ⇒ not diagonalizable": false. The identity has one eigenvalue repeated \(n\) times and is already diagonal. Only \(g_i<a_i\) obstructs, never repetition itself.
  • Conflating algebraic and geometric multiplicity: assuming \(g_i=a_i\) automatically. They are equal for non-defective operators only; the whole criterion is about when they coincide.
  • Testing diagonalizability by the determinant or trace: \(\det A\ne 0\) (invertibility) is unrelated — a Jordan block \(\begin{pmatrix}2&1\\0&2\end{pmatrix}\) is invertible yet not diagonalizable.
  • Forgetting the field: declaring a real rotation "not diagonalizable" without noting it is diagonalizable over \(\mathbb{C}\); conflating the field obstruction with the defect obstruction.
  • Counting \(g_i\) as "number of eigenvectors found" rather than \(\dim\ker(T-\lambda_i I)\): any scalar multiple of an eigenvector is another eigenvector; only the dimension of the eigenspace matters.
Discussion

The criterion decomposes diagonalizability into two independent questions. First, do the eigenvalues exist in the field? This is the content of "\(p\) splits," and it is a question about scalars, resolved by working over \(\mathbb{C}\) where the fundamental theorem of algebra guarantees splitting. Second, is each eigenspace as large as its root demands? This is the geometric question \(g_i=a_i\), and it is genuinely about the operator's action, not about the field. Separating these two is the whole conceptual payoff: a matrix can fail to be diagonalizable for a "cheap" reason (wrong field, fixable) or a "deep" reason (defect, structural).

The defect obstruction is measured exactly by Jordan form. When \(g_i<a_i\), the missing \(a_i-g_i\) dimensions are supplied by generalized eigenvectors — solutions of \((T-\lambda_i I)^m\mathbf{v}=\mathbf{0}\) for \(m>1\) — and the operator restricted to that generalized eigenspace is a nilpotent shift plus \(\lambda_i I\). Diagonalizability is then precisely the case where every nilpotent part vanishes, i.e. the minimal polynomial has no repeated roots. This is the algebraic mirror of the multiplicity criterion: \(T\) is diagonalizable iff its minimal polynomial factors into distinct linear factors.

Physically, defectiveness is the signature of a non-normal operator, and it appears wherever dissipation or coupling breaks the symmetry that would guarantee an eigenbasis. A critically damped oscillator sits exactly at the parameter value where two modes coalesce into a single Jordan block — the reason its solution carries a \(t\,e^{-\gamma t}\) term rather than two clean exponentials. In such cases the eigenvector picture literally runs out of directions, and the extra polynomial-in-\(t\) factors are the fingerprint of the generalized eigenvectors that had to be recruited.

The sharpest statement lives at the level of the minimal polynomial \(m_T\). Because \(V\) decomposes as a direct sum of generalized eigenspaces (primary decomposition), and \(T\) acts on each as scalar-plus-nilpotent, diagonalizability is equivalent to \(m_T(\lambda)=\prod_i(\lambda-\lambda_i)\) with all factors simple. This is strictly stronger information than the characteristic polynomial carries: \(p\) and \(m_T\) share roots but the defect is visible only in the multiplicities of \(m_T\). Two operators can share a characteristic polynomial yet differ in diagonalizability — \(\operatorname{diag}(2,2)\) versus the Jordan block \(J_2(2)\) — and it is the minimal polynomial, \((\lambda-2)\) versus \((\lambda-2)^2\), that tells them apart.

Common misconceptions. (i) Diagonalizability is a property of the operator over a specified field, not an absolute one — always name the field. (ii) The eigenvectors, not the eigenvalues, must span; distinct eigenvalues is a convenient sufficient condition, never a necessary one. (iii) A basis of eigenvectors need not be orthogonal — orthogonality is the extra bonus of the spectral theorem for normal operators, not part of diagonalizability itself.

Worked examples
1
\[ \text{Example 1 — a defective shear.}\quad A=\begin{pmatrix}4&1\\0&4\end{pmatrix} \]
A triangular matrix with a repeated diagonal entry; test whether the eigenspace is large enough. Symbols first.
2
\[ p(\lambda)=\det\begin{pmatrix}4-\lambda&1\\0&4-\lambda\end{pmatrix}=(4-\lambda)^2 \;\Rightarrow\; \lambda=4,\ a=2 \]
The triangular determinant is the product of diagonal entries; the single eigenvalue \(4\) has algebraic multiplicity \(2\).
3
\[ E_4=\ker(A-4I)=\ker\begin{pmatrix}0&1\\0&0\end{pmatrix}=\operatorname{span}\!\Big\{\begin{pmatrix}1\\0\end{pmatrix}\Big\},\quad g=\dim E_4=1 \]
The condition \((A-4I)\mathbf{v}=\mathbf 0\) forces the second component to zero; only one independent eigenvector. Now compare the counts.
\[ g=1 < a=2 \;\Rightarrow\; A \text{ is NOT diagonalizable} \]

Reading. The eigenspace is one-dimensional but the root is double, so the eigenvectors cannot span \(\mathbb{R}^2\). The best reduction is the Jordan block itself; a generalized eigenvector \((0,1)^{\!\top}\) completes the basis.

Units check. Entries are pure numbers, so eigenvalue and multiplicities are dimensionless; the verdict is structural and field-independent (it fails over \(\mathbb{R}\) and over \(\mathbb{C}\) alike).

1
\[ \text{Example 2 — a repeated eigenvalue that IS diagonalizable.}\quad A=\begin{pmatrix}0&1&1\\1&0&1\\1&1&0\end{pmatrix} \]
The adjacency matrix of a triangle (real-symmetric, so expect an eigenbasis). Test the defect criterion on its repeated eigenvalue. Symbols first.
2
\[ A = J - I,\ \ J=\text{all-ones matrix};\quad J \text{ has eigenvalues } 3,0,0 \;\Rightarrow\; A:\ \lambda=2\ (a=1),\ \lambda=-1\ (a=2) \]
Shifting by \(-I\) shifts every eigenvalue by \(-1\). The rank-one matrix \(J\) has eigenvalue \(3\) once and \(0\) twice.
3
\[ E_{-1}=\ker(A+I)=\ker\begin{pmatrix}1&1&1\\1&1&1\\1&1&1\end{pmatrix}=\{\,x+y+z=0\,\},\quad g_{-1}=2 \]
All three rows collapse to the single equation \(x+y+z=0\), a plane; its dimension is \(3-1=2\) by rank–nullity.
4
\[ g_{-1}=2=a_{-1},\quad g_{2}=1=a_{2},\quad g_{-1}+g_{2}=3=n \]
Both eigenvalues meet \(g_i=a_i\) and the geometric multiplicities sum to \(n=3\): the eigenvectors span.
\[ A \text{ IS diagonalizable},\quad P=\begin{pmatrix}1&1&1\\1&-1&0\\1&0&-1\end{pmatrix},\ \ P^{-1}AP=\operatorname{diag}(2,-1,-1) \]

Reading. The columns are the eigenvectors \((1,1,1)\) for \(\lambda=2\) and \((1,-1,0),(1,0,-1)\) spanning the plane \(E_{-1}\). A repeated eigenvalue caused no trouble because its eigenspace was full — the direct contrast with Example 1.

Units check. Dimensionless throughout; had the entries been coupling energies in eV, the diagonal would read \(2,-1,-1\) eV, and \(P\) would remain unitless.

Problems
  1. Decide whether \(A=\begin{pmatrix}2&1\\0&2\end{pmatrix}\) is diagonalizable over \(\mathbb{R}\).
    Solution\(p(\lambda)=(2-\lambda)^2\), so \(\lambda=2\) with \(a=2\). The eigenspace is \(\ker\begin{pmatrix}0&1\\0&0\end{pmatrix}=\operatorname{span}\{(1,0)^{\!\top}\}\), so \(g=1\). Since \(g=1<2=a\), not diagonalizable (a single Jordan block). Note \(\det A=4\ne0\): invertibility is irrelevant.
  2. Show \(A=\begin{pmatrix}1&2\\0&3\end{pmatrix}\) is diagonalizable and give \(P\).
    SolutionTriangular, so \(\lambda=1,3\) — two distinct eigenvalues, hence automatically diagonalizable. Eigenvectors: for \(\lambda=1\), \(\ker\begin{pmatrix}0&2\\0&2\end{pmatrix}=\operatorname{span}\{(1,0)^{\!\top}\}\); for \(\lambda=3\), solve \(\begin{pmatrix}-2&2\\0&0\end{pmatrix}\mathbf v=\mathbf 0\Rightarrow \mathbf v=(1,1)^{\!\top}\). So \(P=\begin{pmatrix}1&1\\0&1\end{pmatrix}\), \(P^{-1}AP=\operatorname{diag}(1,3)\).
  3. Is \(A=\begin{pmatrix}3&0&0\\0&3&0\\1&0&3\end{pmatrix}\) diagonalizable?
    SolutionTriangular \(\Rightarrow\) \(\lambda=3\) with \(a=3\). Geometric multiplicity: \(A-3I=\begin{pmatrix}0&0&0\\0&0&0\\1&0&0\end{pmatrix}\) has rank \(1\), so \(g=\text{nullity}=3-1=2\). Since \(g=2<3=a\), not diagonalizable (one Jordan block of size 2 and one of size 1).
  4. A \(3\times3\) matrix has eigenvalues \(2,2,5\) and satisfies \(\operatorname{rank}(A-2I)=1\). Is it diagonalizable?
    SolutionFor \(\lambda=2\): \(a=2\), and \(g=\dim\ker(A-2I)=3-\operatorname{rank}(A-2I)=3-1=2\), so \(g=a=2\). For \(\lambda=5\): \(a=1\) forces \(g=1=a\). Both multiplicities match and \(g_2+g_5=3=n\), so yes, diagonalizable. Had \(\operatorname{rank}(A-2I)=2\) instead, then \(g=1<2\) and it would fail.
  5. Prove that if an \(n\times n\) matrix over \(\mathbb{C}\) has \(n\) distinct eigenvalues then it is diagonalizable, and show the converse is false.
    SolutionForward: distinct eigenvalues give \(k=n\) eigenvalues each with \(a_i=1\); since \(1\le g_i\le a_i=1\) we get \(g_i=1\) for all \(i\), so \(\sum g_i=n\) and the eigenvectors span (equivalently, eigenvectors of distinct eigenvalues are independent by step 4, and \(n\) independent vectors form a basis). Converse false: the identity \(I_n\) is diagonalizable (it is already diagonal) yet has the single eigenvalue \(1\) repeated \(n\) times — distinctness is sufficient but not necessary.