physics2u
Tier
⌕ Search ⌘K
Derivation

Paraxial Ray-Transfer (ABCD) Matrices

D-199 Home PU-206 Threads light Depends on Laws of Reflection and Refraction from Fermat, matrix-multiplication-and-determinant
Statement

Representing a paraxial ray in a meridional plane by the two-component state \(\mathbf{V}=(y,\,n\theta)^{\mathsf T}\) — transverse height \(y\) and reduced slope \(u=n\theta\) — we derive the three elementary ray-transfer matrices: free propagation through a length \(d\) of index \(n\), refraction at a spherical interface of radius \(R\) between indices \(n_1,n_2\), and the thin lens. We show that traversal of any sequence of elements is the ordered matrix product \(\mathbf{V}_{\text{out}}=M_N\cdots M_2 M_1\,\mathbf{V}_{\text{in}}\), and that in reduced coordinates every elementary matrix — and hence every system matrix — has determinant exactly \(+1\).

Why it matters

The ABCD formalism turns geometrical optics into linear algebra: an entire instrument, from a single interface to a multi-element zoom or a laser resonator, collapses to one \(2\times2\) matrix that can be multiplied, inverted, and iterated. Imaging conditions, cardinal points, and system power are then read directly off the four entries rather than traced ray by ray.

The unit determinant is not a bookkeeping accident. It encodes the conservation of a phase-space area — the Smith–Helmholtz (Lagrange) invariant — which is the ray-optics shadow of Liouville's theorem and of étendue conservation. That single constraint links optical throughput, brightness limits, and why you cannot losslessly concentrate incoherent light beyond the source's phase-space volume.

Assumptions
Paraxial regimeAll rays make small angles with the optical axis and strike surfaces near the axis, so \(\sin\theta\approx\theta\), \(\tan\theta\approx\theta\), \(\cos\theta\approx1\). If dropped, the map from \((y,\theta)\) to output is nonlinear (Seidel aberrations enter at third order) and no \(2\times2\) matrix suffices. Meridional (2D) raysRays lie in a single plane containing the axis, so one \((y,\theta)\) pair captures the state. If dropped, skew rays require the full \(4\times4\) symplectic treatment. Rotationally symmetric, smooth surfacesEach interface is a sphere (or plane) centred on the axis with a well-defined radius \(R\). If dropped, the local curvature — and hence the refraction entry \(-P\) — is not a single number. Homogeneous media between surfacesIndex is piecewise constant, so propagation is a straight line. If dropped (gradient-index media), propagation itself is governed by a differential ray equation and its transfer matrix is no longer the simple shear.
Sign convention fixed onceDistances, angles and radii carry consistent signs (here: light travels left-to-right; \(R>0\) when the centre of curvature lies downstream; angles positive counterclockwise from the axis). If dropped, individual entries flip sign inconsistently and composed systems give wrong cardinal points even when each element looks right.
Derivation
1
\[ \mathbf{V}=\begin{pmatrix} y \\ u \end{pmatrix},\qquad u \equiv n\theta \]
Define the state. Using the reduced slope \(u=n\theta\) (an optical direction cosine to first order) rather than the bare angle \(\theta\) is the choice that makes every elementary determinant equal \(+1\); it is the canonical momentum conjugate to \(y\). A
2
\[ y_2 = y_1 + d\,\theta_1 = y_1 + \frac{d}{n}\,u_1,\qquad \theta_2=\theta_1 \]
Free propagation. In a straight line the slope is unchanged; the height grows by \(\tan\theta_1\approx\theta_1\) per unit length over distance \(d\). Substituting \(\theta_1=u_1/n\) introduces the reduced length \(d/n\). A
3
\[ \begin{pmatrix} y_2\\ u_2\end{pmatrix}=\underbrace{\begin{pmatrix} 1 & d/n \\ 0 & 1\end{pmatrix}}_{T}\begin{pmatrix} y_1\\ u_1\end{pmatrix},\qquad \det T = 1 \]
Read the propagation matrix off the linear relations of Step 2 (\(u_2=u_1\)). It is a shear; its determinant is manifestly \(1\). A
4
\[ n_1 i_1 = n_2 i_2 \]
Refraction at a spherical surface. Snell's law \(n_1\sin i_1=n_2\sin i_2\) (a prior result) linearises to equal reduced angles of incidence and refraction, where \(i_1,i_2\) are measured from the local surface normal. A
5
\[ \beta = \frac{y}{R},\qquad i_1=\theta_1+\beta,\qquad i_2 = \theta_2+\beta \]
Geometry of the normal. The surface normal at height \(y\) passes through the centre of curvature on the axis, so it makes angle \(\beta=y/R\) with the axis (paraxial). Each ray's incidence angle from the normal is its axial angle plus this tilt; \(y\) is continuous across the (thin) interface so \(y_2=y_1\equiv y\). B
6
\[ n_1(\theta_1+\beta)=n_2(\theta_2+\beta)\ \Rightarrow\ n_2\theta_2 = n_1\theta_1 -(n_2-n_1)\beta \]
Insert Step 5 into Step 4 and rearrange symbolically, isolating the output reduced slope. The \(n\beta\) terms collect into the difference \((n_2-n_1)\). A
7
\[ u_2 = n_2\theta_2 = u_1 - \frac{n_2-n_1}{R}\,y \;\equiv\; u_1 - P\,y,\qquad P\equiv \frac{n_2-n_1}{R} \]
Substitute \(\beta=y/R\) and identify the surface power \(P\) (units of inverse length, i.e. dioptres). In reduced coordinates the height enters directly through \(-P\), with no residual index ratio. A
8
\[ \begin{pmatrix} y_2\\ u_2\end{pmatrix}=\underbrace{\begin{pmatrix} 1 & 0\\ -P & 1\end{pmatrix}}_{\mathcal R}\begin{pmatrix} y_1\\ u_1\end{pmatrix},\qquad \det \mathcal R = 1 \]
Assemble the refraction matrix from \(y_2=y_1\) (Step 5) and Step 7. Because the reduced slope absorbs the index jump, the determinant is exactly \(1\); in the bare-angle convention \((y,\theta)\) the analogous matrix has \(\det=n_1/n_2\). B
9
\[ L=\mathcal R_2\,\mathcal R_1=\begin{pmatrix}1&0\\-P_2&1\end{pmatrix}\begin{pmatrix}1&0\\-P_1&1\end{pmatrix}=\begin{pmatrix}1&0\\-(P_1+P_2)&1\end{pmatrix} \]
Thin lens = two refractions with negligible axial gap between them, so the propagation matrix between surfaces is the identity and the two \(\mathcal R\)'s multiply directly. Lower-triangular matrices with unit diagonal add their off-diagonal entries. B
10
\[ P = P_1+P_2 = \frac{n_\ell-n_m}{R_1}+\frac{n_m-n_\ell}{R_2}=(n_\ell-n_m)\!\left(\frac{1}{R_1}-\frac{1}{R_2}\right) \]
Sum the surface powers for a lens of index \(n_\ell\) in a medium \(n_m\): the first surface takes \(n_m\!\to\!n_\ell\), the second \(n_\ell\!\to\!n_m\). This is the Lensmaker's relation; in air \(n_m=1\) and \(P=1/f\). B
11
\[ L=\begin{pmatrix}1&0\\-1/f&1\end{pmatrix},\qquad \det L = 1 \]
Write the thin-lens matrix in air using \(P=1/f\). It has the refraction form with the total power; determinant \(1\) since it is a product of unit-determinant factors. A
12
\[ \mathbf{V}_{\text{out}} = M_N\big(M_{N-1}(\cdots(M_1\mathbf{V}_{\text{in}}))\big)=\big(M_N\cdots M_1\big)\,\mathbf{V}_{\text{in}}\equiv M\,\mathbf{V}_{\text{in}} \]
Composition. Each element acts on the running state in turn; associativity of matrix multiplication (a prior result) lets us collapse the chain into one system matrix. The element the ray meets first stands rightmost. A
13
\[ \det M = \det\!\big(M_N\cdots M_1\big)=\prod_{k=1}^{N}\det M_k = \prod_{k=1}^{N} 1 = 1 \]
Multiplicativity of the determinant (a prior result) applied to Steps 3, 8, 11. Every elementary factor has unit determinant in reduced coordinates, so the whole system does. This is the algebraic face of the Lagrange invariant. B
Result
\[ T=\begin{pmatrix}1&d/n\\[2pt]0&1\end{pmatrix},\quad \mathcal R=\begin{pmatrix}1&0\\[2pt]-\dfrac{n_2-n_1}{R}&1\end{pmatrix},\quad L=\begin{pmatrix}1&0\\[2pt]-\dfrac{1}{f}&1\end{pmatrix},\quad M=\!\!\prod_{k=N}^{1}\!\! M_k,\ \ \det M=1 \]

Reading. A ray's transverse state \((y,\,n\theta)\) transforms linearly through any paraxial system. Propagation is a shear that tilts height without changing reduced slope; refraction and the thin lens are focusing kicks that change slope in proportion to height (with strength the power \(P=1/f\)) without moving the ray. The generic entries \(\left(\begin{smallmatrix}A&B\\C&D\end{smallmatrix}\right)\) mean: \(A\) is transverse (angular) magnification bookkeeping, \(B\) has the dimension of a reduced length, \(C=-P\) is minus the system power, and \(D\) is angular magnification bookkeeping. The constraint \(AD-BC=1\) is the conserved phase-space area.

Units check. With \([y]=\text{m}\) and \([u]=[n\theta]=1\) (dimensionless), the entries carry \([A]=1,\ [B]=\text{m},\ [C]=\text{m}^{-1},\ [D]=1\). Then \(d/n\) is a length ✓, \(P=(n_2-n_1)/R\) and \(1/f\) are \(\text{m}^{-1}\) (dioptres) ✓, and \(\det=AD-BC\) has units \(1-(\text{m})(\text{m}^{-1})=1\), a pure number, consistent with equalling \(1\) ✓.

Limiting cases
  • Plane interface (\(R\to\infty\)): \(P\to0\), \(\mathcal R\to I\) in reduced coordinates — a flat surface bends no reduced slope (all bending is hidden in the \(n\theta\) rescaling; in bare angles it still refracts by \(n_1/n_2\)).
  • Zero-power (afocal) element (\(f\to\infty\)): \(L\to I\); a "window" or a matched afocal telescope leaves reduced slope untouched.
  • Zero gap (\(d\to0\)): \(T\to I\); stacking elements with no space between them just multiplies their kicks (Step 9).
  • Same medium in and out: even in the bare-angle \((y,\theta)\) convention, \(\det M=n_{\text{in}}/n_{\text{out}}=1\), because the internal index jumps telescope.
  • Imaging condition (\(B=0\)): input height maps to output height independent of input slope — a point images to a point, and \(A\) becomes the transverse magnification.
Breaks when
  • Large angles / large apertures. Once \(\sin\theta\neq\theta\) to the accuracy required, the true map is nonlinear: spherical aberration, coma, astigmatism, and distortion appear at third order in \((y,\theta)\) and no single \(2\times2\) matrix reproduces the wavefront. The matrix gives only the Gaussian (first-order) image.
  • Skew rays and non-rotationally-symmetric systems. Cylindrical lenses, tilted or decentred elements, and skew rays couple the two transverse planes; the correct object is a \(4\times4\) symplectic (or \(6\times6\)) matrix, and the meridional \(2\times2\) reduction is invalid.
  • Gradient-index or otherwise inhomogeneous media. When \(n=n(\mathbf r)\) within a segment, propagation is not a straight-line shear — the ray obeys \(\tfrac{d}{ds}\!\left(n\frac{d\mathbf r}{ds}\right)=\nabla n\), and the segment's matrix must be integrated from that equation.
  • Diffraction-dominated / small features. Ray optics itself fails when apertures approach the wavelength; the ABCD entries then survive only as inputs to the diffraction integral (the ABCD Gaussian-beam / Collins-integral extension), not as a complete description.
Failure modes
  • Multiplying in traversal order. Writing \(M=M_1 M_2\cdots M_N\) (first element leftmost). The ray meets element 1 first, so it must be rightmost: \(M=M_N\cdots M_1\). Matrices do not commute, so this silently gives the wrong system.
  • Mixing angle conventions. Deriving \(\mathcal R\) with bare \(\theta\) (giving \(\det=n_1/n_2\)) but then propagating with the reduced-\(u\) matrix, or vice versa. Pick one convention and use its full set consistently.
  • Dropping the index in propagation. Using \(d\) instead of \(d/n\) for a segment inside glass or water; the reduced propagation length is \(d/n\).
  • Sign of \(R\). Forgetting that the second surface of a biconvex lens has \(R_2<0\); using \(|R|\) makes both surfaces "converging" and doubles the apparent power.
  • Reading \(C\) as the power. The entry is \(C=-P=-1/f\); students quote a negative focal length for a converging lens by dropping the sign.
  • Treating a thick lens as thin. Omitting the internal \(T(t/n_\ell)\) between the two surface matrices when the axial thickness \(t\) is not negligible.
Discussion

The formalism separates optics cleanly into two moves: shears (propagation) that change where a ray is without changing where it is going, and kicks (refraction, lenses, mirrors) that change where it is going without moving it. Every paraxial instrument is a word in these two generators. Because the generators are triangular with unit diagonal, their commutators are non-trivial — this non-commutativity is exactly why lens spacing matters and why a telescope is not the same optic as its elements reversed.

The system entries carry direct physical meaning. \(B=0\) is the imaging (conjugate) condition and \(A\) is then the transverse magnification; \(C=0\) is the afocal (telescopic) condition; \(A=0\) means the output height is independent of input height — the input plane sits at the front focal plane; \(D=0\) puts the output plane at the back focal plane. The effective focal length of any system is \(f_{\text{eff}}=-1/C\), so the single entry \(C\) delivers the power of arbitrarily complex optics.

The unit determinant is a genuine conservation law, not a convention. In the reduced coordinates \((y,\,u=n\theta)\), \(u\) is the canonical momentum conjugate to \(y\) in the optical Hamiltonian where the axial coordinate \(z\) plays the role of time. Paraxial elements are then linear canonical (symplectic) transformations, and for a single degree of freedom "symplectic" means precisely "\(\det=1\)". The invariant it protects is the Smith–Helmholtz quantity \(n(y\,\theta'-y'\,\theta)\) for any two rays, which is the ray-optics form of étendue / phase-space area and, through Liouville's theorem, the reason incoherent brightness cannot be increased by any passive optical train.

Common misconceptions. (i) "The lens matrix moves the ray." It does not — \(y\) is continuous through a thin lens; only the slope kicks. (ii) "A bigger determinant means more magnification." Magnification lives in \(A\) and \(D\); the determinant is pinned at \(1\) and says nothing about magnification. (iii) "ABCD gives the real image quality." It gives only the ideal Gaussian image; all aberration information has been discarded at first order.

Worked examples

Example 1 — Back focal distance of a single spherical surface. Air (\(n_1=1\)) meets glass (\(n_2=1.5\)) at a convex surface of radius \(R=+0.10\ \text{m}\). Where do incoming axis-parallel rays cross the axis inside the glass?

1
\[ P=\frac{n_2-n_1}{R}=\frac{1.5-1}{0.10\ \text{m}}=5.0\ \text{m}^{-1} \]
Surface power from Step 7. A
2
\[ \begin{pmatrix}y_2\\u_2\end{pmatrix}=\begin{pmatrix}1&0\\-P&1\end{pmatrix}\begin{pmatrix}y\\0\end{pmatrix}=\begin{pmatrix}y\\-P\,y\end{pmatrix} \]
Parallel input means \(\theta_1=0\Rightarrow u_1=0\). Refraction gives reduced slope \(u_2=-Py\). A
3
\[ \theta_2=\frac{u_2}{n_2}=-\frac{P y}{n_2},\qquad y(\ell)=y+\ell\,\theta_2=0\ \Rightarrow\ \ell=-\frac{y}{\theta_2}=\frac{n_2}{P} \]
Convert reduced slope back to a real angle inside the glass, then propagate a distance \(\ell\) until the height vanishes. B
4
\[ \ell=\frac{n_2}{P}=\frac{1.5}{5.0\ \text{m}^{-1}}=0.30\ \text{m} \]
Insert numbers. A
\[ \boxed{\ \ell_{\text{BFD}} = \dfrac{n_2}{P}=0.30\ \text{m}\ } \]

Reading. Parallel light focuses \(0.30\ \text{m}\) beyond the surface, inside the glass; note the real focal distance is stretched by the factor \(n_2\) relative to \(1/P=0.20\ \text{m}\), because rays travel more slowly (bend less per length) in the denser medium.

Units check. \(n_2/P\) is \(1/(\text{m}^{-1})=\text{m}\) ✓.

Example 2 — Imaging through a thin lens by matrix. A thin lens in air, \(f=0.20\ \text{m}\), with a real object at \(s=0.30\ \text{m}\). Find the image distance and magnification from the system matrix (air throughout, so \(u=\theta\)).

1
\[ M=T(s')\,L\,T(s)=\begin{pmatrix}1&s'\\0&1\end{pmatrix}\begin{pmatrix}1&0\\-1/f&1\end{pmatrix}\begin{pmatrix}1&s\\0&1\end{pmatrix} \]
Build the chain object-plane \(\to\) lens \(\to\) image-plane, first element rightmost (Step 12). A
2
\[ L\,T(s)=\begin{pmatrix}1&s\\-1/f&\,1-s/f\end{pmatrix},\qquad M=\begin{pmatrix}1-\dfrac{s'}{f} & s+s'\!\left(1-\dfrac{s}{f}\right)\\[6pt] -\dfrac{1}{f} & 1-\dfrac{s}{f}\end{pmatrix} \]
Multiply symbolically, right to left. A
3
\[ B=0:\quad s+s'\!\left(1-\frac{s}{f}\right)=0\ \Rightarrow\ s'=\frac{s}{\,s/f-1\,} \]
Impose the imaging condition (output height independent of input slope) and solve for \(s'\). B
4
\[ s'=\frac{0.30}{\,0.30/0.20-1\,}=\frac{0.30}{0.5}=0.60\ \text{m},\qquad m=A=1-\frac{s'}{f}=1-\frac{0.60}{0.20}=-2 \]
Insert numbers; the surviving entry \(A\) is the transverse magnification. A
\[ \boxed{\ s'=0.60\ \text{m},\qquad m=-2\ } \]

Reading. The image is real, \(0.60\ \text{m}\) past the lens, inverted, and twice the size. This reproduces \(1/s'+1/s=1/f\) (\(1/0.6+1/0.3=1.667+3.333=5=1/0.2\)) and \(m=-s'/s=-2\), obtained purely from \(B=0\) and \(A\).

Units check. \(s'\) in metres ✓; \(m\) dimensionless ✓; \(\det M=AD-BC=(-2)(1-1.5)-0\cdot(-5)=(-2)(-0.5)=1\) ✓.

Problems
  1. (A) Reduced propagation. Write the ray-transfer matrix for a \(0.50\ \text{m}\) path through water (\(n=1.33\)) in reduced coordinates.
    Solution The propagation matrix is \(T=\left(\begin{smallmatrix}1&d/n\\0&1\end{smallmatrix}\right)\) with \(d/n=0.50/1.33=0.376\ \text{m}\). So \(T=\left(\begin{smallmatrix}1&0.376\\0&1\end{smallmatrix}\right)\), \(\det=1\). A ray entering at height \(0\) with reduced slope \(u\) exits at height \(0.376\,u\) metres.
  2. (A) Surface refraction matrix. A ray in air meets a concave glass surface, \(n_2=1.52\), \(R=-0.25\ \text{m}\). Give \(\mathcal R\).
    Solution \(P=(n_2-n_1)/R=(1.52-1)/(-0.25)=-2.08\ \text{m}^{-1}\). Then \(\mathcal R=\left(\begin{smallmatrix}1&0\\-P&1\end{smallmatrix}\right)=\left(\begin{smallmatrix}1&0\\+2.08&1\end{smallmatrix}\right)\), \(\det=1\). Negative power: a concave-into-glass surface is diverging for the reduced slope.
  3. (B) Lensmaker. A biconvex crown-glass lens (\(n_\ell=1.52\)) in air has \(R_1=+0.20\ \text{m}\), \(R_2=-0.30\ \text{m}\). Find \(f\) and write \(L\).
    Solution \(P=(n_\ell-1)(1/R_1-1/R_2)=0.52\,(1/0.20-1/(-0.30))=0.52\,(5.00+3.333)=0.52\times8.333=4.33\ \text{m}^{-1}\). Thus \(f=1/P=0.231\ \text{m}\) and \(L=\left(\begin{smallmatrix}1&0\\-4.33&1\end{smallmatrix}\right)\), \(\det=1\).
  4. (B) Two-lens system. Lenses \(f_1=0.10\ \text{m}\) and \(f_2=0.15\ \text{m}\) are separated by \(t=0.05\ \text{m}\) in air. Find the system matrix \(M=L_2\,T(t)\,L_1\) and the effective focal length.
    Solution \(T(t)L_1=\left(\begin{smallmatrix}1&t\\0&1\end{smallmatrix}\right)\left(\begin{smallmatrix}1&0\\-1/f_1&1\end{smallmatrix}\right)=\left(\begin{smallmatrix}1-t/f_1 & t\\-1/f_1&1\end{smallmatrix}\right)=\left(\begin{smallmatrix}0.5&0.05\\-10&1\end{smallmatrix}\right)\). Then \(M=L_2\,(T L_1)=\left(\begin{smallmatrix}1&0\\-1/f_2&1\end{smallmatrix}\right)\left(\begin{smallmatrix}0.5&0.05\\-10&1\end{smallmatrix}\right)=\left(\begin{smallmatrix}0.5&0.05\\-1/0.15\cdot0.5-10 & -1/0.15\cdot0.05+1\end{smallmatrix}\right)\). Compute \(C=-(0.5)/0.15-10=-3.333-10=-13.33\ \text{m}^{-1}\), \(D=-0.05/0.15+1=-0.333+1=0.667\). So \(M=\left(\begin{smallmatrix}0.5&0.05\\-13.33&0.667\end{smallmatrix}\right)\). Check \(\det=0.5\times0.667-0.05\times(-13.33)=0.333+0.667=1\) ✓. Effective focal length \(f_{\text{eff}}=-1/C=1/13.33=0.075\ \text{m}\), matching \(1/f=1/f_1+1/f_2-t/(f_1 f_2)=10+6.667-0.05/0.015=16.667-3.333=13.33\ \text{m}^{-1}\).
  5. (C) Determinant / invariant. Two rays enter a system with states \((y_1,u_1)\) and \((y_2,u_2)\). Show the quantity \(\Lambda=y_1 u_2 - y_2 u_1\) is preserved by any element with \(\det=1\), and evaluate it after Example 2's lens for input rays \((0,\,0.02)\) and \((0.01,\,0)\) (SI, radians).
    Solution Under \(\mathbf V\mapsto M\mathbf V\), the pair of column vectors forms a \(2\times2\) matrix \(\left[\mathbf V_1\ \mathbf V_2\right]\) whose determinant is exactly \(\Lambda=y_1u_2-y_2u_1\). After the element, the matrix becomes \(M\left[\mathbf V_1\ \mathbf V_2\right]\), so its determinant is \(\det M\cdot\Lambda=1\cdot\Lambda=\Lambda\); the invariant is unchanged — this is the Smith–Helmholtz (Lagrange) invariant. Numerically, before the lens \(\Lambda=y_1u_2-y_2u_1=(0)(0)-(0.01)(0.02)=-2.0\times10^{-4}\ \text{m}\). Applying the thin lens \(L=\left(\begin{smallmatrix}1&0\\-5&1\end{smallmatrix}\right)\) (\(1/f=5\)): ray 1 \(\to(0,\,0.02)\) (since \(y=0\) is unkicked), ray 2 \(\to(0.01,\,-5\times0.01+0)=(0.01,-0.05)\). New \(\Lambda=(0)(-0.05)-(0.01)(0.02)=-2.0\times10^{-4}\ \text{m}\) — identical, confirming invariance.