Paraxial Ray-Transfer (ABCD) Matrices
Statement
Representing a paraxial ray in a meridional plane by the two-component state \(\mathbf{V}=(y,\,n\theta)^{\mathsf T}\) — transverse height \(y\) and reduced slope \(u=n\theta\) — we derive the three elementary ray-transfer matrices: free propagation through a length \(d\) of index \(n\), refraction at a spherical interface of radius \(R\) between indices \(n_1,n_2\), and the thin lens. We show that traversal of any sequence of elements is the ordered matrix product \(\mathbf{V}_{\text{out}}=M_N\cdots M_2 M_1\,\mathbf{V}_{\text{in}}\), and that in reduced coordinates every elementary matrix — and hence every system matrix — has determinant exactly \(+1\).
Why it matters
The ABCD formalism turns geometrical optics into linear algebra: an entire instrument, from a single interface to a multi-element zoom or a laser resonator, collapses to one \(2\times2\) matrix that can be multiplied, inverted, and iterated. Imaging conditions, cardinal points, and system power are then read directly off the four entries rather than traced ray by ray.
The unit determinant is not a bookkeeping accident. It encodes the conservation of a phase-space area — the Smith–Helmholtz (Lagrange) invariant — which is the ray-optics shadow of Liouville's theorem and of étendue conservation. That single constraint links optical throughput, brightness limits, and why you cannot losslessly concentrate incoherent light beyond the source's phase-space volume.
Assumptions
Derivation
Result
Reading. A ray's transverse state \((y,\,n\theta)\) transforms linearly through any paraxial system. Propagation is a shear that tilts height without changing reduced slope; refraction and the thin lens are focusing kicks that change slope in proportion to height (with strength the power \(P=1/f\)) without moving the ray. The generic entries \(\left(\begin{smallmatrix}A&B\\C&D\end{smallmatrix}\right)\) mean: \(A\) is transverse (angular) magnification bookkeeping, \(B\) has the dimension of a reduced length, \(C=-P\) is minus the system power, and \(D\) is angular magnification bookkeeping. The constraint \(AD-BC=1\) is the conserved phase-space area.
Units check. With \([y]=\text{m}\) and \([u]=[n\theta]=1\) (dimensionless), the entries carry \([A]=1,\ [B]=\text{m},\ [C]=\text{m}^{-1},\ [D]=1\). Then \(d/n\) is a length ✓, \(P=(n_2-n_1)/R\) and \(1/f\) are \(\text{m}^{-1}\) (dioptres) ✓, and \(\det=AD-BC\) has units \(1-(\text{m})(\text{m}^{-1})=1\), a pure number, consistent with equalling \(1\) ✓.
Limiting cases
- Plane interface (\(R\to\infty\)): \(P\to0\), \(\mathcal R\to I\) in reduced coordinates — a flat surface bends no reduced slope (all bending is hidden in the \(n\theta\) rescaling; in bare angles it still refracts by \(n_1/n_2\)).
- Zero-power (afocal) element (\(f\to\infty\)): \(L\to I\); a "window" or a matched afocal telescope leaves reduced slope untouched.
- Zero gap (\(d\to0\)): \(T\to I\); stacking elements with no space between them just multiplies their kicks (Step 9).
- Same medium in and out: even in the bare-angle \((y,\theta)\) convention, \(\det M=n_{\text{in}}/n_{\text{out}}=1\), because the internal index jumps telescope.
- Imaging condition (\(B=0\)): input height maps to output height independent of input slope — a point images to a point, and \(A\) becomes the transverse magnification.
Breaks when
- Large angles / large apertures. Once \(\sin\theta\neq\theta\) to the accuracy required, the true map is nonlinear: spherical aberration, coma, astigmatism, and distortion appear at third order in \((y,\theta)\) and no single \(2\times2\) matrix reproduces the wavefront. The matrix gives only the Gaussian (first-order) image.
- Skew rays and non-rotationally-symmetric systems. Cylindrical lenses, tilted or decentred elements, and skew rays couple the two transverse planes; the correct object is a \(4\times4\) symplectic (or \(6\times6\)) matrix, and the meridional \(2\times2\) reduction is invalid.
- Gradient-index or otherwise inhomogeneous media. When \(n=n(\mathbf r)\) within a segment, propagation is not a straight-line shear — the ray obeys \(\tfrac{d}{ds}\!\left(n\frac{d\mathbf r}{ds}\right)=\nabla n\), and the segment's matrix must be integrated from that equation.
- Diffraction-dominated / small features. Ray optics itself fails when apertures approach the wavelength; the ABCD entries then survive only as inputs to the diffraction integral (the ABCD Gaussian-beam / Collins-integral extension), not as a complete description.
Failure modes
- Multiplying in traversal order. Writing \(M=M_1 M_2\cdots M_N\) (first element leftmost). The ray meets element 1 first, so it must be rightmost: \(M=M_N\cdots M_1\). Matrices do not commute, so this silently gives the wrong system.
- Mixing angle conventions. Deriving \(\mathcal R\) with bare \(\theta\) (giving \(\det=n_1/n_2\)) but then propagating with the reduced-\(u\) matrix, or vice versa. Pick one convention and use its full set consistently.
- Dropping the index in propagation. Using \(d\) instead of \(d/n\) for a segment inside glass or water; the reduced propagation length is \(d/n\).
- Sign of \(R\). Forgetting that the second surface of a biconvex lens has \(R_2<0\); using \(|R|\) makes both surfaces "converging" and doubles the apparent power.
- Reading \(C\) as the power. The entry is \(C=-P=-1/f\); students quote a negative focal length for a converging lens by dropping the sign.
- Treating a thick lens as thin. Omitting the internal \(T(t/n_\ell)\) between the two surface matrices when the axial thickness \(t\) is not negligible.
Discussion
The formalism separates optics cleanly into two moves: shears (propagation) that change where a ray is without changing where it is going, and kicks (refraction, lenses, mirrors) that change where it is going without moving it. Every paraxial instrument is a word in these two generators. Because the generators are triangular with unit diagonal, their commutators are non-trivial — this non-commutativity is exactly why lens spacing matters and why a telescope is not the same optic as its elements reversed.
The system entries carry direct physical meaning. \(B=0\) is the imaging (conjugate) condition and \(A\) is then the transverse magnification; \(C=0\) is the afocal (telescopic) condition; \(A=0\) means the output height is independent of input height — the input plane sits at the front focal plane; \(D=0\) puts the output plane at the back focal plane. The effective focal length of any system is \(f_{\text{eff}}=-1/C\), so the single entry \(C\) delivers the power of arbitrarily complex optics.
The unit determinant is a genuine conservation law, not a convention. In the reduced coordinates \((y,\,u=n\theta)\), \(u\) is the canonical momentum conjugate to \(y\) in the optical Hamiltonian where the axial coordinate \(z\) plays the role of time. Paraxial elements are then linear canonical (symplectic) transformations, and for a single degree of freedom "symplectic" means precisely "\(\det=1\)". The invariant it protects is the Smith–Helmholtz quantity \(n(y\,\theta'-y'\,\theta)\) for any two rays, which is the ray-optics form of étendue / phase-space area and, through Liouville's theorem, the reason incoherent brightness cannot be increased by any passive optical train.
Common misconceptions. (i) "The lens matrix moves the ray." It does not — \(y\) is continuous through a thin lens; only the slope kicks. (ii) "A bigger determinant means more magnification." Magnification lives in \(A\) and \(D\); the determinant is pinned at \(1\) and says nothing about magnification. (iii) "ABCD gives the real image quality." It gives only the ideal Gaussian image; all aberration information has been discarded at first order.
Worked examples
Example 1 — Back focal distance of a single spherical surface. Air (\(n_1=1\)) meets glass (\(n_2=1.5\)) at a convex surface of radius \(R=+0.10\ \text{m}\). Where do incoming axis-parallel rays cross the axis inside the glass?
Reading. Parallel light focuses \(0.30\ \text{m}\) beyond the surface, inside the glass; note the real focal distance is stretched by the factor \(n_2\) relative to \(1/P=0.20\ \text{m}\), because rays travel more slowly (bend less per length) in the denser medium.
Units check. \(n_2/P\) is \(1/(\text{m}^{-1})=\text{m}\) ✓.
Example 2 — Imaging through a thin lens by matrix. A thin lens in air, \(f=0.20\ \text{m}\), with a real object at \(s=0.30\ \text{m}\). Find the image distance and magnification from the system matrix (air throughout, so \(u=\theta\)).
Reading. The image is real, \(0.60\ \text{m}\) past the lens, inverted, and twice the size. This reproduces \(1/s'+1/s=1/f\) (\(1/0.6+1/0.3=1.667+3.333=5=1/0.2\)) and \(m=-s'/s=-2\), obtained purely from \(B=0\) and \(A\).
Units check. \(s'\) in metres ✓; \(m\) dimensionless ✓; \(\det M=AD-BC=(-2)(1-1.5)-0\cdot(-5)=(-2)(-0.5)=1\) ✓.
Problems
- (A) Reduced propagation. Write the ray-transfer matrix for a \(0.50\ \text{m}\) path through water (\(n=1.33\)) in reduced coordinates.
Solution
The propagation matrix is \(T=\left(\begin{smallmatrix}1&d/n\\0&1\end{smallmatrix}\right)\) with \(d/n=0.50/1.33=0.376\ \text{m}\). So \(T=\left(\begin{smallmatrix}1&0.376\\0&1\end{smallmatrix}\right)\), \(\det=1\). A ray entering at height \(0\) with reduced slope \(u\) exits at height \(0.376\,u\) metres. - (A) Surface refraction matrix. A ray in air meets a concave glass surface, \(n_2=1.52\), \(R=-0.25\ \text{m}\). Give \(\mathcal R\).
Solution
\(P=(n_2-n_1)/R=(1.52-1)/(-0.25)=-2.08\ \text{m}^{-1}\). Then \(\mathcal R=\left(\begin{smallmatrix}1&0\\-P&1\end{smallmatrix}\right)=\left(\begin{smallmatrix}1&0\\+2.08&1\end{smallmatrix}\right)\), \(\det=1\). Negative power: a concave-into-glass surface is diverging for the reduced slope. - (B) Lensmaker. A biconvex crown-glass lens (\(n_\ell=1.52\)) in air has \(R_1=+0.20\ \text{m}\), \(R_2=-0.30\ \text{m}\). Find \(f\) and write \(L\).
Solution
\(P=(n_\ell-1)(1/R_1-1/R_2)=0.52\,(1/0.20-1/(-0.30))=0.52\,(5.00+3.333)=0.52\times8.333=4.33\ \text{m}^{-1}\). Thus \(f=1/P=0.231\ \text{m}\) and \(L=\left(\begin{smallmatrix}1&0\\-4.33&1\end{smallmatrix}\right)\), \(\det=1\). - (B) Two-lens system. Lenses \(f_1=0.10\ \text{m}\) and \(f_2=0.15\ \text{m}\) are separated by \(t=0.05\ \text{m}\) in air. Find the system matrix \(M=L_2\,T(t)\,L_1\) and the effective focal length.
Solution
\(T(t)L_1=\left(\begin{smallmatrix}1&t\\0&1\end{smallmatrix}\right)\left(\begin{smallmatrix}1&0\\-1/f_1&1\end{smallmatrix}\right)=\left(\begin{smallmatrix}1-t/f_1 & t\\-1/f_1&1\end{smallmatrix}\right)=\left(\begin{smallmatrix}0.5&0.05\\-10&1\end{smallmatrix}\right)\). Then \(M=L_2\,(T L_1)=\left(\begin{smallmatrix}1&0\\-1/f_2&1\end{smallmatrix}\right)\left(\begin{smallmatrix}0.5&0.05\\-10&1\end{smallmatrix}\right)=\left(\begin{smallmatrix}0.5&0.05\\-1/0.15\cdot0.5-10 & -1/0.15\cdot0.05+1\end{smallmatrix}\right)\). Compute \(C=-(0.5)/0.15-10=-3.333-10=-13.33\ \text{m}^{-1}\), \(D=-0.05/0.15+1=-0.333+1=0.667\). So \(M=\left(\begin{smallmatrix}0.5&0.05\\-13.33&0.667\end{smallmatrix}\right)\). Check \(\det=0.5\times0.667-0.05\times(-13.33)=0.333+0.667=1\) ✓. Effective focal length \(f_{\text{eff}}=-1/C=1/13.33=0.075\ \text{m}\), matching \(1/f=1/f_1+1/f_2-t/(f_1 f_2)=10+6.667-0.05/0.015=16.667-3.333=13.33\ \text{m}^{-1}\). - (C) Determinant / invariant. Two rays enter a system with states \((y_1,u_1)\) and \((y_2,u_2)\). Show the quantity \(\Lambda=y_1 u_2 - y_2 u_1\) is preserved by any element with \(\det=1\), and evaluate it after Example 2's lens for input rays \((0,\,0.02)\) and \((0.01,\,0)\) (SI, radians).
Solution
Under \(\mathbf V\mapsto M\mathbf V\), the pair of column vectors forms a \(2\times2\) matrix \(\left[\mathbf V_1\ \mathbf V_2\right]\) whose determinant is exactly \(\Lambda=y_1u_2-y_2u_1\). After the element, the matrix becomes \(M\left[\mathbf V_1\ \mathbf V_2\right]\), so its determinant is \(\det M\cdot\Lambda=1\cdot\Lambda=\Lambda\); the invariant is unchanged — this is the Smith–Helmholtz (Lagrange) invariant. Numerically, before the lens \(\Lambda=y_1u_2-y_2u_1=(0)(0)-(0.01)(0.02)=-2.0\times10^{-4}\ \text{m}\). Applying the thin lens \(L=\left(\begin{smallmatrix}1&0\\-5&1\end{smallmatrix}\right)\) (\(1/f=5\)): ray 1 \(\to(0,\,0.02)\) (since \(y=0\) is unkicked), ray 2 \(\to(0.01,\,-5\times0.01+0)=(0.01,-0.05)\). New \(\Lambda=(0)(-0.05)-(0.01)(0.02)=-2.0\times10^{-4}\ \text{m}\) — identical, confirming invariance.