Fermat's Principle from the Eikonal Equation
Statement
Starting from the eikonal equation \( \lvert \nabla S \rvert^2 = n^2 \) of geometrical optics, we show that the light rays — the curves everywhere tangent to \( \nabla S \) — obey the vector ray equation \( \dfrac{d}{ds}\!\left( n\, \dfrac{d\mathbf{r}}{ds} \right) = \nabla n \), where \( s \) is arc length and \( n(\mathbf{r}) \) the refractive index. We then show this same equation is the Euler–Lagrange equation of the optical path length \( \int n\, ds \), so that physical ray paths make the optical path length stationary — Fermat's principle.
Why it matters
Fermat's principle is the variational bedrock of all ray optics: it fixes lens design, fibre-optic waveguiding, atmospheric refraction, mirages and gravitational-lens ray tracing, all from a single stationarity condition \( \delta \int n\, ds = 0 \). Deriving it from the eikonal equation shows it is not an independent postulate but a consequence of the short-wavelength limit of the wave equation.
The derivation also exposes the deep optical–mechanical analogy: the eikonal \( S \) plays the role of Hamilton's characteristic function, \( \lvert \nabla S \rvert = n \) is a Hamilton–Jacobi equation, and \( \int n\, ds \) is the optical analogue of Maupertuis' action \( \int p\, dq \). This bridge is what later let de Broglie and Schrödinger read wave mechanics off optics.
Assumptions
Derivation
Result
Reading. The left form is a local force-like law: the "optical momentum" \( n\,\hat{\mathbf{t}} \) changes along the ray at a rate set by the gradient of the index. A ray bends toward the region of higher \( n \), because \( \nabla n \) supplies the curving push. The right form is the global statement — Fermat's principle — that the actual ray makes the optical path length (equivalently the travel time \( \int n\, ds / c \)) stationary between its endpoints. The two are the same physics: the ray equation is the Euler–Lagrange equation of the optical path integral.
Units check. \( \hat{\mathbf{t}} = d\mathbf{r}/ds \) is dimensionless (metre per metre) and \( n \) is dimensionless, so \( n\,\hat{\mathbf{t}} \) is dimensionless; \( d/ds \) carries \( \mathrm{m^{-1}} \), giving the left side units of \( \mathrm{m^{-1}} \). The right side \( \nabla n \) is (dimensionless)\( / \)metre \( = \mathrm{m^{-1}} \). Both sides balance at \( \mathrm{m^{-1}} \). The integral \( \int n\, ds \) has units of metres (an optical path length), consistent with a length that is stationary.
Limiting cases
- Homogeneous medium (\( n=\text{const} \)): \( \nabla n = 0 \Rightarrow d^2\mathbf{r}/ds^2 = 0 \), so rays are straight lines and \( \int n\, ds = n\,L \) is minimised by the shortest geometric path.
- Stratified medium \( n=n(y) \): the \( x \)-component of \( n\,\hat{\mathbf{t}} \) is conserved, giving the continuum Snell invariant \( n\sin\theta = \text{const} \) along every ray.
- Paraxial / weak-gradient regime: linearising \( \nabla n \) about a reference ray gives Hooke-like transverse equations \( d^2 x/dz^2 = -g^2 x \), i.e. sinusoidal rays (graded-index lenses and fibres).
- Vacuum limit (\( n\to 1 \)): the optical path length reduces to geometric length and rays are the straight null geodesics of flat space.
Breaks when
- Wavelength not small compared with the scale of variation of \( n \): the eikonal expansion fails, diffraction and interference dominate, and no ray-equation solution captures the field (e.g. light through a sub-wavelength aperture).
- Sharp interfaces / index discontinuities: \( \nabla n \) is a delta function, the local ray equation is meaningless at the surface, and one must instead impose Snell/Fresnel boundary conditions across it.
- Anisotropic (birefringent) media: the ray tangent \( d\mathbf{r}/ds \) no longer aligns with \( \nabla S \); \( n \) is a tensor and the scalar derivation collapses.
- At caustics and foci: neighbouring rays cross, the geometrical-optics amplitude diverges, and the ray description must be repaired by diffraction (Airy/uniform) corrections even though the ray paths themselves remain formally defined.
Failure modes
- Dropping the arc-length constraint: treating \( d\mathbf{r}/ds \) as an arbitrary vector rather than a unit tangent (\( \lvert d\mathbf{r}/ds \rvert = 1 \)); this loses the eikonal normalisation and corrupts curvature estimates.
- Pulling \( n \) out of the derivative: writing \( n\, d^2\mathbf{r}/ds^2 = \nabla n \) and dropping the \( \tfrac{dn}{ds}\hat{\mathbf{t}} \) term. The correct expansion is \( n\,\kappa\,\hat{\mathbf{n}} + (\hat{\mathbf{t}}\cdot\nabla n)\hat{\mathbf{t}} = \nabla n \).
- Using the full \( \nabla n \) for curvature: ray curvature depends only on the component of \( \nabla n \) perpendicular to the ray, \( \kappa = \lvert\nabla_\perp n\rvert / n \); the tangential part only rescales the parametrisation.
- Parametrising by \( z \) without the Jacobian: forgetting the \( \lvert\mathbf{r}'\rvert \) factor in \( F = n\lvert\mathbf{r}'\rvert \) when \( \sigma\neq s \), which yields a wrong Euler–Lagrange equation.
- Confusing "stationary" with "minimum": Fermat's principle gives a stationary optical path; near foci and in mirror geometries the true ray can be a saddle or a local maximum, not a minimum.
Discussion
The ray equation \( \frac{d}{ds}(n\,d\mathbf{r}/ds) = \nabla n \) is a statement of "optical inertia". The quantity \( \mathbf{p} = n\,\hat{\mathbf{t}} = \nabla S \) is the optical momentum (the ray analogue of a mechanical momentum), and \( \nabla n \) is the force that turns it. Where the index is uniform the momentum is constant and the ray coasts in a straight line; where \( n \) increases the momentum is deflected toward the gradient, so light curves into the optically denser region. This is exactly why mirages, atmospheric looming and the twinkling of stars occur: vertical index gradients bend otherwise-straight rays.
Recasting the result as \( \delta\int n\, ds = 0 \) reveals the variational structure. Fermat's principle is a genuine stationary-action principle with "action" \( \int n\, ds \), and the eikonal \( S \) is its Hamilton characteristic function: \( \lvert\nabla S\rvert = n \) is precisely the Hamilton–Jacobi equation of a system with Hamiltonian \( H = \tfrac12(\lvert\mathbf{p}\rvert^2 - n^2) = 0 \). Rays are the characteristics (bicharacteristics) of this equation, and wavefronts \( S=\text{const} \) are the surfaces they thread orthogonally in an isotropic medium.
The symmetry content is worth naming explicitly. Whenever \( n \) is independent of a coordinate, the conjugate component of \( n\,\hat{\mathbf{t}} \) is conserved — a Noether conservation law for the optical action. Translational symmetry in \( x \) for a stratified medium gives \( n\sin\theta=\text{const} \) (Snell), and rotational symmetry for a spherically symmetric \( n(r) \) gives Bouguer's invariant \( n\, r\sin\phi = \text{const} \). Fermat's principle thus organises ray optics the way least action organises mechanics.
At the deepest level this derivation is the classical limit of a wave theory taken twice over. The eikonal ansatz \( \psi \sim A\, e^{i k_0 S} \) is the optical WKB expansion; \( \lvert\nabla S\rvert = n \) is its leading order and the ray equation its characteristic flow, exactly paralleling how the quantum WKB phase \( e^{iW/\hbar} \) yields \( \lvert\nabla W\rvert = \sqrt{2m(E-V)} \) and Newtonian trajectories as \( \hbar\to 0 \). The optical path length \( \int n\, ds \) maps onto Maupertuis' abbreviated action \( \int \mathbf{p}\cdot d\mathbf{q} \), with \( n \leftrightarrow \lvert\mathbf{p}\rvert \). Hamiltonian optics — ray transfer matrices, symplectic phase space, Liouville's theorem for étendue — is this analogy made into machinery.
Common misconceptions. Fermat's principle is often stated as "light takes the path of least time." It is really the path of stationary optical path length: at a focus or off a curved mirror the ray can be a saddle point or even a local maximum. Nor does light "try out" all paths and select one — the stationary path is simply the constructive-interference condition of the underlying wave, recovered here as the eikonal characteristic. Finally, the bending of a ray is not caused by \( n \) being large but by \( n \) varying: a uniform slab of glass, however high its index, bends nothing internally.
Worked examples
Example 1 — Optical path length and straight rays in a uniform slab.
Reading. Although the glass is only \( 4.0\ \mathrm{cm} \) thick, light traverses \( 6.0\ \mathrm{cm} \) of optical path — the extra \( 2.0\ \mathrm{cm} \) is the phase delay that makes the slab act like a longer stretch of vacuum, and Fermat's principle confirms this straight line is the stationary path.
Example 2 — Sinusoidal rays in a graded-index (GRIN) fibre.
Reading. Rays launched off-axis oscillate about the fibre axis and return to focus every \( 25\ \mathrm{mm} \). This self-focusing — a direct consequence of \( \nabla n \) pointing toward the axis in the ray equation — is exactly how GRIN rod lenses and multimode graded-index fibres guide light without discrete refracting surfaces.
Problems
- Optical path across two slabs. Light passes normally through \( 3.0\ \mathrm{cm} \) of water (\( n=1.33 \)) then \( 2.0\ \mathrm{cm} \) of glass (\( n=1.52 \)). Find the total optical path length and the transit time.
Solution
Rays are straight (uniform slabs, \( \nabla n = 0 \)). \( \text{OPL} = n_1 L_1 + n_2 L_2 = 1.33(3.0) + 1.52(2.0) = 3.99 + 3.04 = 7.03\ \mathrm{cm} \). Time \( t = \text{OPL}/c = 0.0703/(3.0\times10^8) = 2.34\times10^{-10}\ \mathrm{s} \approx 0.23\ \mathrm{ns} \). - Snell's invariant from the ray equation. For a stratified medium \( n=n(y) \) with rays confined to the \( xy \)-plane, show that \( n\sin\theta \) is constant along a ray, where \( \theta \) is the angle from the \( y \)-axis.
Solution
The \( x \)-component of the ray equation is \( \frac{d}{ds}\!\left(n\frac{dx}{ds}\right)=\frac{\partial n}{\partial x}=0 \) since \( n \) depends only on \( y \). Hence \( n\,\frac{dx}{ds}=\text{const} \). But \( \frac{dx}{ds}=\sin\theta \) (the tangent's angle from the \( y \)-axis), so \( n\sin\theta = \text{const} \). This is Snell's law in continuous form; at a discrete interface it becomes \( n_1\sin\theta_1 = n_2\sin\theta_2 \). - Radius of curvature of a horizontal ray (mirage). Near the ground the index varies vertically with \( \lvert dn/dz \rvert = 1.0\times10^{-7}\ \mathrm{m^{-1}} \) and \( n\approx 1.000 \). For a ray travelling horizontally, find its radius of curvature. Compare with Earth's radius \( R_\oplus = 6.37\times10^{6}\ \mathrm{m} \).
Solution
Write the ray equation as \( n\,\kappa\,\hat{\mathbf{n}} + (\hat{\mathbf{t}}\cdot\nabla n)\hat{\mathbf{t}} = \nabla n \). The normal component gives \( n\kappa = \hat{\mathbf{n}}\cdot\nabla n = \lvert\nabla_\perp n\rvert \). For a horizontal ray the perpendicular gradient is the vertical one, so \( \kappa = \lvert dn/dz\rvert / n \) and \( R = 1/\kappa = n/\lvert dn/dz\rvert = 1.000/(1.0\times10^{-7}) = 1.0\times10^{7}\ \mathrm{m} \). This is \( \approx 1.6\,R_\oplus \); because the ray's curvature is comparable to Earth's, strong gradients let light hug the surface and produce mirages and looming. - GRIN pitch and focal spacing. A GRIN rod has \( n(r)=n_0(1-\tfrac12 g^2 r^2) \) with \( n_0=1.60 \) and \( g=0.40\ \mathrm{mm^{-1}} \). Find (a) the ray pitch \( \Lambda \) and (b) the shortest rod length that images an input face onto the output face (a "quarter-pitch" lens).
Solution
Paraxial transverse equation \( d^2 r/dz^2 = -g^2 r \Rightarrow \Lambda = 2\pi/g = 2\pi/0.40 = 15.7\ \mathrm{mm} \). A quarter-pitch lens converts a point at the input into collimated/imaged output after a quarter period: \( L = \Lambda/4 = 15.7/4 = 3.93\ \mathrm{mm} \). (Half-pitch, \( \Lambda/2 = 7.85\ \mathrm{mm} \), reimages a point to a point.) - Bouguer's theorem (spherical symmetry). For a spherically symmetric medium \( n=n(r) \), show that \( n\, r \sin\phi \) is conserved along a ray, where \( \phi \) is the angle between the ray and the local radial direction. (Hint: use the symmetry / angular-momentum structure of \( n\,\hat{\mathbf{t}} \).)
Solution
Consider the vector \( \mathbf{L} = \mathbf{r}\times(n\,\hat{\mathbf{t}}) \). Its arc-length derivative is \( \frac{d\mathbf{L}}{ds}=\frac{d\mathbf{r}}{ds}\times(n\hat{\mathbf{t}}) + \mathbf{r}\times\frac{d}{ds}(n\hat{\mathbf{t}}) \). The first term is \( \hat{\mathbf{t}}\times(n\hat{\mathbf{t}})=\mathbf{0} \). By the ray equation the second is \( \mathbf{r}\times\nabla n \); for \( n=n(r) \), \( \nabla n = (dn/dr)\hat{\mathbf{r}} \parallel \mathbf{r} \), so \( \mathbf{r}\times\nabla n=\mathbf{0} \). Hence \( \mathbf{L} \) is constant, and its magnitude \( \lvert\mathbf{L}\rvert = r\,n\,\lvert\hat{\mathbf{r}}\times\hat{\mathbf{t}}\rvert = n\,r\sin\phi \) is conserved. This is Bouguer's theorem, the optical analogue of angular-momentum conservation, and it governs ray tracing in lens atmospheres and the solar chromosphere.