physics2u
Tier
⌕ Search ⌘K
Derivation

Fermat's Principle from the Eikonal Equation

D-197 Home PU-206 Threads light · symmetry Depends on The Eikonal Equation as the Short-Wavelength Limit
Statement

Starting from the eikonal equation \( \lvert \nabla S \rvert^2 = n^2 \) of geometrical optics, we show that the light rays — the curves everywhere tangent to \( \nabla S \) — obey the vector ray equation \( \dfrac{d}{ds}\!\left( n\, \dfrac{d\mathbf{r}}{ds} \right) = \nabla n \), where \( s \) is arc length and \( n(\mathbf{r}) \) the refractive index. We then show this same equation is the Euler–Lagrange equation of the optical path length \( \int n\, ds \), so that physical ray paths make the optical path length stationary — Fermat's principle.

Why it matters

Fermat's principle is the variational bedrock of all ray optics: it fixes lens design, fibre-optic waveguiding, atmospheric refraction, mirages and gravitational-lens ray tracing, all from a single stationarity condition \( \delta \int n\, ds = 0 \). Deriving it from the eikonal equation shows it is not an independent postulate but a consequence of the short-wavelength limit of the wave equation.

The derivation also exposes the deep optical–mechanical analogy: the eikonal \( S \) plays the role of Hamilton's characteristic function, \( \lvert \nabla S \rvert = n \) is a Hamilton–Jacobi equation, and \( \int n\, ds \) is the optical analogue of Maupertuis' action \( \int p\, dq \). This bridge is what later let de Broglie and Schrödinger read wave mechanics off optics.

Assumptions
Geometrical-optics (short-wavelength) limit.The wavelength is negligible compared with the scale over which \( n \) varies, so the eikonal equation holds. If dropped, diffraction and interference terms survive and rays are not well defined. Isotropic, scalar refractive index.\( n(\mathbf{r}) \) is a scalar, so the ray tangent is parallel to \( \nabla S \). If dropped (birefringent/anisotropic media), the ray direction \( d\mathbf{r}/ds \) and the wave-normal \( \nabla S \) no longer coincide and \( n \) becomes a tensor. Smoothly varying \( n(\mathbf{r}) \).\( n \) is differentiable, so \( \nabla n \) exists pointwise. If dropped (a sharp interface), the local ray equation is undefined there and one matches solutions across the boundary with Snell's law instead. Monochromatic, non-dispersive treatment.\( n \) is evaluated at one fixed frequency. If dropped, chromatic effects split phase and group behaviour and each colour follows a different ray. Static medium.\( n \) has no explicit time dependence. If dropped (moving or time-varying media), the eikonal equation acquires extra terms and rays are no longer stationary curves of a fixed functional.
Derivation
1
\[ \lvert \nabla S(\mathbf{r}) \rvert^2 = n(\mathbf{r})^2 \]
Prior result (eikonal equation): the leading short-wavelength order of the scalar wave equation for a phase \( k_0 S \). \( S \) is the optical path (eikonal); its level surfaces are the wavefronts. A
2
\[ \hat{\mathbf{t}} \equiv \frac{d\mathbf{r}}{ds} = \frac{\nabla S}{\lvert \nabla S \rvert} = \frac{\nabla S}{n} \quad\Longrightarrow\quad n\,\frac{d\mathbf{r}}{ds} = \nabla S \]
A ray is the curve orthogonal to the wavefronts \( S = \text{const} \); in an isotropic medium its unit tangent \( \hat{\mathbf{t}} \) points along \( \nabla S \). Normalising by \( \lvert \nabla S \rvert = n \) from Step 1 fixes the coefficient. B
3
\[ \frac{d}{ds}\!\left( n\,\frac{d\mathbf{r}}{ds} \right) = \frac{d}{ds}\big( \nabla S \big) \]
Differentiate the identity of Step 2 along the ray with respect to arc length \( s \). Legal because both sides are smooth vector fields evaluated on the ray. A
4
\[ \frac{d}{ds}\big( \nabla S \big) = \left( \frac{d\mathbf{r}}{ds}\cdot \nabla \right)\nabla S = \frac{1}{n}\big( \nabla S \cdot \nabla \big)\nabla S \]
Directional (convective) derivative: for a field \( \mathbf{F}(\mathbf{r}) \) evaluated along a curve, \( d\mathbf{F}/ds = (\hat{\mathbf{t}}\cdot\nabla)\mathbf{F} \). Substitute \( \hat{\mathbf{t}} = \nabla S / n \) from Step 2. B
5
\[ \big( \nabla S \cdot \nabla \big)\nabla S = \nabla\!\left( \tfrac{1}{2}\lvert \nabla S \rvert^2 \right) - \nabla S \times \big( \nabla \times \nabla S \big) = \nabla\!\left( \tfrac{1}{2}\lvert \nabla S \rvert^2 \right) \]
Vector identity \( (\mathbf{a}\cdot\nabla)\mathbf{a} = \nabla(\tfrac12 \lvert\mathbf{a}\rvert^2) - \mathbf{a}\times(\nabla\times\mathbf{a}) \) with \( \mathbf{a}=\nabla S \). The curl of a gradient vanishes, \( \nabla\times\nabla S = \mathbf{0} \), killing the second term. C
6
\[ \nabla\!\left( \tfrac{1}{2}\lvert \nabla S \rvert^2 \right) = \nabla\!\left( \tfrac{1}{2} n^2 \right) = n\,\nabla n \]
Insert the eikonal equation \( \lvert\nabla S\rvert^2 = n^2 \) (Step 1) and differentiate the scalar \( \tfrac12 n^2 \) by the chain rule. B
7
\[ \frac{d}{ds}\!\left( n\,\frac{d\mathbf{r}}{ds} \right) = \frac{1}{n}\,\big( n\,\nabla n \big) = \nabla n \]
Chain Steps 3, 4, 5 and 6: the \( 1/n \) prefactor of Step 4 cancels the \( n \) of Step 6. This is the ray equation. A
8
\[ \delta\!\int n\, ds = 0,\qquad F(\mathbf{r},\mathbf{r}') = n(\mathbf{r})\,\lvert \mathbf{r}' \rvert,\quad \mathbf{r}'=\frac{d\mathbf{r}}{d\sigma} \]
\[ \frac{\partial F}{\partial \mathbf{r}'} = n\,\frac{\mathbf{r}'}{\lvert\mathbf{r}'\rvert} = n\,\hat{\mathbf{t}},\qquad \frac{\partial F}{\partial \mathbf{r}} = \lvert\mathbf{r}'\rvert\,\nabla n \]
Now run the argument in reverse. Write the optical path length with an arbitrary parameter \( \sigma \), so \( ds = \lvert\mathbf{r}'\rvert\, d\sigma \). Compute the two Euler–Lagrange ingredients of \( F = n\lvert\mathbf{r}'\rvert \). C
9
\[ \frac{d}{d\sigma}\!\left( n\,\hat{\mathbf{t}} \right) - \lvert\mathbf{r}'\rvert\,\nabla n = 0 \;\;\xrightarrow{\,d/d\sigma = \lvert\mathbf{r}'\rvert\,d/ds\,}\;\; \frac{d}{ds}\!\left( n\,\frac{d\mathbf{r}}{ds} \right) = \nabla n \]
The Euler–Lagrange equation \( \frac{d}{d\sigma}\frac{\partial F}{\partial\mathbf{r}'} - \frac{\partial F}{\partial\mathbf{r}} = 0 \). Converting the parameter derivative to arc length (\( d/d\sigma = \lvert\mathbf{r}'\rvert\, d/ds \)) and cancelling the common \( \lvert\mathbf{r}'\rvert \) reproduces the ray equation of Step 7 exactly. Hence the rays are precisely the stationary paths of \( \int n\, ds \). C
Result
\[ \frac{d}{ds}\!\left( n\,\frac{d\mathbf{r}}{ds} \right) = \nabla n \qquad\Longleftrightarrow\qquad \delta\!\int_{A}^{B} n\, ds = 0 \]

Reading. The left form is a local force-like law: the "optical momentum" \( n\,\hat{\mathbf{t}} \) changes along the ray at a rate set by the gradient of the index. A ray bends toward the region of higher \( n \), because \( \nabla n \) supplies the curving push. The right form is the global statement — Fermat's principle — that the actual ray makes the optical path length (equivalently the travel time \( \int n\, ds / c \)) stationary between its endpoints. The two are the same physics: the ray equation is the Euler–Lagrange equation of the optical path integral.

Units check. \( \hat{\mathbf{t}} = d\mathbf{r}/ds \) is dimensionless (metre per metre) and \( n \) is dimensionless, so \( n\,\hat{\mathbf{t}} \) is dimensionless; \( d/ds \) carries \( \mathrm{m^{-1}} \), giving the left side units of \( \mathrm{m^{-1}} \). The right side \( \nabla n \) is (dimensionless)\( / \)metre \( = \mathrm{m^{-1}} \). Both sides balance at \( \mathrm{m^{-1}} \). The integral \( \int n\, ds \) has units of metres (an optical path length), consistent with a length that is stationary.

Limiting cases
  • Homogeneous medium (\( n=\text{const} \)): \( \nabla n = 0 \Rightarrow d^2\mathbf{r}/ds^2 = 0 \), so rays are straight lines and \( \int n\, ds = n\,L \) is minimised by the shortest geometric path.
  • Stratified medium \( n=n(y) \): the \( x \)-component of \( n\,\hat{\mathbf{t}} \) is conserved, giving the continuum Snell invariant \( n\sin\theta = \text{const} \) along every ray.
  • Paraxial / weak-gradient regime: linearising \( \nabla n \) about a reference ray gives Hooke-like transverse equations \( d^2 x/dz^2 = -g^2 x \), i.e. sinusoidal rays (graded-index lenses and fibres).
  • Vacuum limit (\( n\to 1 \)): the optical path length reduces to geometric length and rays are the straight null geodesics of flat space.
Breaks when
  • Wavelength not small compared with the scale of variation of \( n \): the eikonal expansion fails, diffraction and interference dominate, and no ray-equation solution captures the field (e.g. light through a sub-wavelength aperture).
  • Sharp interfaces / index discontinuities: \( \nabla n \) is a delta function, the local ray equation is meaningless at the surface, and one must instead impose Snell/Fresnel boundary conditions across it.
  • Anisotropic (birefringent) media: the ray tangent \( d\mathbf{r}/ds \) no longer aligns with \( \nabla S \); \( n \) is a tensor and the scalar derivation collapses.
  • At caustics and foci: neighbouring rays cross, the geometrical-optics amplitude diverges, and the ray description must be repaired by diffraction (Airy/uniform) corrections even though the ray paths themselves remain formally defined.
Failure modes
  • Dropping the arc-length constraint: treating \( d\mathbf{r}/ds \) as an arbitrary vector rather than a unit tangent (\( \lvert d\mathbf{r}/ds \rvert = 1 \)); this loses the eikonal normalisation and corrupts curvature estimates.
  • Pulling \( n \) out of the derivative: writing \( n\, d^2\mathbf{r}/ds^2 = \nabla n \) and dropping the \( \tfrac{dn}{ds}\hat{\mathbf{t}} \) term. The correct expansion is \( n\,\kappa\,\hat{\mathbf{n}} + (\hat{\mathbf{t}}\cdot\nabla n)\hat{\mathbf{t}} = \nabla n \).
  • Using the full \( \nabla n \) for curvature: ray curvature depends only on the component of \( \nabla n \) perpendicular to the ray, \( \kappa = \lvert\nabla_\perp n\rvert / n \); the tangential part only rescales the parametrisation.
  • Parametrising by \( z \) without the Jacobian: forgetting the \( \lvert\mathbf{r}'\rvert \) factor in \( F = n\lvert\mathbf{r}'\rvert \) when \( \sigma\neq s \), which yields a wrong Euler–Lagrange equation.
  • Confusing "stationary" with "minimum": Fermat's principle gives a stationary optical path; near foci and in mirror geometries the true ray can be a saddle or a local maximum, not a minimum.
Discussion

The ray equation \( \frac{d}{ds}(n\,d\mathbf{r}/ds) = \nabla n \) is a statement of "optical inertia". The quantity \( \mathbf{p} = n\,\hat{\mathbf{t}} = \nabla S \) is the optical momentum (the ray analogue of a mechanical momentum), and \( \nabla n \) is the force that turns it. Where the index is uniform the momentum is constant and the ray coasts in a straight line; where \( n \) increases the momentum is deflected toward the gradient, so light curves into the optically denser region. This is exactly why mirages, atmospheric looming and the twinkling of stars occur: vertical index gradients bend otherwise-straight rays.

Recasting the result as \( \delta\int n\, ds = 0 \) reveals the variational structure. Fermat's principle is a genuine stationary-action principle with "action" \( \int n\, ds \), and the eikonal \( S \) is its Hamilton characteristic function: \( \lvert\nabla S\rvert = n \) is precisely the Hamilton–Jacobi equation of a system with Hamiltonian \( H = \tfrac12(\lvert\mathbf{p}\rvert^2 - n^2) = 0 \). Rays are the characteristics (bicharacteristics) of this equation, and wavefronts \( S=\text{const} \) are the surfaces they thread orthogonally in an isotropic medium.

The symmetry content is worth naming explicitly. Whenever \( n \) is independent of a coordinate, the conjugate component of \( n\,\hat{\mathbf{t}} \) is conserved — a Noether conservation law for the optical action. Translational symmetry in \( x \) for a stratified medium gives \( n\sin\theta=\text{const} \) (Snell), and rotational symmetry for a spherically symmetric \( n(r) \) gives Bouguer's invariant \( n\, r\sin\phi = \text{const} \). Fermat's principle thus organises ray optics the way least action organises mechanics.

At the deepest level this derivation is the classical limit of a wave theory taken twice over. The eikonal ansatz \( \psi \sim A\, e^{i k_0 S} \) is the optical WKB expansion; \( \lvert\nabla S\rvert = n \) is its leading order and the ray equation its characteristic flow, exactly paralleling how the quantum WKB phase \( e^{iW/\hbar} \) yields \( \lvert\nabla W\rvert = \sqrt{2m(E-V)} \) and Newtonian trajectories as \( \hbar\to 0 \). The optical path length \( \int n\, ds \) maps onto Maupertuis' abbreviated action \( \int \mathbf{p}\cdot d\mathbf{q} \), with \( n \leftrightarrow \lvert\mathbf{p}\rvert \). Hamiltonian optics — ray transfer matrices, symplectic phase space, Liouville's theorem for étendue — is this analogy made into machinery.

Common misconceptions. Fermat's principle is often stated as "light takes the path of least time." It is really the path of stationary optical path length: at a focus or off a curved mirror the ray can be a saddle point or even a local maximum. Nor does light "try out" all paths and select one — the stationary path is simply the constructive-interference condition of the underlying wave, recovered here as the eikonal characteristic. Finally, the bending of a ray is not caused by \( n \) being large but by \( n \) varying: a uniform slab of glass, however high its index, bends nothing internally.

Worked examples

Example 1 — Optical path length and straight rays in a uniform slab.

1
\[ \nabla n = 0 \;\Rightarrow\; \frac{d}{ds}\!\left( n\,\frac{d\mathbf{r}}{ds}\right)=0 \;\Rightarrow\; \frac{d^2\mathbf{r}}{ds^2}=0 \]
In a homogeneous medium the ray equation collapses: constant \( n \) divides out and the tangent is constant. A
2
\[ \text{OPL} = \int n\, ds = n\, L,\qquad n = 1.50,\;\; L = 4.0\ \mathrm{cm} \]
With straight rays the geometric path is a line of length \( L \); the optical path length is simply \( nL \). Insert the glass values. A
3
\[ \text{OPL} = 1.50 \times 4.0\ \mathrm{cm} = 6.0\ \mathrm{cm};\qquad t = \frac{\text{OPL}}{c} = \frac{0.060\ \mathrm{m}}{3.0\times10^{8}\ \mathrm{m\,s^{-1}}} \]
Evaluate. The travel time is the optical path length divided by \( c \) (since \( n\, ds/c = ds/v \)). A
\[ \text{OPL} = 6.0\ \mathrm{cm},\qquad t = 2.0\times10^{-10}\ \mathrm{s} = 0.20\ \mathrm{ns} \]

Reading. Although the glass is only \( 4.0\ \mathrm{cm} \) thick, light traverses \( 6.0\ \mathrm{cm} \) of optical path — the extra \( 2.0\ \mathrm{cm} \) is the phase delay that makes the slab act like a longer stretch of vacuum, and Fermat's principle confirms this straight line is the stationary path.

Example 2 — Sinusoidal rays in a graded-index (GRIN) fibre.

1
\[ n(x) = n_0\left(1 - \tfrac{1}{2}g^2 x^2\right),\qquad n_0 = 1.50,\;\; g = 0.25\ \mathrm{mm^{-1}} \]
Standard parabolic GRIN profile; \( x \) is the transverse displacement from the axis and \( z \) the propagation direction. A
2
\[ \frac{d}{ds}\!\left(n\,\frac{dx}{ds}\right)=\frac{\partial n}{\partial x}\;\xrightarrow{\text{paraxial } ds\approx dz,\; n\approx n_0}\; n_0\,\frac{d^2 x}{dz^2} = -\,n_0\, g^2 x \]
Take the transverse component of the ray equation. In the paraxial regime the ray is nearly axial, so \( ds \approx dz \) and \( n\approx n_0 \) in the prefactor; differentiate the profile for the right-hand side. B
3
\[ \frac{d^2 x}{dz^2} = -\,g^2 x \;\Rightarrow\; x(z) = x_0\cos(g z) + \frac{x_0'}{g}\sin(g z),\qquad \Lambda = \frac{2\pi}{g} \]
Simple-harmonic transverse equation; its solution is sinusoidal with spatial period (pitch) \( \Lambda = 2\pi/g \). B
4
\[ \Lambda = \frac{2\pi}{0.25\ \mathrm{mm^{-1}}} \]
Insert the profile constant \( g \). The pitch is independent of the launch amplitude — every paraxial ray refocuses together. A
\[ \Lambda = \frac{2\pi}{0.25\ \mathrm{mm^{-1}}} \approx 25.1\ \mathrm{mm} \]

Reading. Rays launched off-axis oscillate about the fibre axis and return to focus every \( 25\ \mathrm{mm} \). This self-focusing — a direct consequence of \( \nabla n \) pointing toward the axis in the ray equation — is exactly how GRIN rod lenses and multimode graded-index fibres guide light without discrete refracting surfaces.

Problems
  1. Optical path across two slabs. Light passes normally through \( 3.0\ \mathrm{cm} \) of water (\( n=1.33 \)) then \( 2.0\ \mathrm{cm} \) of glass (\( n=1.52 \)). Find the total optical path length and the transit time.
    Solution Rays are straight (uniform slabs, \( \nabla n = 0 \)). \( \text{OPL} = n_1 L_1 + n_2 L_2 = 1.33(3.0) + 1.52(2.0) = 3.99 + 3.04 = 7.03\ \mathrm{cm} \). Time \( t = \text{OPL}/c = 0.0703/(3.0\times10^8) = 2.34\times10^{-10}\ \mathrm{s} \approx 0.23\ \mathrm{ns} \).
  2. Snell's invariant from the ray equation. For a stratified medium \( n=n(y) \) with rays confined to the \( xy \)-plane, show that \( n\sin\theta \) is constant along a ray, where \( \theta \) is the angle from the \( y \)-axis.
    Solution The \( x \)-component of the ray equation is \( \frac{d}{ds}\!\left(n\frac{dx}{ds}\right)=\frac{\partial n}{\partial x}=0 \) since \( n \) depends only on \( y \). Hence \( n\,\frac{dx}{ds}=\text{const} \). But \( \frac{dx}{ds}=\sin\theta \) (the tangent's angle from the \( y \)-axis), so \( n\sin\theta = \text{const} \). This is Snell's law in continuous form; at a discrete interface it becomes \( n_1\sin\theta_1 = n_2\sin\theta_2 \).
  3. Radius of curvature of a horizontal ray (mirage). Near the ground the index varies vertically with \( \lvert dn/dz \rvert = 1.0\times10^{-7}\ \mathrm{m^{-1}} \) and \( n\approx 1.000 \). For a ray travelling horizontally, find its radius of curvature. Compare with Earth's radius \( R_\oplus = 6.37\times10^{6}\ \mathrm{m} \).
    Solution Write the ray equation as \( n\,\kappa\,\hat{\mathbf{n}} + (\hat{\mathbf{t}}\cdot\nabla n)\hat{\mathbf{t}} = \nabla n \). The normal component gives \( n\kappa = \hat{\mathbf{n}}\cdot\nabla n = \lvert\nabla_\perp n\rvert \). For a horizontal ray the perpendicular gradient is the vertical one, so \( \kappa = \lvert dn/dz\rvert / n \) and \( R = 1/\kappa = n/\lvert dn/dz\rvert = 1.000/(1.0\times10^{-7}) = 1.0\times10^{7}\ \mathrm{m} \). This is \( \approx 1.6\,R_\oplus \); because the ray's curvature is comparable to Earth's, strong gradients let light hug the surface and produce mirages and looming.
  4. GRIN pitch and focal spacing. A GRIN rod has \( n(r)=n_0(1-\tfrac12 g^2 r^2) \) with \( n_0=1.60 \) and \( g=0.40\ \mathrm{mm^{-1}} \). Find (a) the ray pitch \( \Lambda \) and (b) the shortest rod length that images an input face onto the output face (a "quarter-pitch" lens).
    Solution Paraxial transverse equation \( d^2 r/dz^2 = -g^2 r \Rightarrow \Lambda = 2\pi/g = 2\pi/0.40 = 15.7\ \mathrm{mm} \). A quarter-pitch lens converts a point at the input into collimated/imaged output after a quarter period: \( L = \Lambda/4 = 15.7/4 = 3.93\ \mathrm{mm} \). (Half-pitch, \( \Lambda/2 = 7.85\ \mathrm{mm} \), reimages a point to a point.)
  5. Bouguer's theorem (spherical symmetry). For a spherically symmetric medium \( n=n(r) \), show that \( n\, r \sin\phi \) is conserved along a ray, where \( \phi \) is the angle between the ray and the local radial direction. (Hint: use the symmetry / angular-momentum structure of \( n\,\hat{\mathbf{t}} \).)
    Solution Consider the vector \( \mathbf{L} = \mathbf{r}\times(n\,\hat{\mathbf{t}}) \). Its arc-length derivative is \( \frac{d\mathbf{L}}{ds}=\frac{d\mathbf{r}}{ds}\times(n\hat{\mathbf{t}}) + \mathbf{r}\times\frac{d}{ds}(n\hat{\mathbf{t}}) \). The first term is \( \hat{\mathbf{t}}\times(n\hat{\mathbf{t}})=\mathbf{0} \). By the ray equation the second is \( \mathbf{r}\times\nabla n \); for \( n=n(r) \), \( \nabla n = (dn/dr)\hat{\mathbf{r}} \parallel \mathbf{r} \), so \( \mathbf{r}\times\nabla n=\mathbf{0} \). Hence \( \mathbf{L} \) is constant, and its magnitude \( \lvert\mathbf{L}\rvert = r\,n\,\lvert\hat{\mathbf{r}}\times\hat{\mathbf{t}}\rvert = n\,r\sin\phi \) is conserved. This is Bouguer's theorem, the optical analogue of angular-momentum conservation, and it governs ray tracing in lens atmospheres and the solar chromosphere.