Selection coefficient and allele frequency change
Statement
At one autosomal locus with alleles \(A\) and \(a\) at frequencies \(p\) and \(q=1-p\) in an infinite, randomly mating, diploid population with discrete non-overlapping generations, let constant viability selection act with relative genotype fitnesses \(w_{AA}=1\), \(w_{Aa}=1-hs\) and \(w_{aa}=1-s\), where \(s\in[0,1]\) is the selection coefficient against \(a\) and \(h\) its dominance coefficient. Then one generation of selection changes the frequency of \(A\) by \(\Delta p = pq\,(w_A-w_a)/\bar w\), where \(w_A\) and \(w_a\) are the marginal (allelic) fitnesses and \(\bar w\) the population mean fitness; equivalently \(\Delta p = s\,pq\,[\,ph+q(1-h)\,]/\bar w\), which is exact for any \(s\), vanishes only at \(p=0\), \(p=1\) or where the marginal fitnesses are equal, and reduces to \(\Delta p \approx \tfrac{1}{2}spq\) for weak additive selection.
Why it matters
Natural selection is stated qualitatively as differential survival and reproduction of heritable variation; that statement, on its own, predicts nothing numerical. The selection coefficient converts it into an equation: \(s\) is the fractional reduction in relative fitness borne by one genotype, and the recursion derived below turns that single number into a per-generation allele-frequency change, a trajectory, a waiting time and — when read backwards — an estimator of \(s\) from genotype counts. It is the point at which evolution stops being a narrative and becomes an arithmetic that can be tested against a real cohort.
It is also the deterministic backbone of everything else in population genetics. Hardy-Weinberg supplies the zygote frequencies each generation starts from; mutation-selection balance is exactly this recursion set equal and opposite to a mutational input; genetic drift and neutral theory are the statements of what happens when \(s\) is too small for this recursion to dominate the sampling noise. The comparison \(|s|\) against \(1/(2N_e)\) — the boundary between the deterministic regime treated here and the stochastic one — is the single most used inequality in the field.
Hypotheses
Proof
The generation cycle is: zygotes in Hardy-Weinberg proportions, then differential survival, then random mating among survivors to make the next zygote pool. The whole derivation is the bookkeeping of that cycle, done in symbols before any number is substituted.
Result
Reading. The per-generation gain of the favoured allele is the product of three things: the heterozygosity-like variation term \(pq\), which is zero at both boundaries and maximal at \(p=\tfrac12\); the marginal fitness difference \(w_A-w_a\), which is what “strength of selection” actually means; and the normaliser \(\bar w\). The selection coefficient \(s\) is a dimensionless per-generation fitness ratio, not a rate of frequency change: \(s=0.1\) never means “the allele rises by 10% per generation”.
Scope. Exact for constant-viability selection at one autosomal locus in an infinite randomly mating population with discrete generations, for any \(s\in[0,1]\) and any \(h\). It is the deterministic limit of a stochastic process, and is trustworthy only while \(|s|\gg 1/(2N_e)\).
Corollaries & converses
- Equilibria. \(\Delta p=0\) exactly when \(p=0\), \(p=1\), or \(w_A=w_a\). For \(0\le h\le1\) only the boundaries qualify, so the favoured allele fixes; an internal equilibrium requires \(h\lt0\) (heterozygote advantage) or \(h\gt1\) (heterozygote disadvantage).
- Overdominance. With \(w_{AA}=1-s_1\), \(w_{Aa}=1\), \(w_{aa}=1-s_2\), the condition \(w_A=w_a\) gives \(\hat p=s_2/(s_1+s_2)\). It is stable: both alleles are permanently maintained, and the population pays a segregation load \(1-\bar w=s_1s_2/(s_1+s_2)\).
- Underdominance. The same algebra with the heterozygote worst gives an internal equilibrium that is unstable — a threshold, above which \(A\) fixes and below which it is lost. This is the population-genetic basis of chromosomal rearrangements spreading only after a founder event.
- Estimating \(s\) from data. Rearranging Step 6, \(w_A/w_a=\bigl[p'(1-p)\bigr]/\bigl[p(1-p')\bigr]\): the ratio of marginal fitnesses is the odds ratio of two successive allele frequencies, so \(s\) is measurable from allele counts alone without ever observing a death.
- Haldane’s sieve. The bracket tends to \(1-h\) as \(p\to0\), so the initial spread of a rare favoured allele \(A\) is set entirely by the advantage it expresses in heterozygotes. If \(A\) is recessive (\(h\to1\), heterozygotes no fitter than \(aa\)) the bracket vanishes and selection is nearly blind to it; dominant beneficial alleles pass the sieve, recessive ones mostly do not.
- Additive superposition fails. \(\Delta p\) is not linear in \(s\) except to first order; for \(s\) near 1 the \(\bar w\) denominator matters, and the naive \(\Delta p\approx\frac12 spq\) can be wrong by tens of percent.
Fails without
- Finite population size — the drift regime: the recursion is a deterministic expectation, while drift adds a per-generation variance \(pq/(2N_e)\). When \(|s|\lesssim 1/(2N_e)\) the noise exceeds the signal and the allele behaves as if neutral. Kimura’s diffusion result for genic selection, \(u(p)=\bigl(1-e^{-4N_es_hp}\bigr)/\bigl(1-e^{-4N_es_h}\bigr)\), replaces certain fixation with a probability: a single new mutant with heterozygous advantage \(s_h\) in a population with \(N_e\approx N\) fixes with probability only about \(2s_h\), so a 1% beneficial mutation is lost 98 times out of 100 despite \(\Delta p\gt0\) at every frequency.
- Frequency- or density-dependent fitness: with \(w\) a function of \(p\), Step 9’s gradient reading collapses because \(\bar w\) is no longer a fixed landscape. Under rare-morph advantage the same algebra yields a stable polymorphism at which mean fitness is not maximal, and mean fitness can decrease over time — a possibility rigorously excluded in the constant-fitness model.
- Non-random mating: with inbreeding coefficient \(F\gt0\) the zygote frequencies of Step 1 become \(p^2+Fpq,\ 2pq(1-F),\ q^2+Fpq\), so a deleterious recessive is exposed at \(q^2+Fpq\) rather than \(q^2\). For \(q=0.01\) and \(F=0.05\) the exposed fraction rises from \(1.0\times10^{-4}\) to \(5.95\times10^{-4}\), a nearly sixfold increase in the rate of removal, and the \(1/q_t\) law of Step 12 no longer holds.
- Selection through fertility or in overlapping generations: if fitness attaches to mating pairs rather than to individuals, or generations overlap so that the census is not a clean zygote pool, the state cannot be summarised by \(p\) alone and no single-locus recursion of this form exists.
Common errors
- “\(s=0.02\), so the allele rises by 2% per generation.” \(s\) is a fitness ratio; the frequency change is \(\Delta p\approx\frac12 spq\), which at \(p=0.01\) is about \(10^{-4}\) — two hundred times smaller.
- “Fitness is a property of an allele.” Selection acts on diploid genotypes; the allelic quantities \(w_A,w_a\) are marginal averages over genetic backgrounds (Step 4) and therefore change as \(p\) changes, even with the three genotype fitnesses held fixed.
- “Selection will clear a deleterious recessive from the population.” Step 12 gives hyperbolic, not exponential, decay: even a lethal recessive at \(q=0.01\) needs 100 generations to reach \(q=0.005\), because almost every copy is sheltered in a heterozygote.
- “Dropping \(\bar w\) is always safe.” True to \(O(s^2)\) for weak selection, badly wrong for strong selection: with \(h=0\), \(s=1\), \(q=0.5\) the denominator is \(0.75\), a 33% correction.
- Confusing \(s\) with \(\mu\). Selection coefficients of interest run from \(10^{-4}\) to 1; per-locus mutation rates are of order \(10^{-6}\) to \(10^{-5}\). They enter mutation-selection balance asymmetrically: \(\hat q=\sqrt{\mu/s}\) for a recessive deleterious allele, \(\hat q\approx\mu/(hs)\) once heterozygotes pay a cost.
- Reporting an estimated \(s\) without \(N_e\). A fitted \(s=0.001\) means opposite things in a population of \(N_e=10^6\) (strongly selected) and \(N_e=100\) (effectively neutral).
Discussion
The algebra above is the technical core of the modern synthesis. Haldane worked through the single-locus cases — dominant, recessive, sex-linked, with and without inbreeding — in his series “A mathematical theory of natural and artificial selection” beginning in 1924, and it was he who established the slow, hyperbolic elimination of recessives that Step 12 reproduces. Fisher’s The Genetical Theory of Natural Selection (1930) recast the same content as a statement about variance, and Wright’s gradient form (Step 9) supplied the adaptive-landscape picture that has dominated intuition ever since. That these three routes give the same recursion is the reason the one-locus model is treated as settled ground rather than one model among many.
Empirically, the useful direction is backwards. Given genotype counts at two life stages, or allele frequencies at two time points, the odds-ratio corollary returns \(w_A/w_a\) and hence \(s\); given ancient DNA or serial sampling in an experimental population, the log-odds line of Step 11 is fitted and its slope is \(s/2\). Selection coefficients recovered this way for the strongest known signals in human populations — lactase persistence, several malaria-resistance loci — come out of order \(10^{-2}\) per generation, which is precisely why they are detectable at all: at \(s=10^{-2}\) a sweep takes a few thousand generations, short enough to leave a haplotype signature and long enough to be plausible.
The deterministic model’s real boundary is not weak selection but the product \(N_es\). The diffusion approximation shows that fixation probability depends on \(p\) and \(N_es\) alone, so a “strongly selected” allele in a bottlenecked population and a “nearly neutral” one in a vast population can have identical dynamics. This is Ohta’s nearly neutral theory in one inequality, and it is why comparative genomics reads the efficacy of selection off effective population size: the same \(s\) is visible in a bacterium with \(N_e\sim10^8\) and invisible in a large mammal with \(N_e\sim10^4\).
Common misconceptions. First, that \(\Delta p\gt0\) guarantees fixation — it guarantees fixation only in the infinite-population idealisation, and most beneficial mutations are in fact lost while rare. Second, that the maximum of \(pq\) at \(p=\tfrac12\) means selection is “strongest” in the middle of a sweep; the fitness difference \(w_A-w_a\) is generally largest when the deleterious allele is common, and the two factors trade off, which is why sweeps are slow at both ends and fast in the middle. Third, that fixing a beneficial allele improves the population without cost: Haldane’s substitution load counts the deaths required, and it is this bookkeeping that motivated the neutralist argument in the first place.
Worked examples
Example 1. A fully recessive lethal allele (\(h=0\), \(s=1\)) sits at frequency \(q_0=0.10\) in a large randomly mating population. How many generations of complete selection against homozygotes are needed to halve it, and how many more to halve it again? Compare with the naive expectation from a single-generation calculation.
Reading. Complete lethality of homozygotes — the strongest selection possible — removes a recessive allele only hyperbolically, because heterozygotes shelter a fraction \(2pq/(2pq+2q^{2})=p\to1\) of all copies of \(a\) as \(q\to0\). Eugenic proposals to eliminate recessive disease alleles by preventing affected individuals from reproducing founder on exactly this arithmetic.
Scope. Exact for \(h=0,\ s=1\) under the stated hypotheses; any heterozygote cost (\(h\gt0\)) accelerates removal dramatically, since the exposed fraction jumps from \(q^2\) to \(2hpq+q^2\).
Example 2. A cohort of 1000 newly settled marine invertebrate juveniles is genotyped at a single locus: 360 \(AA\), 480 \(Aa\), 160 \(aa\). At the end of the season the survivors are recounted: 324 \(AA\), 408 \(Aa\), 96 \(aa\). Estimate \(s\) and \(h\), predict \(\Delta p\) from the Result, and check it against the survivor counts.
Reading. A single season of quite strong selection (a third of the fitness of \(aa\) lost) moves the allele frequency by under four percentage points, from 0.600 to 0.638. Selection coefficients large enough to be measured in one field season still produce modest per-generation frequency changes.
Scope. The estimates are point estimates from one cohort with no sampling error attached; with 1000 juveniles the binomial standard error on \(p'\) is about \(\sqrt{p'(1-p')/1656}\approx0.012\), so \(s\) here is resolved only to roughly the first decimal place.
Problems
- An additive allele (\(h=\tfrac12\)) has \(s=0.10\) and is currently at \(p=0.50\). Compute \(\Delta p\) exactly from the Result and from the weak-selection approximation of Step 10, and give the percentage error of the approximation.
Solution
Bracket: \(ph+q(1-h)=0.5(0.5)+0.5(0.5)=0.500\). Denominator: \(\bar w=1-2pq\,hs-q^2s=1-2(0.25)(0.5)(0.1)-(0.25)(0.1)=1-0.0250-0.0250=0.9500\).
Exact: \(\Delta p=(0.10)(0.25)(0.500)/0.9500=0.01250/0.9500=0.013158\), so \(p'=0.51316\).
Approximation: \(\Delta p\approx\tfrac12 spq=\tfrac12(0.10)(0.25)=0.01250\).
Error \(=(0.01250-0.013158)/0.013158=-0.050\), i.e. the approximation is 5.0% low — exactly the \(O(s)\) size of the neglected \(\bar w\), since \(1-\bar w=0.05\).
- A deleterious allele sits at \(q=0.05\) with \(s=0.20\). Compute \(\Delta q\) (i) if it is fully recessive (\(h=0\)) and (ii) if it is fully dominant (\(h=1\)), and explain the ratio in one sentence.
Solution
(i) \(h=0\): \(\bar w=1-sq^2=1-(0.2)(0.0025)=0.99950\); \(\Delta q=-spq^2/\bar w=-(0.2)(0.95)(0.0025)/0.99950=-4.75\times10^{-4}/0.99950=-4.752\times10^{-4}\).
(ii) \(h=1\): \(\bar w=1-2pq s-q^2s=1-sq(2p+q)=1-(0.2)(0.05)(1.95)=0.98050\); \(\Delta q=-sp^2q/\bar w=-(0.2)(0.9025)(0.05)/0.98050=-9.025\times10^{-3}/0.98050=-9.204\times10^{-3}\).
Ratio \(=9.204\times10^{-3}/4.752\times10^{-4}=19.4\), essentially \(p/q=19\). The marginal fitness difference is \(w_A-w_a=s\,[\,ph+q(1-h)\,]\), which is \(sp\) for a fully dominant deleterious allele and \(sq\) for a fully recessive one, so the two \(\Delta q\) values stand in the ratio \(p/q\): a rare dominant is exposed to selection in every heterozygous carrier, a rare recessive only in the far rarer \(q^{2}\) homozygotes.
- A lethal recessive allele (\(h=0\), \(s=1\)) is at \(q_0=0.02\). Find \(q\) after 50 and after 150 generations, the frequency of affected homozygotes at each time, and the number of generations needed to reach \(q=0.001\).
Solution
Step 12: \(1/q_t=1/q_0+t=50+t\).
\(t=50\): \(1/q=100\), \(q=0.0100\), affected \(q^2=1.00\times10^{-4}\) (1 in 10,000).
\(t=150\): \(1/q=200\), \(q=0.00500\), affected \(q^2=2.50\times10^{-5}\) (1 in 40,000).
\(q=0.001\): \(1/q=1000\), so \(t=1000-50=950\) generations.
Halving the allele frequency from 0.02 to 0.01 takes 50 generations; the next halving takes 100; reaching a twentieth of the starting value takes 950. The affected phenotype frequency falls by a factor of 400 over those 950 generations, which is why complete selection against a recessive phenotype looks effective early and stalls indefinitely afterwards.
- Suppose a locus shows heterozygote advantage with illustrative relative fitnesses \(w_{AA}=0.88\), \(w_{AS}=1\), \(w_{SS}=0.20\) (the sickle-cell pattern in a high-malaria environment; the numbers here are round illustrative values, not measured estimates). Find the equilibrium frequency of \(S\), verify it is stable by evaluating \(\Delta p\) on either side, and compute the mean fitness and segregation load at equilibrium.
Solution
Write \(s_1=1-w_{AA}=0.12\) (against \(AA\)) and \(s_2=1-w_{SS}=0.80\) (against \(SS\)). Let \(p\) be the frequency of \(A\), \(q\) that of \(S\). Marginal fitnesses: \(w_A=p(1-s_1)+q\) and \(w_S=p+q(1-s_2)\), so \(w_A-w_S=q s_2-p s_1\).
Equilibrium: \(\hat p s_1=\hat q s_2\Rightarrow \hat p=s_2/(s_1+s_2)=0.80/0.92=0.8696\), \(\hat q=s_1/(s_1+s_2)=0.12/0.92=0.1304\).
Stability: at \(q=0.10\ (\lt\hat q)\), \(w_A-w_S=(0.10)(0.80)-(0.90)(0.12)=0.080-0.108=-0.028\lt0\), so \(\Delta p\lt0\) and \(q\) rises. At \(q=0.20\ (\gt\hat q)\), \(w_A-w_S=(0.20)(0.80)-(0.80)(0.12)=0.160-0.096=+0.064\gt0\), so \(q\) falls. The equilibrium is approached from both sides: it is stable.
Mean fitness: \(\bar w=\hat p^2(0.88)+2\hat p\hat q(1)+\hat q^2(0.20)=0.7561(0.88)+2(0.11342)+0.01701(0.20)=0.6654+0.22684+0.00340=0.8956\). Segregation load \(=1-\bar w=0.1044\), matching \(s_1s_2/(s_1+s_2)=(0.12)(0.80)/0.92=0.1043\).
At equilibrium 1.7% of births are \(SS\) and 22.7% are carriers — the polymorphism is maintained by selection itself, not by mutation, and it costs the population about 10% of its mean fitness.
- An additive beneficial allele (\(h=\tfrac12\)) with \(s=0.005\) starts at \(p_0=0.02\). (a) How many generations until \(p=0.50\)? (b) Repeat for \(s=0.05\). (c) For \(N_e=10^4\), is the deterministic treatment justified, and what is the fixation probability of a single new copy of this allele when \(s=0.005\)?
Solution
(a) Step 11: \(t=\dfrac{2}{s}\ln\!\left[\dfrac{p_t(1-p_0)}{p_0(1-p_t)}\right]=\dfrac{2}{0.005}\ln\!\left[\dfrac{(0.50)(0.98)}{(0.02)(0.50)}\right]=400\ln(49)=400(3.8918)=1557\) generations.
(b) The log-odds term is unchanged, so \(t=\dfrac{2}{0.05}(3.8918)=40(3.8918)=156\) generations — waiting time scales as \(1/s\).
(c) \(2N_es=2(10^4)(0.005)=100\gg1\), so selection dominates drift once the allele is common and the deterministic trajectory is a good description; at \(p_0=0.02\) there are already \(2N_ep_0=400\) copies, far from the stochastic boundary.
Fixation of a single new copy: rescale so the resident \(aa\) has fitness 1. Then the heterozygote has \(w_{Aa}/w_{aa}=(1-s/2)/(1-s)\approx1+s/2\), i.e. \(s_h=0.0025\). Haldane’s approximation with \(N_e\approx N\) gives \(u\approx2s_h=0.0050\): about 1 new beneficial copy in 200 fixes, and 199 are lost to drift while rare, even though \(\Delta p\gt0\) at every frequency.