biology2u
Tier
⌕ Search ⌘K
Concept

Selection coefficient and allele frequency change

T-175Home BU-102Threads structure · energy
Statement

At one autosomal locus with alleles \(A\) and \(a\) at frequencies \(p\) and \(q=1-p\) in an infinite, randomly mating, diploid population with discrete non-overlapping generations, let constant viability selection act with relative genotype fitnesses \(w_{AA}=1\), \(w_{Aa}=1-hs\) and \(w_{aa}=1-s\), where \(s\in[0,1]\) is the selection coefficient against \(a\) and \(h\) its dominance coefficient. Then one generation of selection changes the frequency of \(A\) by \(\Delta p = pq\,(w_A-w_a)/\bar w\), where \(w_A\) and \(w_a\) are the marginal (allelic) fitnesses and \(\bar w\) the population mean fitness; equivalently \(\Delta p = s\,pq\,[\,ph+q(1-h)\,]/\bar w\), which is exact for any \(s\), vanishes only at \(p=0\), \(p=1\) or where the marginal fitnesses are equal, and reduces to \(\Delta p \approx \tfrac{1}{2}spq\) for weak additive selection.

Why it matters

Natural selection is stated qualitatively as differential survival and reproduction of heritable variation; that statement, on its own, predicts nothing numerical. The selection coefficient converts it into an equation: \(s\) is the fractional reduction in relative fitness borne by one genotype, and the recursion derived below turns that single number into a per-generation allele-frequency change, a trajectory, a waiting time and — when read backwards — an estimator of \(s\) from genotype counts. It is the point at which evolution stops being a narrative and becomes an arithmetic that can be tested against a real cohort.

It is also the deterministic backbone of everything else in population genetics. Hardy-Weinberg supplies the zygote frequencies each generation starts from; mutation-selection balance is exactly this recursion set equal and opposite to a mutational input; genetic drift and neutral theory are the statements of what happens when \(s\) is too small for this recursion to dominate the sampling noise. The comparison \(|s|\) against \(1/(2N_e)\) — the boundary between the deterministic regime treated here and the stochastic one — is the single most used inequality in the field.

Hypotheses
Selection acts on viability only, between the zygote census and reproduction.Fitness differences are collapsed into three constants applied to zygote frequencies. If selection instead acts through fertility, the fitness of a mating pair need not factorise into genotype fitnesses at all, and no single-locus recursion of this form exists: a pair-fitness matrix \(F_{ij}\) can produce cycling and internal equilibria that \(\Delta p = pq(w_A-w_a)/\bar w\) cannot represent.
Genotype fitnesses are constants, independent of allele frequency, population density and generation.Drop it and \(\bar w\) ceases to be a function of \(p\) alone, so Wright’s gradient form (Step 8) and the monotone increase of \(\bar w\) both fail. Under negative frequency dependence — rare-morph advantage in prey polymorphisms, or self-incompatibility alleles in plants — the same algebra yields a stable internal equilibrium at which mean fitness is not maximised.
Mating is random and the population is effectively infinite, so zygotes are in Hardy-Weinberg proportions each generation.Random mating restores \(p^2,\,2pq,\,q^2\) at the start of every generation even though selection has just destroyed those proportions among the survivors. Without it, the post-selection genotype frequencies must be carried forward as separate state variables, and with inbreeding (\(F\gt0\)) the effective exposure of a recessive allele to selection rises from \(q^2\) to \(q^2+Fpq\).
The locus is the sole target of selection: no linkage, no epistasis, no mutation or migration during the interval.Fitness is assigned to the genotype at this locus alone. A neutral allele in strong linkage disequilibrium with a selected site changes frequency by hitchhiking, giving an apparent \(s\) that belongs to a neighbour, and epistasis makes the marginal fitness of \(A\) depend on the genetic background rather than on \(p\) alone.
Proof

The generation cycle is: zygotes in Hardy-Weinberg proportions, then differential survival, then random mating among survivors to make the next zygote pool. The whole derivation is the bookkeeping of that cycle, done in symbols before any number is substituted.

1
\[ f(AA)=p^{2},\qquad f(Aa)=2pq,\qquad f(aa)=q^{2},\qquad p+q=1 \]
Random mating in an infinite population makes gametes unite independently, so zygote genotype frequencies are the binomial expansion of \((p+q)^2\) — the Hardy-Weinberg proportions. This is the census point: every generation is entered in this state, whatever selection did to the previous one. A
2
\[ \bar w \;=\; p^{2}w_{AA}+2pq\,w_{Aa}+q^{2}w_{aa} \]
Each genotype survives to reproduce in proportion to its relative fitness, so the surviving fraction of the cohort is the frequency-weighted mean of the three fitnesses. Because the \(w\) are relative (defined up to a common positive factor), \(\bar w\) is a normalising constant, not an absolute survival probability. A
3
\[ f^{*}(AA)=\frac{p^{2}w_{AA}}{\bar w},\qquad f^{*}(Aa)=\frac{2pq\,w_{Aa}}{\bar w},\qquad f^{*}(aa)=\frac{q^{2}w_{aa}}{\bar w} \]
Dividing each selected class by \(\bar w\) renormalises the survivors to sum to one. Note that these are no longer Hardy-Weinberg proportions: selection destroys them within a generation, and Step 1 rebuilds them at the next zygote census. A
4
\[ p' \;=\; f^{*}(AA)+\tfrac{1}{2}f^{*}(Aa) \;=\; \frac{p^{2}w_{AA}+pq\,w_{Aa}}{\bar w} \;=\; \frac{p\,w_{A}}{\bar w},\qquad w_{A}\equiv p\,w_{AA}+q\,w_{Aa} \]
A survivor’s gametes carry \(A\) with probability 1 for \(AA\) and \(\tfrac12\) for \(Aa\) (Mendelian segregation), so the gamete pool’s \(A\) frequency is the survivor frequency of \(AA\) plus half that of \(Aa\). Factoring \(p\) out defines the marginal fitness \(w_A\): the mean fitness of a randomly chosen individual carrying a randomly chosen \(A\) allele. B
5
\[ w_{a}\equiv p\,w_{Aa}+q\,w_{aa},\qquad \bar w \;=\; p\,w_{A}+q\,w_{a} \]
The same construction for \(a\); expanding \(pw_A+qw_a=p^2w_{AA}+2pq\,w_{Aa}+q^2w_{aa}\) recovers Step 2 exactly. So mean fitness is the allele-frequency-weighted mean of the two marginal fitnesses — the identity that makes the next step collapse. B
6
\[ \Delta p \;=\; p'-p \;=\; \frac{p\,w_{A}}{\bar w}-p \;=\; \frac{p\,(w_{A}-\bar w)}{\bar w} \;=\; \frac{p\,\bigl(w_{A}-p\,w_{A}-q\,w_{a}\bigr)}{\bar w} \;=\; \frac{pq\,(w_{A}-w_{a})}{\bar w} \]
Subtracting \(p\), putting over the common denominator \(\bar w\), then substituting the identity of Step 5 and using \(1-p=q\). The result is exact for arbitrarily strong selection: no approximation has been made anywhere. Its structure is already the whole biology — a variation factor \(pq\), a selection factor \(w_A-w_a\), and a normaliser. B
7
\[ w_{AA}=1,\quad w_{Aa}=1-hs,\quad w_{aa}=1-s \;\Longrightarrow\; w_{A}-w_{a}=s\bigl[\,ph+q(1-h)\,\bigr],\quad \bar w = 1-2pq\,hs-q^{2}s \]
Substituting the standard \((s,h)\) parametrisation: \(w_A=p+q(1-hs)=1-qhs\) and \(w_a=p(1-hs)+q(1-s)=1-phs-qs\), whose difference is \(s(ph+q-qh)\). For \(\bar w\), the three terms sum to \(1\) minus the fitness deficits \(2pq\,hs\) and \(q^2 s\). Only now do parameters with named biological meaning enter. A
8
\[ \Delta p \;=\; \frac{s\,pq\,\bigl[\,ph+q(1-h)\,\bigr]}{1-2pq\,hs-q^{2}s},\qquad \Delta q=-\Delta p \]
Combining Steps 6 and 7. Since the bracket is a convex combination of \(h\) and \(1-h\), it is strictly positive whenever \(0\le h\le 1\) and \(0\lt p\lt1\): with \(s\gt0\) and no over- or under-dominance, \(A\) increases every generation without exception. \(\Delta q=-\Delta p\) because the two frequencies sum to one. B
9
\[ \frac{d\bar w}{dp}=2\bigl[p\,w_{AA}+(1-2p)w_{Aa}-q\,w_{aa}\bigr]=2\,(w_{A}-w_{a}) \;\Longrightarrow\; \Delta p=\frac{pq}{2\,\bar w}\,\frac{d\bar w}{dp} \]
Differentiating \(\bar w=p^2w_{AA}+2p(1-p)w_{Aa}+(1-p)^2w_{aa}\) with respect to \(p\) and recognising the marginal fitnesses of Steps 4–5. This is Wright’s form: with constant fitnesses the population climbs the gradient of \(\bar w\) at a rate proportional to the genetic variation \(pq\) available to it, which is why \(\bar w\) never decreases in this model and \(pq\) is called the “fuel” of selection. C
10
\[ s\ll 1 \;\Longrightarrow\; \bar w\to 1,\qquad \Delta p \approx s\,pq\bigl[\,ph+q(1-h)\,\bigr] \;\xrightarrow{\ h=1/2\ }\; \Delta p\approx\tfrac{1}{2}s\,pq \]
For weak selection the denominator differs from 1 by \(O(s)\), so dropping it costs only \(O(s^2)\) per generation. In the additive case the bracket equals \(\tfrac12(p+q)=\tfrac12\) identically, and the recursion loses all dependence on dominance. A
11
\[ \frac{dp}{dt}=\tfrac{1}{2}s\,p(1-p) \;\Longrightarrow\; \ln\!\frac{p_{t}}{1-p_{t}}=\ln\!\frac{p_{0}}{1-p_{0}}+\tfrac{1}{2}s\,t \;\Longrightarrow\; t=\frac{2}{s}\ln\!\left[\frac{p_{t}(1-p_{0})}{p_{0}(1-p_{t})}\right] \]
Treating generations as continuous when \(s\ll1\) turns Step 10 into the logistic equation, whose solution is linear in the log-odds. So an additive allele sweeps along a logistic curve and the log-odds of the favoured allele advance by exactly \(s/2\) per generation — the working formula for every waiting-time calculation. C
12
\[ h=0,\ s=1:\quad q'=\frac{pq+0}{1-q^{2}}=\frac{q(1-q)}{(1-q)(1+q)}=\frac{q}{1+q} \;\Longrightarrow\; \frac{1}{q_{t}}=\frac{1}{q_{0}}+t \]
For a lethal recessive the recursion integrates in closed form: taking reciprocals of \(q'=q/(1+q)\) gives \(1/q'=1/q+1\), so the reciprocal frequency advances by exactly one per generation. Decay is therefore hyperbolic, not exponential — \(q_t=q_0/(1+q_0t)\) — and removing the last copies takes arbitrarily long because selection can only see the vanishing \(q^2\) fraction hidden in homozygotes. C
Result
\[ \Delta p=\frac{pq\,(w_{A}-w_{a})}{\bar w}=\frac{s\,pq\,\bigl[\,ph+q(1-h)\,\bigr]}{1-2pq\,hs-q^{2}s} \]

Reading. The per-generation gain of the favoured allele is the product of three things: the heterozygosity-like variation term \(pq\), which is zero at both boundaries and maximal at \(p=\tfrac12\); the marginal fitness difference \(w_A-w_a\), which is what “strength of selection” actually means; and the normaliser \(\bar w\). The selection coefficient \(s\) is a dimensionless per-generation fitness ratio, not a rate of frequency change: \(s=0.1\) never means “the allele rises by 10% per generation”.

Scope. Exact for constant-viability selection at one autosomal locus in an infinite randomly mating population with discrete generations, for any \(s\in[0,1]\) and any \(h\). It is the deterministic limit of a stochastic process, and is trustworthy only while \(|s|\gg 1/(2N_e)\).

Corollaries & converses
  • Equilibria. \(\Delta p=0\) exactly when \(p=0\), \(p=1\), or \(w_A=w_a\). For \(0\le h\le1\) only the boundaries qualify, so the favoured allele fixes; an internal equilibrium requires \(h\lt0\) (heterozygote advantage) or \(h\gt1\) (heterozygote disadvantage).
  • Overdominance. With \(w_{AA}=1-s_1\), \(w_{Aa}=1\), \(w_{aa}=1-s_2\), the condition \(w_A=w_a\) gives \(\hat p=s_2/(s_1+s_2)\). It is stable: both alleles are permanently maintained, and the population pays a segregation load \(1-\bar w=s_1s_2/(s_1+s_2)\).
  • Underdominance. The same algebra with the heterozygote worst gives an internal equilibrium that is unstable — a threshold, above which \(A\) fixes and below which it is lost. This is the population-genetic basis of chromosomal rearrangements spreading only after a founder event.
  • Estimating \(s\) from data. Rearranging Step 6, \(w_A/w_a=\bigl[p'(1-p)\bigr]/\bigl[p(1-p')\bigr]\): the ratio of marginal fitnesses is the odds ratio of two successive allele frequencies, so \(s\) is measurable from allele counts alone without ever observing a death.
  • Haldane’s sieve. The bracket tends to \(1-h\) as \(p\to0\), so the initial spread of a rare favoured allele \(A\) is set entirely by the advantage it expresses in heterozygotes. If \(A\) is recessive (\(h\to1\), heterozygotes no fitter than \(aa\)) the bracket vanishes and selection is nearly blind to it; dominant beneficial alleles pass the sieve, recessive ones mostly do not.
  • Additive superposition fails. \(\Delta p\) is not linear in \(s\) except to first order; for \(s\) near 1 the \(\bar w\) denominator matters, and the naive \(\Delta p\approx\frac12 spq\) can be wrong by tens of percent.
Fails without
  • Finite population size — the drift regime: the recursion is a deterministic expectation, while drift adds a per-generation variance \(pq/(2N_e)\). When \(|s|\lesssim 1/(2N_e)\) the noise exceeds the signal and the allele behaves as if neutral. Kimura’s diffusion result for genic selection, \(u(p)=\bigl(1-e^{-4N_es_hp}\bigr)/\bigl(1-e^{-4N_es_h}\bigr)\), replaces certain fixation with a probability: a single new mutant with heterozygous advantage \(s_h\) in a population with \(N_e\approx N\) fixes with probability only about \(2s_h\), so a 1% beneficial mutation is lost 98 times out of 100 despite \(\Delta p\gt0\) at every frequency.
  • Frequency- or density-dependent fitness: with \(w\) a function of \(p\), Step 9’s gradient reading collapses because \(\bar w\) is no longer a fixed landscape. Under rare-morph advantage the same algebra yields a stable polymorphism at which mean fitness is not maximal, and mean fitness can decrease over time — a possibility rigorously excluded in the constant-fitness model.
  • Non-random mating: with inbreeding coefficient \(F\gt0\) the zygote frequencies of Step 1 become \(p^2+Fpq,\ 2pq(1-F),\ q^2+Fpq\), so a deleterious recessive is exposed at \(q^2+Fpq\) rather than \(q^2\). For \(q=0.01\) and \(F=0.05\) the exposed fraction rises from \(1.0\times10^{-4}\) to \(5.95\times10^{-4}\), a nearly sixfold increase in the rate of removal, and the \(1/q_t\) law of Step 12 no longer holds.
  • Selection through fertility or in overlapping generations: if fitness attaches to mating pairs rather than to individuals, or generations overlap so that the census is not a clean zygote pool, the state cannot be summarised by \(p\) alone and no single-locus recursion of this form exists.
Common errors
  • “\(s=0.02\), so the allele rises by 2% per generation.” \(s\) is a fitness ratio; the frequency change is \(\Delta p\approx\frac12 spq\), which at \(p=0.01\) is about \(10^{-4}\) — two hundred times smaller.
  • “Fitness is a property of an allele.” Selection acts on diploid genotypes; the allelic quantities \(w_A,w_a\) are marginal averages over genetic backgrounds (Step 4) and therefore change as \(p\) changes, even with the three genotype fitnesses held fixed.
  • “Selection will clear a deleterious recessive from the population.” Step 12 gives hyperbolic, not exponential, decay: even a lethal recessive at \(q=0.01\) needs 100 generations to reach \(q=0.005\), because almost every copy is sheltered in a heterozygote.
  • “Dropping \(\bar w\) is always safe.” True to \(O(s^2)\) for weak selection, badly wrong for strong selection: with \(h=0\), \(s=1\), \(q=0.5\) the denominator is \(0.75\), a 33% correction.
  • Confusing \(s\) with \(\mu\). Selection coefficients of interest run from \(10^{-4}\) to 1; per-locus mutation rates are of order \(10^{-6}\) to \(10^{-5}\). They enter mutation-selection balance asymmetrically: \(\hat q=\sqrt{\mu/s}\) for a recessive deleterious allele, \(\hat q\approx\mu/(hs)\) once heterozygotes pay a cost.
  • Reporting an estimated \(s\) without \(N_e\). A fitted \(s=0.001\) means opposite things in a population of \(N_e=10^6\) (strongly selected) and \(N_e=100\) (effectively neutral).
Discussion

The algebra above is the technical core of the modern synthesis. Haldane worked through the single-locus cases — dominant, recessive, sex-linked, with and without inbreeding — in his series “A mathematical theory of natural and artificial selection” beginning in 1924, and it was he who established the slow, hyperbolic elimination of recessives that Step 12 reproduces. Fisher’s The Genetical Theory of Natural Selection (1930) recast the same content as a statement about variance, and Wright’s gradient form (Step 9) supplied the adaptive-landscape picture that has dominated intuition ever since. That these three routes give the same recursion is the reason the one-locus model is treated as settled ground rather than one model among many.

Empirically, the useful direction is backwards. Given genotype counts at two life stages, or allele frequencies at two time points, the odds-ratio corollary returns \(w_A/w_a\) and hence \(s\); given ancient DNA or serial sampling in an experimental population, the log-odds line of Step 11 is fitted and its slope is \(s/2\). Selection coefficients recovered this way for the strongest known signals in human populations — lactase persistence, several malaria-resistance loci — come out of order \(10^{-2}\) per generation, which is precisely why they are detectable at all: at \(s=10^{-2}\) a sweep takes a few thousand generations, short enough to leave a haplotype signature and long enough to be plausible.

The deterministic model’s real boundary is not weak selection but the product \(N_es\). The diffusion approximation shows that fixation probability depends on \(p\) and \(N_es\) alone, so a “strongly selected” allele in a bottlenecked population and a “nearly neutral” one in a vast population can have identical dynamics. This is Ohta’s nearly neutral theory in one inequality, and it is why comparative genomics reads the efficacy of selection off effective population size: the same \(s\) is visible in a bacterium with \(N_e\sim10^8\) and invisible in a large mammal with \(N_e\sim10^4\).

Common misconceptions. First, that \(\Delta p\gt0\) guarantees fixation — it guarantees fixation only in the infinite-population idealisation, and most beneficial mutations are in fact lost while rare. Second, that the maximum of \(pq\) at \(p=\tfrac12\) means selection is “strongest” in the middle of a sweep; the fitness difference \(w_A-w_a\) is generally largest when the deleterious allele is common, and the two factors trade off, which is why sweeps are slow at both ends and fast in the middle. Third, that fixing a beneficial allele improves the population without cost: Haldane’s substitution load counts the deaths required, and it is this bookkeeping that motivated the neutralist argument in the first place.

Worked examples

Example 1. A fully recessive lethal allele (\(h=0\), \(s=1\)) sits at frequency \(q_0=0.10\) in a large randomly mating population. How many generations of complete selection against homozygotes are needed to halve it, and how many more to halve it again? Compare with the naive expectation from a single-generation calculation.

1
\[ \Delta q=-\frac{s\,p\,q^{2}}{1-sq^{2}}=-\frac{(1)(0.90)(0.10)^{2}}{1-(1)(0.10)^{2}}=-\frac{0.00900}{0.9900}=-9.09\times10^{-3} \]
Setting \(h=0\) in the Result and writing it for \(q\): the bracket becomes \(q\), so \(\Delta q=-spq^{2}/\bar w\) with \(\bar w=1-sq^{2}\). Only the \(q^2=1\%\) of the population that is homozygous is exposed, so a lethal allele at 10% falls by under one percentage point in the first generation. A
2
\[ \text{naive: } \frac{0.10-0.05}{9.09\times10^{-3}}\approx 5.5\ \text{generations to halve} \]
Extrapolating the first generation’s change linearly. This is wrong, and it is wrong in a predictable direction: \(\Delta q\) shrinks as \(q^2\) as the allele becomes rarer, so a constant-rate extrapolation always underestimates the time. A
3
\[ \frac{1}{q_{t}}=\frac{1}{q_{0}}+t \;\Longrightarrow\; t=\frac{1}{q_{t}}-\frac{1}{q_{0}} \]
Step 12 of the Proof gives the exact integral of the recursion for \(h=0,\ s=1\); rearranging for \(t\) turns the whole trajectory into subtraction of two reciprocals, with no iteration required. B
4
\[ t_{0.10\to0.05}=\frac{1}{0.05}-\frac{1}{0.10}=20-10=10\ \text{generations} \]
Substituting \(q_0=0.10\) and \(q_t=0.05\). The true answer is 10 generations, nearly twice the naive 5.5, and the frequency of affected homozygotes falls from \(1.00\times10^{-2}\) to \(2.50\times10^{-3}\) over that interval. A
5
\[ t_{0.05\to0.025}=\frac{1}{0.025}-\frac{1}{0.05}=40-20=20\ \text{generations},\qquad t_{0.01\to0.005}=200-100=100 \]
Each successive halving costs twice as long as the last, because the reciprocal frequency advances by a fixed one unit per generation. For a species with a 25-year generation time, the last of these halvings is 2500 years of complete lethality. B
\[ q_{t}=\frac{q_{0}}{1+q_{0}t}:\quad 0.10\to0.05\ \text{in }10\ \text{generations};\quad 0.01\to0.005\ \text{in }100 \]

Reading. Complete lethality of homozygotes — the strongest selection possible — removes a recessive allele only hyperbolically, because heterozygotes shelter a fraction \(2pq/(2pq+2q^{2})=p\to1\) of all copies of \(a\) as \(q\to0\). Eugenic proposals to eliminate recessive disease alleles by preventing affected individuals from reproducing founder on exactly this arithmetic.

Scope. Exact for \(h=0,\ s=1\) under the stated hypotheses; any heterozygote cost (\(h\gt0\)) accelerates removal dramatically, since the exposed fraction jumps from \(q^2\) to \(2hpq+q^2\).

Example 2. A cohort of 1000 newly settled marine invertebrate juveniles is genotyped at a single locus: 360 \(AA\), 480 \(Aa\), 160 \(aa\). At the end of the season the survivors are recounted: 324 \(AA\), 408 \(Aa\), 96 \(aa\). Estimate \(s\) and \(h\), predict \(\Delta p\) from the Result, and check it against the survivor counts.

1
\[ p=\frac{2(360)+480}{2000}=\frac{1200}{2000}=0.600,\qquad q=0.400 \]
Counting allele copies directly: each \(AA\) contributes two \(A\) copies and each \(Aa\) one, out of \(2\times1000\) copies. The expected Hardy-Weinberg counts \(p^2,2pq,q^2\) are \(360,480,160\) — identical to the observed juveniles, so the census is at the zygote stage required by Step 1. A
2
\[ \ell_{AA}=\frac{324}{360}=0.900,\qquad \ell_{Aa}=\frac{408}{480}=0.850,\qquad \ell_{aa}=\frac{96}{160}=0.600 \]
Genotype-specific survival probabilities, the absolute viabilities. These are the raw measurement; relative fitness is what the model uses, so they must be rescaled next. A
3
\[ w_{AA}=1,\qquad w_{Aa}=\frac{0.850}{0.900}=0.9444,\qquad w_{aa}=\frac{0.600}{0.900}=0.6667 \]
Dividing by the largest viability sets the fittest genotype to 1, the convention used in Step 7. Relative fitnesses are defined only up to a common factor, so this rescaling changes nothing in \(\Delta p\) — the factor cancels between numerator and \(\bar w\). B
4
\[ s=1-w_{aa}=0.3333,\qquad h=\frac{1-w_{Aa}}{s}=\frac{0.05556}{0.3333}=0.1667 \]
Reading the parametrisation of Step 7 backwards. So selection against \(aa\) removes a third of its relative fitness, and the allele is largely but not completely recessive: heterozygotes carry only one sixth of the homozygous penalty. B
5
\[ \bar w=1-2pq\,hs-q^{2}s=1-2(0.6)(0.4)(0.1667)(0.3333)-(0.16)(0.3333)=1-0.02667-0.05333=0.9200 \]
Mean relative fitness from Step 7. Equivalently \(828/1000\) survivors divided by the reference viability \(0.900\), which gives \(0.920\) — the two routes agree, confirming the rescaling of Step 3. A
6
\[ \Delta p=\frac{s\,pq\,[\,ph+q(1-h)\,]}{\bar w}=\frac{(0.3333)(0.24)\bigl[(0.6)(0.1667)+(0.4)(0.8333)\bigr]}{0.9200}=\frac{(0.08000)(0.4333)}{0.9200}=0.03768 \]
Substituting into the Result, symbols first: the bracket is \(0.1000+0.3333=0.4333\). Predicted new frequency \(p'=0.600+0.038=0.6377\). B
7
\[ p'_{\text{obs}}=\frac{2(324)+408}{2(828)}=\frac{1056}{1656}=0.6377 \;\Longrightarrow\; \Delta p_{\text{obs}}=0.0377 \]
Counting alleles among the 828 survivors directly. Prediction and observation agree to four figures, as they must: the Result is not an approximation but an exact rewriting of the same counting. The agreement is a check on the algebra, not evidence about the population. A
8
\[ \frac{w_{A}}{w_{a}}=\frac{p'(1-p)}{p(1-p')}=\frac{(0.6377)(0.400)}{(0.600)(0.3623)}=\frac{0.2551}{0.2174}=1.173 \]
The odds-ratio corollary, applied without using the genotype-level survival data at all. It recovers the marginal fitness advantage of \(A\) from two allele frequencies alone — the estimator used on time-series and ancient-DNA data where genotype-specific mortality is never observed. B
\[ s=0.333,\quad h=0.167,\quad \bar w=0.920,\quad \Delta p=+0.0377\ \text{per generation} \]

Reading. A single season of quite strong selection (a third of the fitness of \(aa\) lost) moves the allele frequency by under four percentage points, from 0.600 to 0.638. Selection coefficients large enough to be measured in one field season still produce modest per-generation frequency changes.

Scope. The estimates are point estimates from one cohort with no sampling error attached; with 1000 juveniles the binomial standard error on \(p'\) is about \(\sqrt{p'(1-p')/1656}\approx0.012\), so \(s\) here is resolved only to roughly the first decimal place.

Problems
  1. An additive allele (\(h=\tfrac12\)) has \(s=0.10\) and is currently at \(p=0.50\). Compute \(\Delta p\) exactly from the Result and from the weak-selection approximation of Step 10, and give the percentage error of the approximation.
    Solution

    Bracket: \(ph+q(1-h)=0.5(0.5)+0.5(0.5)=0.500\). Denominator: \(\bar w=1-2pq\,hs-q^2s=1-2(0.25)(0.5)(0.1)-(0.25)(0.1)=1-0.0250-0.0250=0.9500\).

    Exact: \(\Delta p=(0.10)(0.25)(0.500)/0.9500=0.01250/0.9500=0.013158\), so \(p'=0.51316\).

    Approximation: \(\Delta p\approx\tfrac12 spq=\tfrac12(0.10)(0.25)=0.01250\).

    Error \(=(0.01250-0.013158)/0.013158=-0.050\), i.e. the approximation is 5.0% low — exactly the \(O(s)\) size of the neglected \(\bar w\), since \(1-\bar w=0.05\).

  2. A deleterious allele sits at \(q=0.05\) with \(s=0.20\). Compute \(\Delta q\) (i) if it is fully recessive (\(h=0\)) and (ii) if it is fully dominant (\(h=1\)), and explain the ratio in one sentence.
    Solution

    (i) \(h=0\): \(\bar w=1-sq^2=1-(0.2)(0.0025)=0.99950\); \(\Delta q=-spq^2/\bar w=-(0.2)(0.95)(0.0025)/0.99950=-4.75\times10^{-4}/0.99950=-4.752\times10^{-4}\).

    (ii) \(h=1\): \(\bar w=1-2pq s-q^2s=1-sq(2p+q)=1-(0.2)(0.05)(1.95)=0.98050\); \(\Delta q=-sp^2q/\bar w=-(0.2)(0.9025)(0.05)/0.98050=-9.025\times10^{-3}/0.98050=-9.204\times10^{-3}\).

    Ratio \(=9.204\times10^{-3}/4.752\times10^{-4}=19.4\), essentially \(p/q=19\). The marginal fitness difference is \(w_A-w_a=s\,[\,ph+q(1-h)\,]\), which is \(sp\) for a fully dominant deleterious allele and \(sq\) for a fully recessive one, so the two \(\Delta q\) values stand in the ratio \(p/q\): a rare dominant is exposed to selection in every heterozygous carrier, a rare recessive only in the far rarer \(q^{2}\) homozygotes.

  3. A lethal recessive allele (\(h=0\), \(s=1\)) is at \(q_0=0.02\). Find \(q\) after 50 and after 150 generations, the frequency of affected homozygotes at each time, and the number of generations needed to reach \(q=0.001\).
    Solution

    Step 12: \(1/q_t=1/q_0+t=50+t\).

    \(t=50\): \(1/q=100\), \(q=0.0100\), affected \(q^2=1.00\times10^{-4}\) (1 in 10,000).

    \(t=150\): \(1/q=200\), \(q=0.00500\), affected \(q^2=2.50\times10^{-5}\) (1 in 40,000).

    \(q=0.001\): \(1/q=1000\), so \(t=1000-50=950\) generations.

    Halving the allele frequency from 0.02 to 0.01 takes 50 generations; the next halving takes 100; reaching a twentieth of the starting value takes 950. The affected phenotype frequency falls by a factor of 400 over those 950 generations, which is why complete selection against a recessive phenotype looks effective early and stalls indefinitely afterwards.

  4. Suppose a locus shows heterozygote advantage with illustrative relative fitnesses \(w_{AA}=0.88\), \(w_{AS}=1\), \(w_{SS}=0.20\) (the sickle-cell pattern in a high-malaria environment; the numbers here are round illustrative values, not measured estimates). Find the equilibrium frequency of \(S\), verify it is stable by evaluating \(\Delta p\) on either side, and compute the mean fitness and segregation load at equilibrium.
    Solution

    Write \(s_1=1-w_{AA}=0.12\) (against \(AA\)) and \(s_2=1-w_{SS}=0.80\) (against \(SS\)). Let \(p\) be the frequency of \(A\), \(q\) that of \(S\). Marginal fitnesses: \(w_A=p(1-s_1)+q\) and \(w_S=p+q(1-s_2)\), so \(w_A-w_S=q s_2-p s_1\).

    Equilibrium: \(\hat p s_1=\hat q s_2\Rightarrow \hat p=s_2/(s_1+s_2)=0.80/0.92=0.8696\), \(\hat q=s_1/(s_1+s_2)=0.12/0.92=0.1304\).

    Stability: at \(q=0.10\ (\lt\hat q)\), \(w_A-w_S=(0.10)(0.80)-(0.90)(0.12)=0.080-0.108=-0.028\lt0\), so \(\Delta p\lt0\) and \(q\) rises. At \(q=0.20\ (\gt\hat q)\), \(w_A-w_S=(0.20)(0.80)-(0.80)(0.12)=0.160-0.096=+0.064\gt0\), so \(q\) falls. The equilibrium is approached from both sides: it is stable.

    Mean fitness: \(\bar w=\hat p^2(0.88)+2\hat p\hat q(1)+\hat q^2(0.20)=0.7561(0.88)+2(0.11342)+0.01701(0.20)=0.6654+0.22684+0.00340=0.8956\). Segregation load \(=1-\bar w=0.1044\), matching \(s_1s_2/(s_1+s_2)=(0.12)(0.80)/0.92=0.1043\).

    At equilibrium 1.7% of births are \(SS\) and 22.7% are carriers — the polymorphism is maintained by selection itself, not by mutation, and it costs the population about 10% of its mean fitness.

  5. An additive beneficial allele (\(h=\tfrac12\)) with \(s=0.005\) starts at \(p_0=0.02\). (a) How many generations until \(p=0.50\)? (b) Repeat for \(s=0.05\). (c) For \(N_e=10^4\), is the deterministic treatment justified, and what is the fixation probability of a single new copy of this allele when \(s=0.005\)?
    Solution

    (a) Step 11: \(t=\dfrac{2}{s}\ln\!\left[\dfrac{p_t(1-p_0)}{p_0(1-p_t)}\right]=\dfrac{2}{0.005}\ln\!\left[\dfrac{(0.50)(0.98)}{(0.02)(0.50)}\right]=400\ln(49)=400(3.8918)=1557\) generations.

    (b) The log-odds term is unchanged, so \(t=\dfrac{2}{0.05}(3.8918)=40(3.8918)=156\) generations — waiting time scales as \(1/s\).

    (c) \(2N_es=2(10^4)(0.005)=100\gg1\), so selection dominates drift once the allele is common and the deterministic trajectory is a good description; at \(p_0=0.02\) there are already \(2N_ep_0=400\) copies, far from the stochastic boundary.

    Fixation of a single new copy: rescale so the resident \(aa\) has fitness 1. Then the heterozygote has \(w_{Aa}/w_{aa}=(1-s/2)/(1-s)\approx1+s/2\), i.e. \(s_h=0.0025\). Haldane’s approximation with \(N_e\approx N\) gives \(u\approx2s_h=0.0050\): about 1 new beneficial copy in 200 fixes, and 199 are lost to drift while rare, even though \(\Delta p\gt0\) at every frequency.