Hardy-Weinberg equilibrium
Statement
Consider one autosomal locus carrying alleles \(A\) and \(a\) in a diploid population with discrete, non-overlapping generations, and suppose meiosis is fair, fertility is genotype-independent, allele frequencies are equal in the two sexes, mating is random with respect to the locus, and no mutation, migration, selection or sampling drift acts between the gamete pool and the census of the next generation. Let \(p\) and \(q=1-p\) be the frequencies of \(A\) and \(a\) in the parental gamete pool. Then the zygotes of the next generation carry \(AA\), \(Aa\), \(aa\) in the frequencies \(p^2\), \(2pq\), \(q^2\) — whatever the parental genotype frequencies were — the allele frequency is unchanged, \(p'=p\), and both allele and genotype frequencies are then fixed in every subsequent generation. The equilibrium is attained in exactly one generation of random mating, and it is a one-parameter family of neutrally stable fixed points indexed by \(p\), not an attractor towards which a disturbed population is pulled back.
Why it matters
Population genetics is the study of how allele frequencies change, and no change can be measured without a statement of what “no change” looks like. Hardy–Weinberg supplies it, and supplies it in an unusually strong form: it does not merely assert that frequencies stay put, it computes the exact genotype frequencies that Mendelian segregation (The law of segregation) plus random union of gametes is obliged to produce, from any starting genotype distribution, in a single generation. Everything else in the subject is then written as a departure from it. Genetic drift is what happens when the population is finite, Gene flow when it is not closed, Mutation–selection balance when alleles are created and removed, Inbreeding and heterozygosity when mating is not random, and Natural selection when genotypes differ in fitness. Each of these is quantified by a coefficient that is defined as a deviation from the proportions derived below.
The practical consequences run in both directions. Forwards, the result converts an allele frequency into predicted genotype frequencies, which is how a carrier frequency is estimated from the incidence of a recessive disease, how a genetic counsellor turns a population statistic into a couple’s risk, and how the expected heterozygosity of a marker is computed before an experiment is run (Genomic medicine). Backwards, it converts observed genotype counts into a test: a locus whose genotype counts deviate significantly from \(p^2, 2pq, q^2\) is telling you that one of the hypotheses fails there, and modern genome-wide studies run exactly this test at every one of a million markers as their first quality filter — a marker out of Hardy–Weinberg proportions in healthy controls is usually not evidence of biology but evidence that the genotyping assay is misreading heterozygotes. The same statistic, read in the other direction, is the standard detector of cryptic population structure and of recent admixture.
Hypotheses
Proof
The argument has three parts. Steps 1–5 push one generation forward: adults \(\to\) gamete pool \(\to\) zygotes. Steps 6–8 show that the resulting state is fixed, identify the whole equilibrium set, and give the single coordinate \(F\) that measures departure from it. Steps 9–12 remove the idealisations that were assumed for convenience — equal frequencies in the sexes, autosomal inheritance, two alleles, an infinite sample — and produce the statistic by which the result is tested on real counts.
Result
Reading. Random union of gametes throws away every trace of the parental genotype distribution and keeps only the allele frequency; the genotype frequencies are then whatever the binomial expansion of \((p+q)^2\) says they are. The population does not drift towards this state over many generations — it arrives in one, and then stays, because the map that produces it is a projection.
Scope. All quantities are dimensionless frequencies; counts are in individuals and time is in generations. One autosomal locus, diploid, discrete generations, sexes equal in allele frequency, fair meiosis, genotype-independent fertility, random union of gametes, and — for the constancy across generations, not for the proportions themselves — an infinite closed population free of mutation and selection. Two alleles were assumed only for readability (Step 11); unequal sexes cost one extra generation (Step 9); X linkage replaces the one-generation result by geometric convergence with ratio \(-\tfrac12\) (Step 10).
Corollaries & converses
- Carrier excess for a rare recessive. \(\dfrac{2pq}{q^2}=\dfrac{2p}{q}\approx\dfrac{2}{q}\) for small \(q\): an allele causing disease in \(1\) in \(10^4\) births has \(q=10^{-2}\) and sits in about \(200\) unaffected carriers for every affected homozygote. Equivalently, a fraction \(p\) of all copies of a recessive allele is hidden in heterozygotes, which is why selection against recessives is so slow (Mutation–selection balance).
- Maximum heterozygosity. \(2pq\le\tfrac12\), with equality only at \(p=q=\tfrac12\); with \(k\) alleles the ceiling is \(1-1/k\). No diploid locus can be more than half heterozygous on two alleles, however the frequencies are arranged.
- Estimating \(q\) from a phenotype. If \(aa\) is the only genotype expressing the recessive phenotype, \(\hat q=\sqrt{\text{incidence}}\) — the only route to an allele frequency when heterozygotes are phenotypically invisible, and one that is valid only under the hypotheses, since the square root of an incidence inflated by inbreeding overestimates \(q\).
- X-linked sex ratio of affection. At equilibrium the affected male : affected female ratio is \(q:q^2\), i.e. \(1/q\). For \(q=0.01\) that is \(100:1\), which is why pedigree evidence of a strong male bias is itself evidence of X linkage.
- The Wahlund effect. Pooling subpopulations that are each internally in Hardy–Weinberg proportions gives \(F=\operatorname{Var}(p)/\left(\bar p\bar q\right)\ge0\): a heterozygote deficit is manufactured by the pooling alone, with no inbreeding and no selection anywhere in the system.
- The test is the coefficient. \(\chi^2=n\hat F^{\,2}\) exactly (Step 12), so the smallest departure a sample of \(n\) can reject at the five per cent level is \(|F|=\sqrt{3.841/n}\) — about \(0.062\) at \(n=1000\). “Consistent with Hardy–Weinberg” from a small sample is a weak statement, and the sample size fixes exactly how weak.
- Converse fails. Observing \(p^2,2pq,q^2\) does not establish the hypotheses. Selection acting only on the zygote-to-adult transition leaves newborns in exact Hardy–Weinberg proportions every generation while \(p\) marches steadily (Problem 5); two departures of opposite sign can cancel; and a locus with a very common allele has almost no power to detect anything. The implication runs one way only.
- Contrast with linkage equilibrium. Randomisation within a locus is complete after one meiosis, but randomisation between loci is not: gametic disequilibrium decays as \(D_t=(1-r)^tD_0\), taking many generations for tightly linked loci (Linkage and recombination). A population can therefore sit exactly on \(\mathcal{H}\) at every locus and still carry strong associations between them, which is precisely what association mapping exploits.
Fails without
- Finite population — the drift regime: with \(N\) diploids the next generation’s allele frequency is a binomial sample of \(2N\) gametes, so \(\mathbb{E}\left(p'\right)=p\) but \(\operatorname{Var}\left(p'\right)=pq/(2N)\), and expected heterozygosity decays as \(H_t=H_0\left(1-1/(2N)\right)^t\), with every locus eventually fixed or lost. The genotype proportions survive — each generation is still binomial at whatever \(p\) currently is — but the constancy does not, and for \(N=50\) heterozygosity halves in about \(70\) generations. This is the regime of every real conservation population and the reason effective population size is the central parameter of Genetic drift and The neutral theory.
- Population subdivision — the Wahlund regime: take two subpopulations of equal size in Hardy–Weinberg proportions at \(p_1=0.9\) and \(p_2=0.3\) and pool them. The pooled sample has \(\bar p=0.6\), \(\operatorname{Var}(p)=0.09\), hence \(F=0.09/(0.6\times0.4)=0.375\) and genotype frequencies \(0.45, 0.30, 0.25\) against expectations \(0.36,0.48,0.16\) — a \(37.5\) per cent heterozygote deficit produced by geography alone. Nothing inside either subpopulation is out of equilibrium; the deficit is an artefact of the sampling boundary, and it is the basis of \(F_{ST}\) as a measure of differentiation (Gene flow).
- Non-random mating — the inbreeding regime: mating between relatives makes the two uniting gametes positively correlated, which is what \(F\) measures in its original sense. Genotype frequencies move to \(p^2+Fpq,\ 2pq(1-F),\ q^2+Fpq\) while \(p\) does not move at all — so inbreeding changes who carries the alleles, not how many there are, and its cost is the exposure of recessive homozygotes, whose frequency rises from \(q^2\) to \(q^2+Fpq\). For a rare allele with \(q=0.001\) and first-cousin mating (\(F=1/16\)), homozygote frequency rises from \(10^{-6}\) to about \(6.3\times10^{-5}\), a \(63\)-fold increase (Inbreeding and heterozygosity).
- Selection between census points — the mismatched-stage regime: if genotypes differ in viability, the newborn cohort is in Hardy–Weinberg proportions and the adult cohort is not, so the answer depends on when you sample. With relative fitnesses \(1,1,1-s\) acting on Hardy–Weinberg zygotes, the zygote cohort has mean fitness \(\bar w=1-sq^2\) — which is also the fraction of it that survives — and the survivors have fixation index \(F=1-\bar w/(1-sq)=-spq/(1-sq)\) — a heterozygote excess, because only homozygotes were culled, and one of exactly the size worked out numerically in Problem 5. A study that genotypes adults and reports a deviation has not necessarily found non-random mating.
- Assay error — the technical regime: a null allele at a primer site, or a probe intensity cluster that merges heterozygotes into a homozygote call, mimics inbreeding perfectly: a heterozygote deficit at one locus while neighbouring loci are clean. Genome-wide studies therefore discard markers failing the test in controls rather than interpreting them, since a genuine biological signal would be expected to show in cases and controls alike and to be shared by markers in linkage disequilibrium with it.
Common errors
- “The population is in Hardy–Weinberg equilibrium, so it is not evolving.” The statement is about one locus, and it is an implication in one direction only. Selection acting after the zygote stage leaves newborns exactly on \(\mathcal{H}\) every generation while \(p\) changes monotonically (Problem 5); the test at \(n=1000\) cannot even detect \(s=0.2\).
- “\(p^2+2pq+q^2=1\) is the theorem.” That identity is \((p+q)^2=1\), true of any two numbers summing to one and carrying no biology at all. The theorem is the assertion that the genotype frequencies equal those three terms — which is Steps 2–4, and which can be false.
- “Take the square root of the heterozygote frequency to get \(q\).” Only the homozygote frequencies are squares. From \(H\) one recovers \(q\) by solving \(2q(1-q)=H\), which has two roots; from countable genotypes one should simply count alleles, \(\hat q=(2n_{aa}+n_{Aa})/(2n)\).
- “The chi-square test has two degrees of freedom because there are three classes.” One degree of freedom is spent estimating \(\hat p\) from the same data, leaving \(1\); using \(2\) inflates the critical value from \(3.84\) to \(5.99\) and hides real departures. With \(k\) alleles the count is \(k(k-1)/2\).
- “Hardy–Weinberg requires an infinite population, so it never applies.” Infinity is used only to make the realised frequencies equal the probabilities (Step 5). A finite population produces the proportions in expectation and deviates by \(O(N^{-1/2})\), which is exactly the noise the test is calibrated against; what finiteness really costs is the constancy of \(p\).
- “Random mating is required.” Random mating with respect to this locus is required (Step 3). Human populations mate highly assortatively on many traits and still sit on \(\mathcal{H}\) at the great majority of loci.
- “A dominant allele will spread.” Dominance appears nowhere in Step 6. This was the actual question put to Hardy, and the answer is that dominance affects which phenotypes appear, never the arithmetic of transmission.
- “Pool the sexes, or pool the sampling sites, to increase \(n\).” Pooling groups with different allele frequencies manufactures a heterozygote deficit (Wahlund) whose size grows with the variance in \(p\); the extra sample size buys power to detect an artefact of the pooling.
Discussion
The result was published twice in 1908. Wilhelm Weinberg, a physician in Stuttgart, gave it in a German paper on inheritance in twins and human traits; G. H. Hardy, who was not a biologist, wrote a short letter to Science after R. C. Punnett described to him a claim made in discussion — that brachydactyly, being dominant, ought steadily to increase until three quarters of the population were affected. Hardy’s letter is a page of algebra amounting to Steps 2, 4 and 6, and it opens by apologising for the triviality of the mathematics. The apology is well taken and beside the point: the difficulty was never the algebra but the framing, and the theorem’s real content is the identification of the correct null model. Once that model exists, deviation from it becomes measurable, and the entire quantitative apparatus of population genetics — \(F\)-statistics, effective population size, tests for selection — is built on measuring deviations from it.
Two structural features deserve more emphasis than they usually receive. First, the one-generation convergence is a consequence of \(T\) being a projection (Step 7), and it makes Hardy–Weinberg qualitatively unlike almost every other equilibrium in biology: there is no relaxation time to wait out, no transient, and no basin of attraction. Second, the equilibrium set is a curve of neutrally stable fixed points rather than a single stable one, so the theorem says nothing whatever about which \(p\) a population should have. That is a feature: it cleanly separates the part of the dynamics fixed by Mendelian mechanics (the genotype proportions) from the part that requires evolutionary explanation (the allele frequency), and it is exactly this separation that lets a single number \(F\) carry all information about departures.
The modern use of the result is almost entirely as a diagnostic instrument, and its power properties matter more than its truth. Because \(\chi^2=n\hat F^{\,2}\), power depends on sample size and on \(\hat p\) but not otherwise on the mechanism, and at a locus with a minor allele frequency of one per cent even large samples detect only gross departures — which is why rare-variant genotyping is filtered on other criteria. Conversely, at a million markers the test is performed a million times, so the significance threshold used in practice is severe (typically \(10^{-6}\) or lower in controls) and is chosen to control the false-discovery burden rather than to reflect any biological expectation. Exact tests based on the conditional distribution of \(n_{Aa}\) given the allele counts are preferred to the chi-square approximation whenever an expected cell count falls below about five, which for rare alleles is the normal situation.
Common misconceptions. That the equilibrium describes a population rather than a locus in a population at a stage of its life cycle: the same cohort can be on \(\mathcal{H}\) as zygotes and off it as adults, with no assumption violated except the reader’s about when the census was taken. And that failing the test identifies which hypothesis failed — it does not. A significant \(\chi^2\) says only that the joint model is wrong; distinguishing subdivision from inbreeding from null alleles requires several loci, and the pattern across loci is what carries the information, since a genotyping fault affects one marker while population structure affects them all.
Worked examples
Example 1. The MN blood group is determined by a single autosomal locus with two codominant alleles, so all three genotypes are directly countable. A sample of \(n=1000\) adults from one population gives \(298\) \(MM\), \(489\) \(MN\) and \(213\) \(NN\). Estimate the allele frequencies, test the sample against Hardy–Weinberg proportions, and state what the outcome does and does not license.
Reading. The \(1000\) individuals are distributed across the three genotypes almost exactly as random union of gametes at \(p=0.5425\) requires; the residual heterozygote deficit is one and a half per cent of the expected heterozygosity and is well within sampling noise.
Scope. A single autosomal codominant locus in one sample at one time. The test constrains \(|F|\) to lie below about \(0.06\) and constrains nothing else: it cannot distinguish which hypothesis might be failing, and a departure of the size produced by moderate inbreeding or modest subdivision would pass unnoticed at this sample size.
Example 2. Red–green colour vision deficiency is X-linked recessive, and its frequency among males of European ancestry is close to \(8\) per \(100\). (a) Estimate the allele frequency and predict the frequencies of affected and carrier females and the sex ratio of affection. (b) Suppose such a population were founded from males with allele frequency \(q_m^{(0)}=0.20\) and females with \(q_f^{(0)}=0.02\). Find the equilibrium frequency and the number of generations of random mating needed before the sexes differ by less than \(0.005\).
Reading. An X-linked allele at \(8\) per cent affects \(8\) per cent of men but only \(0.64\) per cent of women, while \(15\) per cent of women carry it unexpressed; and a founding population with the sexes badly mismatched converges on its permanent frequency, oscillating, with the mismatch halving every generation.
Scope. Red–green deficiency is in reality a composite of several variants of the \(OPN1LW\) and \(OPN1MW\) genes; the arithmetic here treats them as a single recessive allele class, which is legitimate for aggregate frequencies but not for predicting a specific molecular diagnosis. The \(8\) per cent figure is an approximate, ancestry-specific value, and the frequency differs substantially between populations.
Problems
- A recessive condition affects \(1\) in \(2500\) newborns in a large, randomly mating population. Estimate the allele frequency and the carrier frequency, find the ratio of carriers to affected individuals, and compute what fraction of all copies of the recessive allele is carried by unaffected heterozygotes. Comment on the consequence for selection against the allele.
Solution
Under the Result the incidence of the recessive phenotype is \(q^2\), so \(\hat q=\sqrt{1/2500}=\sqrt{4.0\times10^{-4}}=0.020\) and \(\hat p=0.980\).
Carrier frequency: \(2pq=2(0.980)(0.020)=0.0392\), i.e. \(3.92\) per cent, or about \(1\) in \(25.5\) people.
Ratio, in symbols before numbers: \(\dfrac{2pq}{q^2}=\dfrac{2p}{q}=\dfrac{2(0.980)}{0.020}=98\). Ninety-eight carriers for every affected individual.
Fraction of \(a\) copies in heterozygotes: a heterozygote carries one \(a\) copy and a homozygote two, so the fraction is \(\dfrac{\tfrac12(2pq)}{q}=\dfrac{pq}{q}=p=0.980\). Ninety-eight per cent of the recessive alleles in the population are invisible to selection, sheltered in carriers.
Consequence: selection can only act on the \(2\) per cent of copies that are exposed as homozygotes, so even lethal recessives are removed extremely slowly. Removing every affected individual each generation changes \(q\) by only \(\Delta q=-q^2(1-q)/(1-q^2)\approx-q^2\approx-4\times10^{-4}\) per generation at this frequency — a relative reduction of about \(2\) per cent per generation, and slower still as \(q\) falls. This is why deleterious recessives persist at appreciable frequencies (Mutation–selection balance).
- A sample of \(1000\) individuals genotyped at one autosomal locus gives \(450\) \(AA\), \(300\) \(Aa\) and \(250\) \(aa\). (a) Test the sample against Hardy–Weinberg proportions. (b) Supposing the sample is an equal-sized pooling of two subpopulations that are each internally in Hardy–Weinberg proportions, find their allele frequencies and verify the decomposition. (c) State what happens if the pooled population then mates at random for one generation.
Solution
(a) Count alleles: \(\hat p=(2\times450+300)/2000=1200/2000=0.600\), \(\hat q=0.400\). Expected: \(n\hat p^2=360\), \(2n\hat p\hat q=480\), \(n\hat q^2=160\). Observed minus expected is \(+90,\ -180,\ +90\) — the \(1:-2:1\) pattern of Step 8.
\(\hat F=1-300/480=1-0.625=0.375\), so by Step 12, \(\chi^2=n\hat F^{\,2}=1000(0.375)^2=140.6\) on \(1\) degree of freedom. Against \(3.841\) this is overwhelming (\(P\approx10^{-32}\)): a \(37.5\) per cent deficit of heterozygotes.
(b) Wahlund: \(F=\operatorname{Var}(p)/(\bar p\bar q)\), so \(\operatorname{Var}(p)=0.375\times0.600\times0.400=0.0900\). For two equal-sized groups with frequencies \(\bar p\pm\delta\) the variance is \(\delta^2\), hence \(\delta=\sqrt{0.0900}=0.300\) and \(p_1=0.900\), \(p_2=0.300\).
Verification. Subpopulation 1 (\(500\) individuals at \(p=0.9\)): \(0.81,0.18,0.01\), i.e. \(405,\ 90,\ 5\). Subpopulation 2 (\(500\) at \(p=0.3\)): \(0.09,0.42,0.49\), i.e. \(45,\ 210,\ 245\). Totals: \(450,\ 300,\ 250\) — exactly the observed counts. Each subpopulation satisfies Hardy–Weinberg exactly; the pooled sample fails it decisively.
(c) One generation of random mating across the merged population sets \(F\) to zero (Step 7) at the unchanged allele frequency \(\bar p=0.600\), giving \(360,\ 480,\ 160\). The deficit disappears in a single generation, which is itself a diagnostic: a heterozygote deficit that persists after admixture indicates continuing structure or non-random mating rather than a historical merger.
- A population is founded with allele frequencies \(p_f=0.80\) in females and \(p_m=0.40\) in males at an autosomal locus. Find the genotype frequencies among the offspring, the allele frequency in that generation, the value of \(F\), and the genotype frequencies one generation later. Verify the identity \(F=-(p_f-p_m)^2/(4\bar p\bar q)\), and say whether a sample of \(500\) offspring would detect the departure.
Solution
Offspring genotype frequencies are products of the two sex-specific gamete pools (Step 9): \(P'=p_fp_m=0.80\times0.40=0.320\); \(H'=p_fq_m+q_fp_m=(0.80)(0.60)+(0.20)(0.40)=0.480+0.080=0.560\); \(Q'=q_fq_m=(0.20)(0.60)=0.120\). Sum \(=1.000\).
Allele frequency: \(p'=0.320+\tfrac12(0.560)=0.600=\bar p=\tfrac12(0.80+0.40)\), as required, and it is the same in both sexes because autosomal transmission is symmetric.
\(F=1-H'/(2p'q')=1-0.560/(2\times0.600\times0.400)=1-0.560/0.480=1-1.1667=-0.1667\): a heterozygote excess of one sixth.
Identity check: \(-(p_f-p_m)^2/(4\bar p\bar q)=-(0.40)^2/(4\times0.600\times0.400)=-0.160/0.960=-0.1667\). Agreement.
Next generation: both sexes now have \(p=0.600\), so one round of random mating gives \(0.360,\ 0.480,\ 0.160\) — Hardy–Weinberg proportions attained in generation two, not generation one.
Detection: \(\chi^2=n F^2=500\times(0.1667)^2=500\times0.02778=13.9\) on \(1\) degree of freedom, far above \(3.841\) (\(P\approx2\times10^{-4}\)). Yes, easily detected — and a heterozygote excess of this size in a founder population is a much better clue to sex-biased founding than to any exotic biology.
- The ABO locus carries three alleles: \(I^A\) and \(I^B\), codominant with each other, and \(i\), recessive to both, at frequencies \(p,q,r\) with \(p+q+r=1\). A sample of \(1000\) gives phenotypes \(O:490\), \(A:320\), \(B:150\), \(AB:40\). (a) Write the four phenotype frequencies in terms of \(p,q,r\). (b) Estimate the three allele frequencies. (c) State the number of degrees of freedom a Hardy–Weinberg test would carry here, and compute the expected heterozygosity. (d) Say what a non-zero discrepancy \(D=1-(\hat p+\hat q+\hat r)\) would mean.
Solution
(a) From Step 11, genotype frequencies are \(p^2, q^2, r^2, 2pq, 2pr, 2qr\). Grouping by phenotype: \(O=r^2\); \(A=p^2+2pr\); \(B=q^2+2qr\); \(AB=2pq\). These sum to \((p+q+r)^2=1\).
(b) Note that \(A+O=p^2+2pr+r^2=(p+r)^2\) and \(B+O=(q+r)^2\). Solve in symbols first: \(r=\sqrt{O}\), \(p=1-\sqrt{B+O}\), \(q=1-\sqrt{A+O}\).
Numerically: \(\hat r=\sqrt{0.490}=0.700\); \(\hat p=1-\sqrt{0.150+0.490}=1-\sqrt{0.640}=1-0.800=0.200\); \(\hat q=1-\sqrt{0.320+0.490}=1-\sqrt{0.810}=1-0.900=0.100\). Check: \(0.200+0.100+0.700=1.000\).
Independent check on the model: predicted \(AB=2pq=2(0.200)(0.100)=0.040\), i.e. \(40\) individuals, matching the observed \(40\) exactly. The \(AB\) class was not used in estimating the frequencies, so this is a genuine test of fit, and here the fit is exact.
(c) Four phenotype classes, minus one for the total, minus two free allele frequencies estimated (three frequencies constrained to sum to one) leaves \(4-1-2=1\) degree of freedom. Equivalently, by Step 11 with \(k=3\) alleles there are \(k(k-1)/2=3\) degrees of freedom for a genotype-level test, but phenotypes here merge \(AA\) with \(AO\) and \(BB\) with \(BO\), removing two.
Expected heterozygosity: \(H=1-\sum_ip_i^2=1-(0.040+0.010+0.490)=1-0.540=0.460\).
(d) The square-root estimators are consistent only if the population really is in Hardy–Weinberg proportions, and they are not maximum-likelihood estimators, so \(\hat p+\hat q+\hat r\) need not equal one in a real sample. A small \(D\) reflects sampling error; a large one indicates departure from the model — population structure, non-random mating, or misclassified phenotypes. Bernstein’s correction rescales the three estimates by \(1+D/2\) (adding \(D/2\) to \(\hat r\) first) to restore the constraint; a maximum-likelihood fit by expectation–maximisation over the unobserved \(AA\)/\(AO\) and \(BB\)/\(BO\) split is the modern alternative.
- At an autosomal locus with \(p=0.600\) and \(q=0.400\), zygotes are formed in Hardy–Weinberg proportions but \(aa\) individuals have relative viability \(1-s\) with \(s=0.200\), the other two genotypes having viability \(1\). (a) Find the genotype and allele frequencies among the surviving adults. (b) Compute \(F\) for the adults and the \(\chi^2\) a sample of \(1000\) adults would give; state whether the selection would be detected. (c) Find the genotype frequencies of the next generation of zygotes and explain the general lesson.
Solution
(a) Zygotes: \(p^2=0.360\), \(2pq=0.480\), \(q^2=0.160\). Survivors, before normalisation: \(0.360,\ 0.480,\ 0.160(0.800)=0.128\). Mean fitness \(\bar w=1-sq^2=1-0.200(0.160)=0.968\), which is also the sum \(0.360+0.480+0.128\).
Adult frequencies: \(0.360/0.968=0.37190\); \(0.480/0.968=0.49587\); \(0.128/0.968=0.13223\). Sum \(=1.00000\).
Adult allele frequencies, in symbols first: \(p'=(p^2+pq)/\bar w=p/\bar w\) and \(q'=q(1-sq)/\bar w\). Numerically \(p'=0.600/0.968=0.61983\) and \(q'=0.400(1-0.080)/0.968=0.368/0.968=0.38017\); the change is \(\Delta q=-0.01983\) per generation.
(b) \(2p'q'=2(0.61983)(0.38017)=0.47128\), so \(F=1-0.49587/0.47128=1-1.05217=-0.05217\): a heterozygote excess, because the culling removed only homozygotes. The closed form is \(F=-spq/(1-sq)=-(0.200)(0.600)(0.400)/(1-0.080)=-0.048/0.920=-0.05217\), in agreement.
\(\chi^2=nF^2=1000(0.05217)^2=2.72\) on \(1\) degree of freedom, below the critical \(3.841\). The selection would not be detected: a twenty per cent viability cost, which is enormous by the standards of natural populations, is invisible to a Hardy–Weinberg test on \(1000\) adults. The sample size required is \(n\ge3.841/(0.05217)^2\approx1411\) individuals for even marginal significance.
(c) Surviving adults mate at random, so their gamete pool has allele frequency \(p'=0.61983\) and their offspring are in Hardy–Weinberg proportions at that frequency: \(0.38419,\ 0.47128,\ 0.14453\). The lesson is that selection acting between zygote and adult leaves every newborn cohort exactly on \(\mathcal{H}\) while the allele frequency marches steadily downwards. Hardy–Weinberg equilibrium is therefore not evidence that a population is not evolving: the equilibrium is restored each generation by random mating, and the evolution shows up in the drift of \(p\) between cohorts, not in the proportions within one. Detecting it requires either sampling the same cohort at two life stages or comparing allele frequencies across generations.