Counting how many organisms are AA, Aa and aa is the easy half of any population genetics question. The half that carries the marks is the one after it: are those counts consistent with Hardy–Weinberg equilibrium? Answering that properly needs a chi-square goodness-of-fit test with a p-value, and a p-value is not something you get from the critical value table in the back of the book, because that table only covers p = 0.05.
Counting genotypes is the easy half. The question that actually gets marks is whether the counts are consistent with Hardy–Weinberg, and that needs a proper chi-square test with a p-value, not a table lookup. This one works at any significance level and tells you which assumption to suspect when it fails.
Enter raw counts, not percentages. Any labelling works for the two homozygotes — which one is “dominant” only sets the letter, it does not change the maths. Leave alpha at 0.05 unless your question states otherwise.
What the test actually does
Hardy–Weinberg predicts the genotype ratios you should see given the allele frequencies in the population. You measure the allele frequencies from your own counts, predict what the genotype counts ought to be, and then ask how far reality is from that prediction.
The two steps people get wrong:
- Calculate p from the data, not from the ratio. p counts the A alleles, so each AA contributes 2 and each Aa contributes 1. p = (2×AA + Aa) / 2N, and q = 1 − p.
- Expected counts use p², 2pq and q². Not p, q and 1−p. Under random mating you get two homozygotes and a lot of heterozygotes, which is the whole point of the equilibrium.
Then chi-square adds up (observed − expected)² / expected across all three classes. A statistic of zero means a perfect fit. A large statistic means the counts and the prediction disagree.
Why one degree of freedom, not two
This is the most common slip in the topic, and it matters because it doubles your p-value.
You have three genotype classes, so a naive count says three degrees of freedom. But one thing is already fixed: the three counts have to sum to N. And you have not been given p — you estimated it from the data, which costs another one. So 3 − 1 − 1 = 1.
That is where the famous 3.841 comes from. If you use 2 degrees of freedom you compare against 5.991 instead, a much more forgiving threshold, and a genuinely non-equilibrium sample will pass your test when it should have failed.
Reading the result honestly
Compare the statistic against the critical value at your chosen alpha, or read the p-value directly:
- p greater than 0.05 — fail to reject. The data are consistent with equilibrium.
- p less than 0.05 — reject. Something in the assumptions is not holding.
Two things worth saying out loud in an exam. First, failing to reject is not proving equilibrium. It means these counts are not far enough off to rule it out, and a small sample sits inside the noise whether or not it is truly in equilibrium. Second, you only get to reject after you have specified alpha before looking at the data, otherwise you can always find a level that flatters your answer.
Which assumption failed?
When the test rejects, one of five conditions is not being met. In rough order of how often each is the answer:
- Random mating is not happening. Inbreeding, assortative mating, or a small isolated population with limited mating partners.
- Selection against one allele. The homozygote you are undercounting is the one that leaves fewer offspring.
- The population is small. Genetic drift moves allele frequencies a long way by chance alone, and a small population is exactly where you notice.
- Migration. Gene flow from a population with different allele frequencies, or the Wahlund effect when you pool two populations that differ.
- Mutation. Real, but almost never the right answer, because the rate is far too slow to detect in one sample.
Worked example
You count 100 AA, 200 Aa and 700 aa, so N = 1000.
p = (2×100 + 200) / 2000 = 0.2, and q = 0.8. Expected counts are p²N = 40, 2pqN = 320 and q²N = 640.
Chi-square = (100−40)²/40 + (200−320)²/320 + (700−640)²/640 = 90 + 45 + 5.625 = 140.625.
With one degree of freedom the critical value at 0.05 is 3.841, and 140.625 is nowhere near it. The p-value is effectively zero. Reject equilibrium decisively — and the direction tells you more than the test does. There are 200 heterozygotes where 320 were expected, a heterozygote deficit, which points at inbreeding or selection rather than anything else.
You can read the same result as an inbreeding coefficient. Observed heterozygosity is 200/1000 = 0.2, expected is 2pq = 0.32, so Fis = 1 − 0.2/0.32 = 0.375. A third of the expected heterozygosity has been lost. Genetics courses report this as Fis because it says how much inbreeding the population has actually experienced, which a bare p-value does not.
When the test does not apply
If an expected count falls below 5, that cell breaks the approximation the test relies on. Standard practice is to pool the small classes together, which is why the degrees of freedom can rise above 1, and this calculator reports that it has done so rather than hiding it.
If enough classes pool out that no degrees of freedom remain, the test is meaningless and no verdict is given. That happens with tiny samples and with a rare recessive allele, and the honest response is to report a larger sample rather than to quote a number. This calculator shows you the expected counts and tells you the test cannot be run, instead of inventing a verdict.
Every figure here comes from the standard single-locus Hardy–Weinberg model. Real populations are usually a mixture of many loci with linkage between them, and the single-locus test is a first approximation at best.
Frequently Asked Questions
What degrees of freedom does a Hardy-Weinberg chi-square test use?
One, when all three genotype classes are populated. There are three classes, minus one because the counts must sum to N, minus one more because the allele frequency was estimated from the data. Using two is the most common mistake in the topic and it doubles your p-value.
Why does it matter that p is estimated from the data?
Because estimating a parameter from the same data you are testing costs a degree of freedom. The model predicts counts from p, and p came out of your counts, so p is not independent information.
What does a p-value of 0.5 mean here?
The observed counts are perfectly consistent with equilibrium. The test is measuring how far the data are from the prediction, and half an expected deviation is well within normal sampling noise.
What does failing to reject actually prove?
Almost nothing on its own. It means these counts are not far enough from equilibrium to rule it out. A small sample sits inside the noise whether or not it is truly in equilibrium.
Which Hardy-Weinberg assumption is most often the reason for rejection?
Usually random mating not happening, or selection against one allele. A heterozygote deficit points at inbreeding; a surplus points at mixing two genetically different populations, which is the Wahlund effect.
Why are expected counts below 5 a problem?
The chi-square approximation breaks down when a cell has almost no expected individuals, because the sampling distribution stops looking like chi-square. Standard practice is to pool those classes, which raises the degrees of freedom above 1.
What is the inbreeding coefficient Fis?
Fis is 1 minus observed heterozygosity divided by expected heterozygosity. It measures how much inbreeding a population has actually experienced, which tells you more than a bare p-value does.
