Free tools. No account. Nothing stored.

Population Genetics Calculator

Every genetics course teaches Hardy-Weinberg the same way. Here is p, here is q, run the equation, here is your answer. It is a fine demonstration and it is not what you actually do in a lab.

In a lab you count the flies. Or the peas, or the people, or whatever organism is in the bottle. You come out with three numbers — how many were homozygous dominant, how many heterozygous, how many homozygous recessive — and from those you have to work out the allele frequencies yourself, and then decide whether the population is doing what the equation claims it should. This calculator does both directions, and the backward one comes with a chi-square test, because that is the part that tells you whether the textbook conditions are actually holding in the population in front of you.

Two ways in. If you know the allele frequencies, this works out the genotype split they predict. If you counted the organisms instead — which is what a genetics lab actually makes you do — it works backward to the allele frequencies and then tests whether that population is really in equilibrium. The test is a chi-square goodness of fit, and it is the part that tells you whether the textbook conditions actually hold.

What the equation actually says

One allele, call it A. Its alternative version in the population is a. The frequency of A across the whole gene pool is p, and the frequency of a is q. Because those are the only two options, p + q = 1 exactly. There is no third allele hiding somewhere, at least not in this model.

An organism gets two alleles, one from each parent. So the possible genotypes are:

  • AA — both alleles are A, so p × p = p2
  • Aa — one of each, and there are two orderings (A then a, or a then A), so 2pq
  • aa — two a alleles, so q × q = q2

And those three have to add up to 1: p2 + 2pq + q2 = 1. That identity is worth noticing, because it means the equation always works out cleanly with no remainder. Try p = 0.6, q = 0.4:

  • AA = 0.6 × 0.6 = 0.36, so 36%
  • Aa = 2 × 0.6 × 0.4 = 0.48, so 48%
  • aa = 0.4 × 0.4 = 0.16, so 16%

36 + 48 + 16 = 100. Out of a thousand offspring you would expect 360, 480 and 160.

Working backward from your counts

This is the direction the lab actually uses, and the reason is simple. You cannot walk up to a fruit fly and ask for its allele frequencies. Allele frequency is a property of the entire population, so you estimate it from a sample by counting.

Say you counted 100 organisms and found 36 homozygous dominant, 48 heterozygous and 16 homozygous recessive. Every AA organism carries two A alleles. Every Aa carries one. Every aa carries none. So:

p = (2 × 36 + 48) ÷ 200 = 120 ÷ 200 = 0.6

The 200 is 2 × 100 — two alleles per organism across the whole sample. Then q = 1 − p = 0.4, and the expected split comes out at 36 / 48 / 16, which matches your count exactly. That is a population sitting in equilibrium.

The chi-square test, which is the real point

Matching perfectly is unusual. So the question is not “is it exactly right” but “is the difference bigger than chance alone would explain.”

Chi-square adds up (observed − expected) ÷ expected, squared, across all three classes. Then you compare it to a critical value. For one degree of freedom at the 0.05 level, that value is 3.841. One degree of freedom, not two, because the three genotype classes are not independent — fix p and q is fixed automatically.

Chi-square you calculate What it means
Below 3.841 Not significantly different from expected. Consistent with equilibrium.
3.841 or above Significantly different. At least one equilibrium condition is not holding.

Being under the threshold does not prove you are in equilibrium. It means your data does not contradict it. With 100 organisms you have very little power to detect a small deviation, and a textbook sample set is usually tidy enough to pass comfortably.

Zero heterozygotes is a real result, not a mistake

If you count 60 AA, 0 Aa and 40 aa, the allele frequencies come out as clean as you like: p = 0.6, q = 0.4. But that predicts 36 / 48 / 16, and you found none of the heterozygotes at all. The chi-square comes out at 100, which is nowhere near the threshold.

Do not re-check your counts first. A complete absence of heterozygotes is the classic signature of inbreeding, or of mating only between relatives. Random mating is not happening in this population, which is one of the five conditions failing in a very visible way. It is a finding, not an error.

The five conditions, and what breaking each one does

Hardy-Weinberg only holds if all five of these are true at once. Tick them off in the calculator when you have a population in front of you.

Condition What happens without it
Very large population Genetic drift. Chance alone starts shifting allele frequencies, which matters most in small populations.
Random mating Inbreeding. Heterozygotes drop, which is what the example above shows.
No migration Gene flow. Alleles move in or out and the frequencies no longer describe the original population.
No mutations One allele converts into another, so p and q are not constant across generations.
No natural selection One allele gives a survival or reproductive advantage, so the frequencies drift toward it generation after generation.

A chi-square failure cannot tell you which condition broke. It only tells you that one of them did. The count pattern is what narrows it down — and a missing heterozygote class points straight at random mating.

Two limits worth knowing

Dominance is a labelling convention. You cannot see an allele. A dominant phenotype is AA or Aa, so counting phenotypes does not give you genotypes. Working out which of the two you are looking at needs a test cross or a chi-square of its own. If your assignment gives you phenotype counts, check whether it wants a dominance question answered first.

Two alleles only. p + q = 1 assumes there are exactly two versions of the gene. Real traits often have more, and then the model needs a letter for each. Some textbooks quietly ignore the extra versions, which is where a lot of confusing marks come from.

Frequently asked questions

What is the Hardy-Weinberg equation?

p2 + 2pq + q2 = 1. It gives the expected genotype frequencies in a population from its allele frequencies, provided the five conditions hold.

How do I find p and q from a count?

p = (2 × AA + Aa) ÷ 2N, and q = (2 × aa + Aa) ÷ 2N. Every heterozygote contributes one of each allele, which is why it appears in both.

What chi-square value means the population is in equilibrium?

At 0.05 significance with one degree of freedom, anything below 3.841. Above that, the deviation is too large to blame on chance alone.

Why is it one degree of freedom and not two?

There are three genotype classes, but they are not independent. Because p + q is fixed at 1, knowing one class determines the others. That constraint costs one degree of freedom, leaving 1.

What if my counts are nowhere near the expected ratio?

That is a result, not a failure. Work out p and q from your own counts, compare against the expected split, and look at which class is off. Heterozygotes low usually means inbreeding. One class dominating usually means selection is acting.

Can a population be in equilibrium without being stable?

Yes. Equilibrium only means the genotype frequencies match what the allele frequencies predict right now. Natural selection can still be acting on the alleles. The equation is a description, not a mechanism.

What does a Hardy-Weinberg calculator actually do?

Two directions, and you need the second one far more often. Given p and q, it returns the expected genotype split. Given your three observed counts, it works p and q back out, predicts what the split should have been, and runs the chi-square so you can tell whether the difference is just sampling noise. The backward direction is the one that matches a lab practical, because allele frequency is a property of the whole population and the only way to get it is to count a sample.

How many organisms do I need to count?

Enough that a real deviation stands out from the noise, and there is no clean answer because it depends on how big the deviation is. The general shape: a small sample will happily pass a population that is clearly out of equilibrium, and a large one will flag tiny differences that do not matter in practice. Textbook exercises usually use 100 because the numbers divide nicely, not because 100 is the correct threshold. If your assignment tells you a sample size, use it. If it does not, count what you have and treat a pass as “not contradicted” rather than “proved”.

Related tools

A Punnett square belongs next to this one, and the two answer opposite questions. The square asks what will this cross produce. This page asks does this population match what its own allele frequencies predict. One works forward from parents, the other backward from counts you already have.