You have counts in four buckets and a theory about how they should have split. Twenty-five each, you said, or a 50/50 coin, or a 3:1 ratio from a genetics cross. Then you counted, and the numbers came out uneven, and now you have to decide whether that means anything or whether you just got unlucky.
That decision is a chi-square test, and this is the arithmetic for it. Observed values in one box, expected in the other, comma-separated, and it returns the statistic, the degrees of freedom, a p-value and a plain call at the 0.05 level.
One honest detail: the statistic is the easy part, pure multiplication. What makes chi-square confusing is deciding what your expected values should have been.
Calculate the chi-square statistic from observed and expected values. Enter comma-separated values for each.
The formula is shorter than it looks
For each pair of numbers you take the gap, square it, and divide by what you expected:
chi-square = sum of (observed − expected)² ÷ expected
Squaring kills the sign, so a count 10 too high and one 10 too low both push the statistic up instead of cancelling. Dividing by the expected value makes the gaps comparable: being 10 short in a bucket of 25 is different evidence from being 10 short in a bucket of 1,000.
Worked through, with the example from the placeholder, 10, 20, 30, 40 against 25, 25, 25, 25:
(10−25)²÷25 = 9, (20−25)²÷25 = 1, (30−25)²÷25 = 1, (40−25)²÷25 = 9. Total: chi-square = 20.0.
Degrees of freedom is the number of categories minus one, so four values gives df = 3. The minus one is because the categories have to add up to your total: fix three of them and the fourth has no freedom left to move.
Reading the answer
Under the statistic you get a p-value and a sentence saying whether the result is significant at the 0.05 level. What that means in plain terms: if your expected split were exactly right, how often would random counting produce a spread this bad? Below 5% the convention is to call the pattern real. Above it, you have nothing to report but noise.
The standard cross-check is a table of critical values — the point where the statistic crosses into “significant” for each degrees of freedom:
| df | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Critical value (0.05) | 3.84 | 5.99 | 7.81 | 9.49 | 11.07 | 12.59 |
Our 20.0 on 3 degrees of freedom is well past 7.81, so the answer is significant, and it is not close. Now a smaller one: 46 and 54 observed against 50 and 50 expected gives 0.64 on 1 df, comfortably under 3.84, and the honest conclusion is that you observed nothing.
Keep the statistic and the degrees of freedom as the parts you trust: arithmetic you can redo on paper. For the significance call, use the critical values above — published, fixed, and independent of how any code approximates a distribution.
What breaks the test
Percentages instead of counts. The statistic scales with your numbers, so 10%, 30%, 60% against 33.3% each gives a different value than 5, 15, 30 against 16.7 each, even though the shape of the data is identical. Enter the raw counts.
Bad expected values. The tool requires every expected number to be above zero and refuses anything else, but it cannot tell whether your expectations were sensible. Cells under about 5 make the approximation shaky, so a test on 3, 1, 2 versus 4, 1, 1 is arithmetic without much meaning. Group the small categories or collect more data.
Independence. Each observation must be counted once and belong to exactly one category. The same person in two cells, or repeated measures from one subject treated as separate ones, inflates the statistic and hands you a confident answer to a question you did not ask.
And a significant result is narrower than it sounds. It says the split is unlikely to be pure chance. It does not say the difference is large, which cell caused it, or why. With a big enough sample almost everything comes back significant, so look at the size of the gaps next to the p-value rather than reading “significant” as “important”.
Other tools
- class conflict checker — for the timetable data you would rather test before registration closes
- course retake GPA — what a second attempt does to the average you were trying to protect
- course retake impact — the longer view of the same decision
Frequently asked questions
How many values can I enter?
As many categories as you like, as long as the two lists are the same length, which the tool checks. Each extra category adds one to the degrees of freedom, raising the bar the statistic has to clear.
What should my expected values be?
Whatever your hypothesis predicts before you looked at the data: a uniform split, a known population ratio, a 3:1 genetic ratio. Computing expected values from the observed data itself is the classic mistake, because it pulls the statistic towards zero and guarantees a non-significant answer.
Is 0.05 a magic number?
It is a convention, not a law. Some fields use 0.01, some report the exact p and let the reader decide. What matters is choosing your threshold before seeing the results, and reporting the statistic and degrees of freedom alongside it.
What does degrees of freedom mean here?
The number of categories minus one. Because the categories must sum to your total, knowing all but one of them fixes the last automatically, so only n−1 were ever free to vary.
Why is my p-value different from the textbook?
Check the counts first, then the degrees of freedom. Percentages instead of raw counts, or categories set up so a parameter was estimated from the data, both change the answer. Compare the statistic against the table above before assuming anything else is wrong.
