Testing Counts Against a Claim
What you will be able to do
The learner can carry out a chi-square goodness-of-fit test and a test of independence from a contingency table, determine the degrees of freedom from the table shape and the number of estimated parameters, check the expected-count condition, and select a nonparametric alternative when the assumptions of a parametric test fail.
Orientation
Counts against a prediction
A hypothesis about categories makes a prediction that can be written down: if the die is fair, each face should appear about a sixth of the time; if two classifications are unrelated, the joint counts should follow from the margins alone. The data arrive as counts, and the question is whether they sit further from the prediction than chance would explain.
One statistic serves both cases. It adds up the discrepancies between observed and expected counts, each divided by the expected count so that cells of different size can be compared, and refers the total to a chi-square distribution.
Computing it is straightforward. The two things most often got wrong sit either side of the arithmetic.
The degrees of freedom depend on what was done to obtain the expected counts. Fitting a distribution's parameter to the same data costs a degree of freedom, and omitting that correction makes the model appear to fit better than the evidence supports.
The reference distribution is an approximation with a condition. It is accurate when the expected counts are large enough, and the threshold usually quoted is a rule of thumb rather than a guarantee. This unit works a case sitting exactly on that threshold where the approximation and an exact test disagree by a factor of three.
Definition
Where each quantity in the statistic comes from
The canonical statements above give the statistic and the degrees of freedom. What follows is where each piece originates, since that is what decides the right value in an unfamiliar case.
The expected counts come from the null, and nowhere else. For a fully specified hypothesis they are
Degrees of freedom count unconstrained discrepancies.
| Situation | Constraints | df |
|---|---|---|
| counts sum to | ||
| the sum, plus each fitted parameter | ||
| both sets of margins |
The table row follows from the same counting: fixing all row and column totals leaves
Why dividing by
Fisher's exact test computes rather than approximates. For a
Rank-based procedures change the hypothesis. Replacing observations with ranks removes any assumption about distributional shape and also changes what is being tested, typically to a statement about location shift rather than about means. A difference in means can exist while the rank-based null holds, and a rank test can reject where means are equal, so the two are not interchangeable.
Intuition
Why fitting a parameter uses up evidence
Suppose a model is proposed and its parameter is chosen to make the model match the data as closely as possible. Then the match is asked to serve as evidence that the model is right. Part of the agreement was arranged, and the test has to discount it.
That is what subtracting a degree of freedom does. Degrees of freedom count how many independent ways the observed counts could have departed from the expected ones. Every constraint removes one: the counts must total
The direction of the error is worth knowing, because it is the reverse of what most people expect. Using too many degrees of freedom compares the statistic against a distribution that is too spread out, so the tail probability comes out larger and the fit looks better. With a Poisson model fitted to six categories, the correct
Why the discrepancies are divided by the expected count. A cell expecting 10 and observing 15 is behaving unusually; a cell expecting 200 and observing 205 is not, though both are 5 away. Under the null a count has variance roughly equal to its expectation, so dividing by
What independence means in a table. Build the table that the row and column totals alone would predict, assuming each classification says nothing about the other. Compare it with the observed table. A large statistic means the joint pattern carries information the margins do not, which is association, and only that. The test cannot say which variable moved, and a third variable influencing both produces exactly the same signal.
Why the approximation depends on the cells rather than on
Example
Five questions, and whether this test answers them
Do the digits of a measurement follow Benford's law? A fully specified null: the first-digit probabilities are
Do these counts follow a Poisson distribution? The mean has to come from somewhere. If it is estimated from the same counts, the degrees of freedom drop by one, and the test asks the narrower question of whether the data are compatible with some Poisson distribution rather than with a particular one. If the mean is given by theory, the question is sharper and the degrees of freedom stay.
Are treatment and outcome associated in this table? The test answers this. It does not answer whether treatment caused the outcome, and nothing in the table distinguishes association produced by a causal path from association produced by a third variable related to both. Where assignment was randomised, that separate design fact licenses the causal reading. The test contributes the same evidence either way.
Is the sample representative of the population? Partly. Comparing observed demographic counts with census proportions is a goodness-of-fit test with a specified null, so it can detect a discrepancy on the variables compared. Passing it says nothing about variables not tabulated, and a sample matching the census on age and region can still be unrepresentative on income.
Do two groups have the same mean response? Not this test. Chi-square on counts addresses categories; comparing means of a continuous response calls for a
---
What separates the first two from the rest. In the first two the null predicts counts, which is the situation the statistic is built for. In the third the arithmetic applies and the interpretation is the difficult part. In the fourth the test is sound and its scope is narrower than the question. In the fifth the data are the wrong type, and converting them to counts to fit the method is choosing the tool before the question.
Procedure
Running the test, and deciding whether to trust it
To test goodness of fit against a specified distribution.
- State the null as probabilities
summing to 1, fixed before looking at the data. - Compute expected counts
. - Check every
. The working rule asks for at least 5 in each cell. If some fall short, combine adjacent categories. A decision made on substantive grounds, and recorded, since combining until the test passes is selecting on the outcome. - Compute
. - Degrees of freedom:
, minus one for each parameter estimated from this data. - Refer to the chi-square distribution and report the statistic, the degrees of freedom and the
-value together. The statistic alone is uninterpretable without the degrees of freedom.
To test independence in an
- Compute the margins
, and the total . - Expected counts
. - Check them as above. For a
table with marginal counts, prefer Fisher's exact test outright rather than checking a threshold. - Compute
over all cells. - Degrees of freedom
. - If the result is significant, look at the cells. The standardised residual
shows where the departure sits; a significant statistic says only that the table as a whole departs from independence.
To choose between a parametric test and a rank-based one.
- Name the assumption in doubt, normality, equal variances, or the influence of extreme values.
- Ask whether it matters at this sample size. A two-sample
test is robust to moderate non-normality at reasonable ; it is not robust to a heavy tail with few observations. - If switching, state what the alternative tests. A rank test addresses a shift in location, not a difference in means, and the two coincide only under extra assumptions about shape.
- Decide before seeing the result. Running both and reporting the smaller
-value invalidates both.
Checks. Confirm the expected counts sum to
Worked example
Four tests, and two places the answer turns
(a) Is the die fair? 300 rolls of a six-sided die:
| Face | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Observed | 43 | 52 | 38 | 61 | 47 | 59 |
| Expected | 50 | 50 | 50 | 50 | 50 | 50 |
The null is fully specified, each
With
(b) Are the two classifications independent? A
| A | B | C | Total | |
|---|---|---|---|---|
| Row 1 | 30 | 45 | 25 | 100 |
| Row 2 | 20 | 25 | 55 | 100 |
| Total | 50 | 70 | 80 | 200 |
Expected counts from the margins,
| A | B | C | |
|---|---|---|---|
| Row 1 | |||
| Row 2 |
The rows have equal totals, so each column's count splits evenly under independence. Then
The smallest expected count is
(c) A case sitting exactly on the condition. A
| Total | |||
|---|---|---|---|
| Row 1 | 2 | 8 | 10 |
| Row 2 | 9 | 3 | 12 |
| Total | 11 | 11 | 22 |
Expected counts are
Fisher's exact test, which sums hypergeometric probabilities and approximates nothing, gives
The two differ by a factor of about three, and they fall either side of
(d) The degrees of freedom, when a parameter was fitted. Frequencies of an event count over 90 observations:
| Count | 0 | 1 | 2 | 3 | 4 | 5+ |
|---|---|---|---|---|---|---|
| Observed | 18 | 26 | 20 | 14 | 8 | 4 |
A Poisson model is proposed, its mean estimated from this same data as
| Count | 0 | 1 | 2 | 3 | 4 | 5+ |
|---|---|---|---|---|---|---|
| Expected |
Six categories, one constraint from the total, one parameter estimated:
Had the estimation been ignored,
What (a) and (d) do not establish. Neither large
Contrast
Pairs that differ in one decision
The chi-square approximation against Fisher's exact test, on one table.
| Fisher exact | ||
|---|---|---|
| significant at 5% | yes | yes |
| significant at 1% | yes | no |
| approximates | the null distribution | nothing |
Same table, same margins, expected counts of
Degrees of freedom with and without the fitted parameter.
| df | ||
|---|---|---|
| correct, one parameter estimated | 4 | |
| estimation ignored | 5 |
The error inflates the
A specified null against a fitted one.
Testing against Benford's law uses probabilities fixed by the hypothesis, and the full
Association against causation, in the same table.
A significant test of independence establishes that the joint counts depart from what the margins predict. Whether that licenses a causal reading depends on how the data arose, randomised assignment, or an observational comparison where a third variable could produce the identical pattern. The table cannot distinguish the two, and reporting the
A parametric test against its rank-based counterpart.
The second is not the first with weaker assumptions. It tests a different null, usually about a shift in location, so a difference in means can exist where the rank null holds and a rank test can reject where means agree. Switching because an assumption looks shaky changes the question being asked, which should be stated rather than treated as a repair.
Warning
The errors that read as good news
A large
The distinction has a practical test. Genuine agreement survives more data; agreement arising from a lack of resolution does not. A goodness-of-fit test on 20 observations fails to reject almost every candidate, so quoting its
Forgetting that a parameter was estimated inflates the
The threshold of 5 is a convention, not a guarantee. In the worked
Combining categories until the condition is met is a decision about the data. Merging cells to raise expected counts is legitimate when the merged categories mean something together, and it is selection on the outcome when the merge is chosen because it produces a publishable result. Decide the grouping before computing, and record it.
---
Two errors of type rather than of arithmetic.
Percentages are not counts.
Each observation must occupy exactly one cell. Repeated measurements on the same subject, or units that can be counted twice, inflate the apparent sample size and the statistic with it. This breaks the test rather than weakening it, and no correction to the degrees of freedom repairs it.
---
A significant table says the table departs from independence, and no more. It does not identify which cells carry the departure, standardised residuals do that, nor which variable moved, nor whether a third variable produced the pattern. Reporting only the
Application
Where counts carry the argument
Audit and fraud screening. First-digit distributions of reported figures are compared against Benford's law, a fully specified null needing no fitted parameter. A significant departure is a reason to look, not a finding of fraud: legitimate processes with bounded ranges or assigned numbering also violate the law, so the test selects where to spend attention rather than concluding anything.
Genetics. Observed offspring counts across phenotypic classes are compared with the ratios a proposed inheritance model predicts. Because the ratios come from the model rather than the data, the full degrees of freedom are available and the test is sharp. Where a recombination parameter has to be estimated from the same cross, the correction applies and the question narrows to compatibility with a family of models.
Survey weighting. Sampled demographic counts are compared with population margins to decide whether weighting is needed. The test is sound and its scope is exactly the variables tabulated; a sample matching on age and region can be unrepresentative on anything not compared, so passing is not a certificate of representativeness.
Clinical trial baselines with small strata. Comparing adverse-event counts across arms produces tables where some expected counts are small, which is where the exact test earns its place. Regulatory analyses specify the procedure in advance precisely because choosing between the approximation and the exact test after seeing which gives the smaller
A/B testing on conversion counts. A
---
The recurring shape. In each case the test contributes one thing, whether the counts depart from what a stated hypothesis predicts, and the interpretation is supplied by how the data arose. Randomisation licenses a causal reading; a prespecified analysis licenses the