Testing Counts Against a Claim
Comparing observed counts with the counts a hypothesis predicts, how the degrees of freedom follow from the table shape and from any parameter estimated along the way, when the chi-square approximation can be trusted, and which procedure remains available when it cannot.
Definition
A goodness-of-fit test compares observed category counts
Under the null,
Degrees of freedom. For
A test of independence applies the same statistic to an
which is the count predicted if row and column classifications were unrelated. Here
The approximation condition. The chi-square distribution is a large-sample approximation to the distribution of
Rank-based alternatives replace the observations by their ranks, so no distributional shape is assumed. They test a different hypothesis from their parametric counterparts, typically about a shift in location rather than about a difference in means.
Assumptions and scope
The observations must be independent and each must fall in exactly one category. Repeated measurements on the same unit, or units that can appear in two cells, break the test rather than weaken it.
The counts entering
must be counts, not percentages or rates. Applying the statistic to proportions multiplies it by an arbitrary factor and the reference distribution no longer applies. The expected-count condition concerns the expected counts, not the observed ones. A cell observing zero is unremarkable if its expected count is large.
A test of independence establishes association only. It cannot identify a direction, and a third variable related to both classifications produces the same signal.
Failing to reject is not evidence that the hypothesised distribution is correct. A goodness-of-fit test with few observations fails to reject nearly everything, so a large
reports a lack of resolution as readily as a good fit. A rank-based test is not the same test with weaker assumptions. It answers a different question, usually about location shift, and a difference in means can exist where its null holds.
Worked material
Example
Five questions, and whether this test answers them
Do the digits of a measurement follow Benford's law? A fully specified null: the first-digit probabilities are
Do these counts follow a Poisson distribution? The mean has to come from somewhere. If it is estimated from the same counts, the degrees of freedom drop by one, and the test asks the narrower question of whether the data are compatible with some Poisson distribution rather than with a particular one. If the mean is given by theory, the question is sharper and the degrees of freedom stay.
Are treatment and outcome associated in this table? The test answers this. It does not answer whether treatment caused the outcome, and nothing in the table distinguishes association produced by a causal path from association produced by a third variable related to both. Where assignment was randomised, that separate design fact licenses the causal reading. The test contributes the same evidence either way.
Is the sample representative of the population? Partly. Comparing observed demographic counts with census proportions is a goodness-of-fit test with a specified null, so it can detect a discrepancy on the variables compared. Passing it says nothing about variables not tabulated, and a sample matching the census on age and region can still be unrepresentative on income.
Do two groups have the same mean response? Not this test. Chi-square on counts addresses categories; comparing means of a continuous response calls for a
---
What separates the first two from the rest. In the first two the null predicts counts, which is the situation the statistic is built for. In the third the arithmetic applies and the interpretation is the difficult part. In the fourth the test is sound and its scope is narrower than the question. In the fifth the data are the wrong type, and converting them to counts to fit the method is choosing the tool before the question.
Contrast
Pairs that differ in one decision
The chi-square approximation against Fisher's exact test, on one table.
| Fisher exact | ||
|---|---|---|
| significant at 5% | yes | yes |
| significant at 1% | yes | no |
| approximates | the null distribution | nothing |
Same table, same margins, expected counts of
Degrees of freedom with and without the fitted parameter.
| df | ||
|---|---|---|
| correct, one parameter estimated | 4 | |
| estimation ignored | 5 |
The error inflates the
A specified null against a fitted one.
Testing against Benford's law uses probabilities fixed by the hypothesis, and the full
Association against causation, in the same table.
A significant test of independence establishes that the joint counts depart from what the margins predict. Whether that licenses a causal reading depends on how the data arose, randomised assignment, or an observational comparison where a third variable could produce the identical pattern. The table cannot distinguish the two, and reporting the
A parametric test against its rank-based counterpart.
The second is not the first with weaker assumptions. It tests a different null, usually about a shift in location, so a difference in means can exist where the rank null holds and a rank test can reject where means agree. Switching because an assumption looks shaky changes the question being asked, which should be stated rather than treated as a repair.
Common errors
Common misconception
That the degrees of freedom for a goodness-of-fit test are always the number of categories minus one, whatever was done to obtain the expected counts. Each parameter estimated from the same data costs a further degree of freedom, because fitting pulls the expected counts toward the observed ones and manufactures part of the agreement being tested. For a Poisson model fitted to six categories of counts, the correct degrees of freedom are
Common misconception
That a large
Related units
Requires
- Conditional Probability, Total Probability and Bayes' Rule
- Hypothesis Tests for Experimental Research
Connected
- ANOVA for Experimental Research (contrasts with)