Testing Counts Against a Claim

Comparing observed counts with the counts a hypothesis predicts, how the degrees of freedom follow from the table shape and from any parameter estimated along the way, when the chi-square approximation can be trusted, and which procedure remains available when it cannot.

Definition

A goodness-of-fit test compares observed category counts O 1 , … , O k against the counts E 1 , … , E k a hypothesised distribution predicts, through

X 2 = ∑ i = 1 k ( O i − E i ) 2 E i .

Under the null, X 2 has approximately a chi-square distribution. Dividing by E i is what makes the terms comparable: a discrepancy of 5 matters more where 10 were expected than where 200 were.

Degrees of freedom. For k categories with a fully specified null, df = k − 1 , the single constraint being that the counts sum to n . Each parameter estimated from the same data costs one more, so fitting a distribution with p estimated parameters gives df = k − 1 − p .

A test of independence applies the same statistic to an r × c contingency table, with expected counts built from the margins,

E i j = ( row  i  total ) ( column  j  total ) N ,

which is the count predicted if row and column classifications were unrelated. Here df = ( r − 1 ) ( c − 1 ) .

The approximation condition. The chi-square distribution is a large-sample approximation to the distribution of X 2 , and its accuracy depends on the expected counts rather than on n . The usual working rule asks that every expected count be at least 5, with Fisher's exact test available for a 2 × 2 table, which computes the null probability directly from the hypergeometric distribution and needs no approximation at all.

Rank-based alternatives replace the observations by their ranks, so no distributional shape is assumed. They test a different hypothesis from their parametric counterparts, typically about a shift in location rather than about a difference in means.

Assumptions and scope

  • The observations must be independent and each must fall in exactly one category. Repeated measurements on the same unit, or units that can appear in two cells, break the test rather than weaken it.

  • The counts entering X 2 must be counts, not percentages or rates. Applying the statistic to proportions multiplies it by an arbitrary factor and the reference distribution no longer applies.

  • The expected-count condition concerns the expected counts, not the observed ones. A cell observing zero is unremarkable if its expected count is large.

  • A test of independence establishes association only. It cannot identify a direction, and a third variable related to both classifications produces the same signal.

  • Failing to reject is not evidence that the hypothesised distribution is correct. A goodness-of-fit test with few observations fails to reject nearly everything, so a large p reports a lack of resolution as readily as a good fit.

  • A rank-based test is not the same test with weaker assumptions. It answers a different question, usually about location shift, and a difference in means can exist where its null holds.

Worked material

Example

Five questions, and whether this test answers them

Do the digits of a measurement follow Benford's law? A fully specified null: the first-digit probabilities are log 10 ⁡ ( 1 + 1 / d ) for d = 1 , … , 9 , fixed in advance by the hypothesis rather than fitted. Eight degrees of freedom, and the test is exactly what it was designed for.

Do these counts follow a Poisson distribution? The mean has to come from somewhere. If it is estimated from the same counts, the degrees of freedom drop by one, and the test asks the narrower question of whether the data are compatible with some Poisson distribution rather than with a particular one. If the mean is given by theory, the question is sharper and the degrees of freedom stay.

Are treatment and outcome associated in this table? The test answers this. It does not answer whether treatment caused the outcome, and nothing in the table distinguishes association produced by a causal path from association produced by a third variable related to both. Where assignment was randomised, that separate design fact licenses the causal reading. The test contributes the same evidence either way.

Is the sample representative of the population? Partly. Comparing observed demographic counts with census proportions is a goodness-of-fit test with a specified null, so it can detect a discrepancy on the variables compared. Passing it says nothing about variables not tabulated, and a sample matching the census on age and region can still be unrepresentative on income.

Do two groups have the same mean response? Not this test. Chi-square on counts addresses categories; comparing means of a continuous response calls for a t test or ANOVA, and forcing the response into bins to obtain counts discards information and makes the answer depend on the bin boundaries chosen.

---

What separates the first two from the rest. In the first two the null predicts counts, which is the situation the statistic is built for. In the third the arithmetic applies and the interpretation is the difficult part. In the fourth the test is sound and its scope is narrower than the question. In the fifth the data are the wrong type, and converting them to counts to fit the method is choosing the tool before the question.

Contrast

Pairs that differ in one decision

The chi-square approximation against Fisher's exact test, on one table.

χ 2 Fisher exact
p -value 0.0102 0.029973
significant at 5%yesyes
significant at 1%yesno
approximatesthe null distributionnothing

Same table, same margins, expected counts of 5.0 , 5.0 , 6.0 , 6.0 satisfying the conventional rule. The two procedures disagree by a factor of about three and straddle a 1% threshold. Neither was misapplied; the first is an approximation and the second is not.

Degrees of freedom with and without the fitted parameter.

df p
correct, one parameter estimated4 0.7517
estimation ignored5 0.8610

The error inflates the p -value, so the model appears to fit better. An error that made the fit look worse would be caught by the author defending the model; this one reads as confirmation.

A specified null against a fitted one.

Testing against Benford's law uses probabilities fixed by the hypothesis, and the full k − 1 degrees of freedom are available. Testing against "some Poisson distribution" spends one on the estimated mean, and the question answered narrows correspondingly: compatibility with a family, rather than with a stated distribution. The statistic is computed identically and the claims differ.

Association against causation, in the same table.

A significant test of independence establishes that the joint counts depart from what the margins predict. Whether that licenses a causal reading depends on how the data arose, randomised assignment, or an observational comparison where a third variable could produce the identical pattern. The table cannot distinguish the two, and reporting the p -value without saying which situation obtains leaves the reader to supply the stronger reading.

A parametric test against its rank-based counterpart.

The second is not the first with weaker assumptions. It tests a different null, usually about a shift in location, so a difference in means can exist where the rank null holds and a rank test can reject where means agree. Switching because an assumption looks shaky changes the question being asked, which should be stated rather than treated as a repair.

Common errors

Common misconception

That the degrees of freedom for a goodness-of-fit test are always the number of categories minus one, whatever was done to obtain the expected counts. Each parameter estimated from the same data costs a further degree of freedom, because fitting pulls the expected counts toward the observed ones and manufactures part of the agreement being tested. For a Poisson model fitted to six categories of counts, the correct degrees of freedom are 6 − 1 − 1 = 4 , giving p = 0.7517 , while using 6 − 1 = 5 gives p = 0.8610 . The error makes the fit look better than the evidence supports, which is the opposite of the direction most people expect and is why it is rarely caught by inspection.

Common misconception

That a large p -value from a goodness-of-fit test establishes that the data follow the hypothesised distribution. The test can only find evidence against the null, so failing to reject reports that the observed counts are compatible with the model, not that the model is right. Several distributions are typically compatible with the same counts, and a test on few observations fails to reject nearly everything, so a large p is produced both by a good fit and by a study with no resolution. The distinction shows in what changes with more data: genuine agreement persists, while agreement arising from a lack of power disappears as soon as the sample is large enough to detect the discrepancy.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.