Module 2 of 2 · Lesson 1 of 1

Testing Counts Against a Claim

The two things most often got wrong sit either side of the arithmetic.

What you will be able to do

The learner can carry out a chi-square goodness-of-fit test and a test of independence from a contingency table, determine the degrees of freedom from the table shape and the number of estimated parameters, check the expected-count condition, and select a nonparametric alternative when the assumptions of a parametric test fail.

Orientation

Counts against a prediction

A hypothesis about categories makes a prediction that can be written down: if the die is fair, each face should appear about a sixth of the time; if two classifications are unrelated, the joint counts should follow from the margins alone. The data arrive as counts, and the question is whether they sit further from the prediction than chance would explain.

One statistic serves both cases. It adds up the discrepancies between observed and expected counts, each divided by the expected count so that cells of different size can be compared, and refers the total to a chi-square distribution.

Computing it is straightforward. The two things most often got wrong sit either side of the arithmetic.

The degrees of freedom depend on what was done to obtain the expected counts. Fitting a distribution's parameter to the same data costs a degree of freedom, and omitting that correction makes the model appear to fit better than the evidence supports.

The reference distribution is an approximation with a condition. It is accurate when the expected counts are large enough, and the threshold usually quoted is a rule of thumb rather than a guarantee. This unit works a case sitting exactly on that threshold where the approximation and an exact test disagree by a factor of three.

Definition

Where each quantity in the statistic comes from

The canonical statements above give the statistic and the degrees of freedom. What follows is where each piece originates, since that is what decides the right value in an unfamiliar case.

The expected counts come from the null, and nowhere else. For a fully specified hypothesis they are E i = n p i with the p i stated in advance. For a contingency table they are E i j = R i C j / N , which is N times the product of the marginal proportions. The count implied by the margins if the classifications were unrelated. Both are derived quantities, not summaries of the data, and the second uses the data only through its margins.

Degrees of freedom count unconstrained discrepancies.

SituationConstraintsdf
k categories, null fully specifiedcounts sum to n k − 1
k categories, p parameters fittedthe sum, plus each fitted parameter k − 1 − p
r × c table of independenceboth sets of margins ( r − 1 ) ( c − 1 )

The table row follows from the same counting: fixing all row and column totals leaves ( r − 1 ) ( c − 1 ) cells free, the rest being determined.

Why dividing by E i rather than by anything else. Under the null, the count in cell i has variance approximately E i for a Poisson-like count, so ( O i − E i ) / E i is a standardised discrepancy and its square is approximately a squared standard normal. Summing k of them, subject to the constraints, gives the chi-square distribution with the degrees of freedom above. The divisor is a variance, which is why proportions cannot be substituted for counts: rescaling changes the variance and the reference distribution no longer applies.

Fisher's exact test computes rather than approximates. For a 2 × 2 table with fixed margins, the null distribution of a single cell is hypergeometric, so the probability of a table at least as extreme can be summed directly. Nothing is approximated, which is why it remains valid where the chi-square approximation is doubtful; the cost is that it conditions on the observed margins, which is a different null from the one the chi-square test uses.

Rank-based procedures change the hypothesis. Replacing observations with ranks removes any assumption about distributional shape and also changes what is being tested, typically to a statement about location shift rather than about means. A difference in means can exist while the rank-based null holds, and a rank test can reject where means are equal, so the two are not interchangeable.

Intuition

Why fitting a parameter uses up evidence

Suppose a model is proposed and its parameter is chosen to make the model match the data as closely as possible. Then the match is asked to serve as evidence that the model is right. Part of the agreement was arranged, and the test has to discount it.

That is what subtracting a degree of freedom does. Degrees of freedom count how many independent ways the observed counts could have departed from the expected ones. Every constraint removes one: the counts must total n , and each fitted parameter pulls the expected counts toward the observed ones in a further direction.

The direction of the error is worth knowing, because it is the reverse of what most people expect. Using too many degrees of freedom compares the statistic against a distribution that is too spread out, so the tail probability comes out larger and the fit looks better. With a Poisson model fitted to six categories, the correct df = 4 gives p = 0.7517 ; ignoring the estimated parameter gives df = 5 and p = 0.8610 . The mistake raises the p -value rather than lowering it, so the output gives no sign of the error.

Why the discrepancies are divided by the expected count. A cell expecting 10 and observing 15 is behaving unusually; a cell expecting 200 and observing 205 is not, though both are 5 away. Under the null a count has variance roughly equal to its expectation, so dividing by E converts each raw discrepancy into a standardised one, and only standardised quantities can be added.

What independence means in a table. Build the table that the row and column totals alone would predict, assuming each classification says nothing about the other. Compare it with the observed table. A large statistic means the joint pattern carries information the margins do not, which is association, and only that. The test cannot say which variable moved, and a third variable influencing both produces exactly the same signal.

Why the approximation depends on the cells rather than on n . The chi-square shape appears when each standardised discrepancy behaves like a normal variable. A cell with a small expected count produces a lumpy, discrete, skewed distribution instead, and no total sample size repairs it: a study of 10,000 observations spread over categories where one expects 1.5 has the same problem as a small study.

Example

Five questions, and whether this test answers them

Do the digits of a measurement follow Benford's law? A fully specified null: the first-digit probabilities are log 10 ⁡ ( 1 + 1 / d ) for d = 1 , … , 9 , fixed in advance by the hypothesis rather than fitted. Eight degrees of freedom, and the test is exactly what it was designed for.

Do these counts follow a Poisson distribution? The mean has to come from somewhere. If it is estimated from the same counts, the degrees of freedom drop by one, and the test asks the narrower question of whether the data are compatible with some Poisson distribution rather than with a particular one. If the mean is given by theory, the question is sharper and the degrees of freedom stay.

Are treatment and outcome associated in this table? The test answers this. It does not answer whether treatment caused the outcome, and nothing in the table distinguishes association produced by a causal path from association produced by a third variable related to both. Where assignment was randomised, that separate design fact licenses the causal reading. The test contributes the same evidence either way.

Is the sample representative of the population? Partly. Comparing observed demographic counts with census proportions is a goodness-of-fit test with a specified null, so it can detect a discrepancy on the variables compared. Passing it says nothing about variables not tabulated, and a sample matching the census on age and region can still be unrepresentative on income.

Do two groups have the same mean response? Not this test. Chi-square on counts addresses categories; comparing means of a continuous response calls for a t test or ANOVA, and forcing the response into bins to obtain counts discards information and makes the answer depend on the bin boundaries chosen.

---

What separates the first two from the rest. In the first two the null predicts counts, which is the situation the statistic is built for. In the third the arithmetic applies and the interpretation is the difficult part. In the fourth the test is sound and its scope is narrower than the question. In the fifth the data are the wrong type, and converting them to counts to fit the method is choosing the tool before the question.

Procedure

Running the test, and deciding whether to trust it

To test goodness of fit against a specified distribution.

  1. State the null as probabilities p 1 , … , p k summing to 1, fixed before looking at the data.
  2. Compute expected counts E i = n p i .
  3. Check every E i . The working rule asks for at least 5 in each cell. If some fall short, combine adjacent categories. A decision made on substantive grounds, and recorded, since combining until the test passes is selecting on the outcome.
  4. Compute X 2 = ∑ i ( O i − E i ) 2 / E i .
  5. Degrees of freedom: k − 1 , minus one for each parameter estimated from this data.
  6. Refer to the chi-square distribution and report the statistic, the degrees of freedom and the p -value together. The statistic alone is uninterpretable without the degrees of freedom.

To test independence in an r × c table.

  1. Compute the margins R i , C j and the total N .
  2. Expected counts E i j = R i C j / N .
  3. Check them as above. For a 2 × 2 table with marginal counts, prefer Fisher's exact test outright rather than checking a threshold.
  4. Compute X 2 over all r c cells.
  5. Degrees of freedom ( r − 1 ) ( c − 1 ) .
  6. If the result is significant, look at the cells. The standardised residual ( O i j − E i j ) / E i j shows where the departure sits; a significant statistic says only that the table as a whole departs from independence.

To choose between a parametric test and a rank-based one.

  1. Name the assumption in doubt, normality, equal variances, or the influence of extreme values.
  2. Ask whether it matters at this sample size. A two-sample t test is robust to moderate non-normality at reasonable n ; it is not robust to a heavy tail with few observations.
  3. If switching, state what the alternative tests. A rank test addresses a shift in location, not a difference in means, and the two coincide only under extra assumptions about shape.
  4. Decide before seeing the result. Running both and reporting the smaller p -value invalidates both.

Checks. Confirm the expected counts sum to n , which catches most arithmetic errors at once. Confirm the degrees of freedom by counting constraints rather than recalling a formula. Where the smallest expected count is near the threshold and a 2 × 2 table is involved, compute the exact test as well: agreement settles the matter, and disagreement is itself the finding, as in the worked case where 0.0102 and 0.029973 fall either side of a 1% level.

Worked example

Four tests, and two places the answer turns

(a) Is the die fair? 300 rolls of a six-sided die:

Face123456
Observed435238614759
Expected505050505050

The null is fully specified, each p i = 1 / 6 , so E i = 300 / 6 = 50 throughout.

X 2 = ( − 7 ) 2 + 2 2 + ( − 12 ) 2 + 11 2 + ( − 3 ) 2 + 9 2 50 = 49 + 4 + 144 + 121 + 9 + 81 50 = 408 50 = 8.1600 .

With df = 6 − 1 = 5 , the critical value at the 5% level is 11.0705 and p = 0.147635 . Do not reject. Face 4 appeared 61 times against 50 expected, which looks large in isolation; across six faces a spread of this size is ordinary.

(b) Are the two classifications independent? A 2 × 3 table:

ABCTotal
Row 1304525100
Row 2202555100
Total507080200

Expected counts from the margins, E i j = R i C j / N :

ABC
Row 1 25.0 35.0 40.0
Row 2 25.0 35.0 40.0

The rows have equal totals, so each column's count splits evenly under independence. Then

X 2 = 18.9643 , df = ( 2 − 1 ) ( 3 − 1 ) = 2 , p = 0.00007620 .

The smallest expected count is 25.00 , comfortably above the threshold, so the approximation is sound. Reject. Column C carries the bulk of it: 25 observed against 40 expected in row 1, 55 against 40 in row 2.

(c) A case sitting exactly on the condition. A 2 × 2 table:

Total
Row 12810
Row 29312
Total111122

Expected counts are 5.0 , 5.0 , 6.0 , 6.0 . The smallest is exactly 5.00 , at the conventional threshold, not below it, so by the usual rule the approximation is admissible.

X 2 = 6.6000 , df = 1 , p ≈ 0.0102 .

Fisher's exact test, which sums hypergeometric probabilities and approximates nothing, gives

p = 0.029973 .

The two differ by a factor of about three, and they fall either side of 0.01 . Both procedures were applied correctly, the expected counts satisfy the stated rule, and the conclusion at a 1% level depends on which was used. The threshold of 5 is a working convention, not a guarantee that the approximation is accurate.

(d) The degrees of freedom, when a parameter was fitted. Frequencies of an event count over 90 observations:

Count012345+
Observed1826201484

A Poisson model is proposed, its mean estimated from this same data as λ ^ = 1.7778 . Expected counts follow:

Count012345+
Expected 15.211 27.042 24.037 14.244 6.331 3.134
X 2 = 1.9132 .

Six categories, one constraint from the total, one parameter estimated:

df = 6 − 1 − 1 = 4 , p = 0.7517 .

Had the estimation been ignored, df = 5 would give p = 0.8610 . The error does not cause a false rejection; it makes the model look like a better fit than the data support. That direction is why it survives review. An inflated p -value reads as reassurance.

What (a) and (d) do not establish. Neither large p -value shows the model is correct. Both report that the observed counts are compatible with it, which is a weaker claim that several other distributions would also satisfy.

Contrast

Pairs that differ in one decision

The chi-square approximation against Fisher's exact test, on one table.

χ 2 Fisher exact
p -value 0.0102 0.029973
significant at 5%yesyes
significant at 1%yesno
approximatesthe null distributionnothing

Same table, same margins, expected counts of 5.0 , 5.0 , 6.0 , 6.0 satisfying the conventional rule. The two procedures disagree by a factor of about three and straddle a 1% threshold. Neither was misapplied; the first is an approximation and the second is not.

Degrees of freedom with and without the fitted parameter.

df p
correct, one parameter estimated4 0.7517
estimation ignored5 0.8610

The error inflates the p -value, so the model appears to fit better. An error that made the fit look worse would be caught by the author defending the model; this one reads as confirmation.

A specified null against a fitted one.

Testing against Benford's law uses probabilities fixed by the hypothesis, and the full k − 1 degrees of freedom are available. Testing against "some Poisson distribution" spends one on the estimated mean, and the question answered narrows correspondingly: compatibility with a family, rather than with a stated distribution. The statistic is computed identically and the claims differ.

Association against causation, in the same table.

A significant test of independence establishes that the joint counts depart from what the margins predict. Whether that licenses a causal reading depends on how the data arose, randomised assignment, or an observational comparison where a third variable could produce the identical pattern. The table cannot distinguish the two, and reporting the p -value without saying which situation obtains leaves the reader to supply the stronger reading.

A parametric test against its rank-based counterpart.

The second is not the first with weaker assumptions. It tests a different null, usually about a shift in location, so a difference in means can exist where the rank null holds and a rank test can reject where means agree. Switching because an assumption looks shaky changes the question being asked, which should be stated rather than treated as a repair.

Warning

The errors that read as good news

A large p -value does not confirm the distribution. The test can only find evidence against the null. Failing to reject reports compatibility, and several distributions are typically compatible with the same counts. The Poisson fit in the worked example gives p = 0.7517 , which establishes that these 90 observations do not contradict a Poisson model, not that the counts are Poisson.

The distinction has a practical test. Genuine agreement survives more data; agreement arising from a lack of resolution does not. A goodness-of-fit test on 20 observations fails to reject almost every candidate, so quoting its p -value as support for a chosen model reports the study's weakness as if it were a finding.

Forgetting that a parameter was estimated inflates the p -value. df = 4 gives 0.7517 and df = 5 gives 0.8610 on the same statistic. Both look like reassurance, so nothing prompts a recheck. The rule is mechanical: count the constraints, and every parameter fitted from this data is one.

The threshold of 5 is a convention, not a guarantee. In the worked 2 × 2 case the expected counts are 5.0 , 5.0 , 6.0 , 6.0 , the rule is satisfied, and the approximation still returns 0.0102 where the exact test returns 0.029973 . For a 2 × 2 table the exact test is cheap, so there is little reason to rely on the approximation at all near the boundary.

Combining categories until the condition is met is a decision about the data. Merging cells to raise expected counts is legitimate when the merged categories mean something together, and it is selection on the outcome when the merge is chosen because it produces a publishable result. Decide the grouping before computing, and record it.

---

Two errors of type rather than of arithmetic.

Percentages are not counts. X 2 divides by the expected count because that is the variance of a count. Feeding it proportions or rates rescales the statistic by an arbitrary factor and the reference distribution no longer describes it. The computation completes and returns a number that means nothing.

Each observation must occupy exactly one cell. Repeated measurements on the same subject, or units that can be counted twice, inflate the apparent sample size and the statistic with it. This breaks the test rather than weakening it, and no correction to the degrees of freedom repairs it.

---

A significant table says the table departs from independence, and no more. It does not identify which cells carry the departure, standardised residuals do that, nor which variable moved, nor whether a third variable produced the pattern. Reporting only the p -value invites the reader to supply a direction the test never established.

Application

Where counts carry the argument

Audit and fraud screening. First-digit distributions of reported figures are compared against Benford's law, a fully specified null needing no fitted parameter. A significant departure is a reason to look, not a finding of fraud: legitimate processes with bounded ranges or assigned numbering also violate the law, so the test selects where to spend attention rather than concluding anything.

Genetics. Observed offspring counts across phenotypic classes are compared with the ratios a proposed inheritance model predicts. Because the ratios come from the model rather than the data, the full degrees of freedom are available and the test is sharp. Where a recombination parameter has to be estimated from the same cross, the correction applies and the question narrows to compatibility with a family of models.

Survey weighting. Sampled demographic counts are compared with population margins to decide whether weighting is needed. The test is sound and its scope is exactly the variables tabulated; a sample matching on age and region can be unrepresentative on anything not compared, so passing is not a certificate of representativeness.

Clinical trial baselines with small strata. Comparing adverse-event counts across arms produces tables where some expected counts are small, which is where the exact test earns its place. Regulatory analyses specify the procedure in advance precisely because choosing between the approximation and the exact test after seeing which gives the smaller p -value would invalidate both.

A/B testing on conversion counts. A 2 × 2 table of converted against not converted, by variant. The arithmetic is the least of it: the observations must be independent, so a user appearing in both arms or converting twice breaks the test, and running it repeatedly as data accumulate inflates the error rate regardless of how carefully each individual test is computed.

---

The recurring shape. In each case the test contributes one thing, whether the counts depart from what a stated hypothesis predicts, and the interpretation is supplied by how the data arose. Randomisation licenses a causal reading; a prespecified analysis licenses the p -value; independent observations license the reference distribution. None of those comes from the table.

Next step

Practice Testing Counts Against a Claim

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lesson

This is the last lesson in Inferential Statistics. Review the course map to see what is left.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.