Practice: Testing Counts Against a Claim

Direct application

A six-sided die is rolled 300 times, giving the counts

Face123456
Observed435238614759

Under the hypothesis that the die is fair, compute X 2 = ∑ ( O − E ) 2 / E . Give your answer to four decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

Under fairness each face is equally likely, so each expected count is 300 / 6 .

Hint 2: Next step

Square each difference from 50, add the six squares, and divide the total by 50.

Direct application

A 2 × 3 table has row totals 100 and 100 , column totals 50 , 70 and 80 , and grand total 200 . Under the hypothesis that the row and column classifications are independent, what is the expected count in row 1, column 3?

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

The expected count uses only the margins: the row total, the column total and the grand total.

Hint 2: Next step

E i j = R i C j / N .

Direct application

Counts are grouped into six categories and compared against a Poisson model whose mean was estimated from those same counts. How many degrees of freedom does the chi-square test have?

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

Start from the number of categories and subtract one constraint for the total.

Hint 2: Concept cue

Estimating a parameter from the same data pulls the expected counts toward the observed ones, which costs a further degree of freedom.

Error diagnosis · Evaluation

A 2 × 2 table has expected counts 5.0 , 5.0 , 6.0 and 6.0 . The chi-square test returns p = 0.0102 and Fisher's exact test returns p = 0.029973 . An analyst reports the result as significant at the 1% level, citing the chi-square value and noting that all expected counts meet the conventional minimum of 5. What should be said?

Method selection

A study compares whether a rare complication occurred, against which of two surgical approaches was used. Thirty patients received each approach; the complication occurred four times in total. Which procedure should be used to test for association, and why?

Interpretation

A goodness-of-fit test of 90 counts against a Poisson model, with the mean estimated from the same counts, gives X 2 = 1.9132 on 4 degrees of freedom and p = 0.7517 . A report concludes: "The data are Poisson distributed." What is the accurate statement?

Transfer · Evaluation

An A/B test dashboard shows a 2 × 2 table of conversions by variant and recomputes a chi-square test of independence every time new visitors arrive. The team watches it and stops the experiment the first time p falls below 0.05 , reporting that result. Each visitor is counted once and expected counts are in the hundreds. What is the defect?

Construction · Evaluation · Explanation

A hospital records the number of emergency admissions in each of 120 consecutive night shifts, grouped as follows:

Admissions012345 or more
Shifts14313322146

An analyst proposes that admissions follow a Poisson distribution, estimates its mean from these very counts, computes X 2 = 3.84 , looks it up on 5 degrees of freedom to obtain p = 0.57 , and reports that "admissions are Poisson."

Separately, the analyst has a 2 × 2 table comparing whether a senior doctor was on duty against whether a shift was busy, with expected counts of 4.8 , 5.2 , 6.1 and 6.9 , and plans to use a chi-square test on it.

Work through the following.

  1. The expected counts. Say where the expected counts for the Poisson test come from and what must be checked about them before the statistic is used.
  2. The degrees of freedom. State the correct value, justify it by counting constraints, and say what the analyst's choice does to the reported p -value and in which direction.
  3. The conclusion. Assess the sentence "admissions are Poisson", and give the statement the test actually supports.
  4. The second table. Say whether the chi-square test is appropriate there, what you would do instead, and why the choice matters at these counts.
  5. A different question. The analyst now wants to know whether median admissions differ between winter and summer shifts, and proposes a rank-based test because the counts are skewed. Say what that test's null hypothesis is and whether it answers the question asked.

Write your answer, then compare it with the worked solution.

3 hints available, least help first.

Hint 1: Retrieval cue

For part 2, count the constraints: the total, and anything fitted from the data.

Hint 2: Concept cue

For part 3, ask what a test can establish and what it can only fail to find.

Hint 3: Strategy cue

For part 5, write down the null of the rank test explicitly before deciding whether it matches the question.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

1. The expected counts. They are E i = n p ^ i , where p ^ i is the Poisson probability of category i evaluated at the estimated mean, and n = 120 . The final category is "5 or more", so its probability is 1 minus the sum of the others rather than a single Poisson term. A tail category, not a point mass at 5. Before using the statistic, every expected count must be checked against the conventional minimum of about 5. The tail category is the one at risk here: with a mean near 2, P ( X ≥ 5 ) is small, and 120 times a small probability can fall below the threshold. If it does, the remedy is to merge "4" and "5 or more" into "4 or more", decided on substantive grounds and recorded, not chosen after seeing which grouping gives the preferred answer. 2. The degrees of freedom. Six categories. One constraint because the expected counts must total 120 . One more because the mean was estimated from these same counts, which pulled the expected counts toward the observed ones. So

df = 6 − 1 − 1 = 4 ,

not 5. Using 5 compares X 2 against a distribution that is too spread out, which makes the tail probability larger. The reported p = 0.57 therefore overstates the agreement; the correct degrees of freedom give a smaller p -value on the same statistic. The direction matters: the error raises the p -value rather than lowering it, so the output gives no sign of the mistake and the mistake survives review. 3. The conclusion. "Admissions are Poisson" is not supported. A goodness-of-fit test accumulates evidence only against the null, so failing to reject reports compatibility, not truth. Several distributions, a negative binomial with modest overdispersion, for one, would also be compatible with 120 shifts of counts like these. The supported statement is: these counts show no evidence of departure from a Poisson distribution, at the sample size available. Worth adding that 120 observations spread over six categories gives limited power, so a large p -value here partly reports the study's resolution rather than the model's quality. A Poisson model also implies variance equal to mean, which can be checked directly and is a sharper diagnostic than the omnibus test. 4. The second table. The chi-square test is a poor choice. Two expected counts, 4.8 and 5.2 , sit right at the conventional threshold, and that threshold is a working convention rather than a guarantee of accuracy. Use Fisher's exact test. For a 2 × 2 table it computes the null probability directly from the hypergeometric distribution with no approximation, and it is cheap. It should be chosen in advance, not after comparing both p -values. Why it matters at these counts: the approximation and the exact test can differ substantially in this range. In the unit's worked case, expected counts of 5.0 , 5.0 , 6.0 and 6.0 , satisfying the rule, gave p = 0.0102 by chi-square and p = 0.029973 exactly, a factor of about three, falling either side of a 1% level. The two nulls differ: the exact test conditions on the observed margins. 5. A different question. A rank-based two-sample test, such as Wilcoxon-Mann-Whitney, does not have "the medians are equal" as its null. Its null is that a randomly chosen observation from one group is equally likely to exceed or fall below one from the other, that the two distributions are the same in stochastic ordering. It coincides with a statement about medians only under an extra assumption, typically that the two distributions differ by a shift and have the same shape. So it does not directly answer the question as posed. If the medians are the quantity of interest, estimate them with a confidence interval for the difference, and if a rank test is used, report what it tested rather than relabelling it a test of medians. The skew is a legitimate reason to avoid assuming normality; it is not a reason to treat the rank test as the same test with weaker requirements.

A complete answer does each of these:

  • computes expected counts
  • assembles chi square
  • derives degrees of freedom
  • checks approximation condition
  • selects nonparametric alternative
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.