Practice: Hypothesis Tests for Experimental Research
Question
Recognition · Interpretation
Thirty staff each work one week with an old form and one week with a new form. Which test compares the mean times?
2 hints available, least help first.
Hint 1: Retrieval cue
Ask how many independent units the design produced.
Hint 2: Concept cue
If one person appears in both columns, are the columns independent?
Method selection · Direct application · Interpretation
A clinic randomizes 50 patients to a new intake process and 50 to the standard one. Mean waiting time is 23.4 minutes (new,
Choose the test, compute the statistic, name the reference distribution, and state the conclusion at
A colleague proposes first testing whether the two variances are equal and pooling if that test does not reject. Say whether to do that.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask whether the two groups contain the same people or different people.
Hint 2: Strategy cue
Write the standard error first; the rest of the test follows from it.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The test. Two independent groups of different patients, so a two-sample test. Welch's version is the default, since nothing here justifies assuming equal variances, and the sample standard deviations do differ.
The standard error.
The statistic.
Reference distribution. A
Conclusion. The two-sided critical value at
What to add. An interval for the difference, roughly
On pre-testing the variances. Do not. Welch's test makes no equal-variance assumption and is the safer default, and here
A complete answer does each of these:
- matches test to design
Comparison · Evaluation
Two analyses of the same 20 people measured under conditions A and B give
2 hints available, least help first.
Hint 1: Retrieval cue
The numerator is identical. Which part of the statistic could differ?
Hint 2: Concept cue
Ask what variation each standard error includes.
Direct application · Interpretation · Explanation
Two independent groups are compared. Group A:
(a) Compute the test statistic against a null of no difference, and name the reference distribution with its degrees of freedom.
(b) Say whether you pooled the variances, and why.
(c) The result is not significant at
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
The standard error of a difference of independent means combines the two variances.
Hint 2: Strategy cue
Before pooling, look at the two standard deviations and the two sample sizes together.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) The statistic. The standard error of the difference, without pooling:
Referred to a
So
(b) No pooling. The sample standard deviations differ by nearly a factor of two and the groups are unequal in size. The case in which pooling is least safe, because the larger group's smaller variance dominates the pooled estimate and understates the standard error. Welch's test is the default regardless: it costs almost nothing when variances are equal and protects the error rate when they are not.
Note what was NOT done: a preliminary
(c) What the result establishes.
A complete answer does each of these:
- assembles statistic
- interprets non rejection
- justifies variance handling
Recognition · Interpretation
A trial of 40 patients reports a difference in means of
2 hints available, least help first.
Hint 1: Retrieval cue
What does a p-value condition on?
Hint 2: Concept cue
Ask which effect sizes this trial could have detected at all.
Direct application
Two independent groups give sample means differing by
Under the null hypothesis of no difference in population means, report the value of the test statistic.
Enter the value. It is checked against the answer and the precision this task asks for.
Error diagnosis · Explanation · Evaluation
A team reports:
We compared conversion rates between the two variants. B converted at 4.1% versus 3.8% for A, with
. We conclude the variants perform equivalently and recommend choosing on development cost.
Identify the error and say what should be reported instead.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
State what a large p-value licenses, in one sentence.
Hint 2: Concept cue
Ask which other hypotheses these data are equally compatible with.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The error. Failure to reject is being converted into a positive claim of equivalence. A p-value of 0.31 says the data would not be surprising if the variants performed identically. It does not say they do.
What the result is also compatible with. A real difference the study was too small to detect. Data consistent with 'no difference' are equally consistent with 'B is 0.5 points better' and with 'B is 0.4 points worse', if the sample cannot separate those. The test asks about compatibility with one hypothesis, and compatibility does not single out the null.
Why this matters commercially. If a 0.5-percentage-point gain would change the decision, and the data cannot rule it in or out, then the study has not answered the question the recommendation depends on.
What should be reported. The estimated difference of 0.3 percentage points with a confidence interval, which shows the range of differences the data are compatible with. If that interval runs from about
If equivalence is the actual claim. It requires an equivalence test against a pre-specified margin of practical indifference. A different procedure that must be planned in advance, not inferred from a non-significant result.
A further note on the standard error. Nothing in the report says how the two proportions' variability was combined. With unequal group sizes or unequal variances, a pooled standard error can understate uncertainty; the default should be the unpooled form, chosen in advance rather than after inspecting the variances.
A complete answer does each of these:
- assembles statistic
- interprets non rejection
- justifies variance handling
Transfer · Evaluation · Explanation
A logistics firm tests two routing algorithms. Each of 18 depots runs algorithm X for a fortnight and algorithm Y for the next fortnight. The analyst compares the 18 X-fortnight averages against the 18 Y-fortnight averages with a two-sample
Assess the analysis, say what it should be, and name a design concern the test cannot fix.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Count the independent units the design produced, not the measurements.
Hint 2: Strategy cue
Ask what every depot has in common between its two fortnights, and what differs besides the algorithm.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The structure. Each depot is measured under both algorithms, so these are paired measurements with the depot as the unit. This is a crossover design, and the vocabulary is the only thing unfamiliar about it.
Why the proposed test is wrong. A two-sample test computes its standard error from the spread across all 36 depot-fortnights, which includes the large differences between depots, size, geography, traffic, staffing. Those differences are identical for a depot's two fortnights and cancel completely in a within-depot comparison. Including them inflates the standard error and will usually understate what the study established.
The correct analysis. Form
The design concern no test can fix. Every depot ran X first and Y second, so the algorithm is completely confounded with the time period. Any seasonal change in demand, any learning by depot staff, any unrelated operational change between the fortnights enters the estimate as though it were an algorithm effect. The remedy is design, not analysis: randomize the order so that half the depots receive Y first, which turns period into something that can be estimated and separated rather than something inseparable from the treatment.
Why the variance question does not arise here. Once the analysis is paired, each depot contributes one difference and there is a single sample of 18 differences with one variance. The Welch-versus-pooled question belongs to two independent samples; pairing removes it entirely, which is another reason the unit of analysis has to be settled before any variance comparison is contemplated.
A complete answer does each of these:
- assembles statistic
- interprets non rejection
- justifies variance handling
Method selection · Classification
For each study, name the test the design supports and say what the unit of analysis is. Do not compute anything.
(a) Sixty patients each have blood pressure measured before and after a drug.
(b) Sixty patients are randomized, thirty to a drug and thirty to placebo, with blood pressure measured once.
(c) Twelve clinics are randomized, six to a new protocol; forty patients are measured at each clinic.
(d) One sample of sixty patients is compared against a published population mean of 140.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
For each design, ask what was randomized or what varies.
Hint 2: Concept cue
Two readings on one person are not two independent observations.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) Paired. The unit of analysis is the PATIENT, and each supplies one difference. A paired
(b) Two independent samples. The unit is the patient, and each patient appears in one arm only. Welch's two-sample
(c) Clustered. The treatment was assigned to CLINICS, so the unit of analysis is the clinic and the effective sample size is twelve, not four hundred and eighty. A
(d) One sample against a fixed value. The unit is the patient; the comparison value is a constant, not an estimate. A one-sample
The through-line. In every case the question is the same: what could have come out differently under the design? That is the unit of analysis, and it fixes the standard error before any arithmetic begins.
A complete answer does each of these:
- matches test to design
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.