Practice: Hypothesis Tests for Experimental Research

Recognition · Interpretation

Thirty staff each work one week with an old form and one week with a new form. Which test compares the mean times?

2 hints available, least help first.

Hint 1: Retrieval cue

Ask how many independent units the design produced.

Hint 2: Concept cue

If one person appears in both columns, are the columns independent?

Method selection · Direct application · Interpretation

A clinic randomizes 50 patients to a new intake process and 50 to the standard one. Mean waiting time is 23.4 minutes (new, s = 8.1 ) and 27.9 minutes (standard, s = 9.6 ).

Choose the test, compute the statistic, name the reference distribution, and state the conclusion at α = 0.05 .

A colleague proposes first testing whether the two variances are equal and pooling if that test does not reject. Say whether to do that.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask whether the two groups contain the same people or different people.

Hint 2: Strategy cue

Write the standard error first; the rest of the test follows from it.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The test. Two independent groups of different patients, so a two-sample test. Welch's version is the default, since nothing here justifies assuming equal variances, and the sample standard deviations do differ.

The standard error.

S E = 8.1 2 50 + 9.6 2 50 = 65.61 50 + 92.16 50 = 1.312 + 1.843 = 3.155 ≈ 1.776 .

The statistic.

T = 23.4 − 27.9 1.776 = − 4.5 1.776 ≈ − 2.53 .

Reference distribution. A t distribution with degrees of freedom from the Welch–Satterthwaite approximation, here roughly 95.

Conclusion. The two-sided critical value at α = 0.05 is about 1.99, and | − 2.53 | exceeds it, so the null of no difference is rejected; p ≈ 0.013 . The new process is associated with a mean waiting time about 4.5 minutes shorter.

What to add. An interval for the difference, roughly − 4.5 ± 1.99 × 1.776 = [ − 8.0 ,   − 1.0 ] minutes, which reports the magnitude the data support rather than only that a threshold was crossed. Because patients were randomized, the comparison supports a causal reading.

On pre-testing the variances. Do not. Welch's test makes no equal-variance assumption and is the safer default, and here s = 8.1 against 9.6 gives no reason to pool anyway. The deeper objection is procedural: a preliminary variance test uses the same data to choose the test and then to run it, so the reported error rate no longer describes what was done. Pool only on substantive grounds known before the data, never on the strength of a pre-test.

A complete answer does each of these:

  • matches test to design

Comparison · Evaluation

Two analyses of the same 20 people measured under conditions A and B give T = 1.30 and T = 3.83 from the identical mean difference of 3.6. What explains the discrepancy?

2 hints available, least help first.

Hint 1: Retrieval cue

The numerator is identical. Which part of the statistic could differ?

Hint 2: Concept cue

Ask what variation each standard error includes.

Direct application · Interpretation · Explanation

Two independent groups are compared. Group A: n 1 = 18 , x ¯ 1 = 52.4 , s 1 = 11.2 . Group B: n 2 = 22 , x ¯ 2 = 47.9 , s 2 = 6.1 .

(a) Compute the test statistic against a null of no difference, and name the reference distribution with its degrees of freedom.

(b) Say whether you pooled the variances, and why.

(c) The result is not significant at 0.05 . State precisely what that does and does not establish.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

The standard error of a difference of independent means combines the two variances.

Hint 2: Strategy cue

Before pooling, look at the two standard deviations and the two sample sizes together.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) The statistic. The standard error of the difference, without pooling:

S E = 11.2 2 18 + 6.1 2 22 = 6.969 + 1.691 = 8.660 ≈ 2.943
t = 52.4 − 47.9 2.943 = 4.5 2.943 ≈ 1.53

Referred to a t distribution with Welch's degrees of freedom:

ν = 8.660 2 6.969 2 17 + 1.691 2 21 = 74.99 2.857 + 0.136 ≈ 25.1

So t ≈ 1.53 on about 25 degrees of freedom.

(b) No pooling. The sample standard deviations differ by nearly a factor of two and the groups are unequal in size. The case in which pooling is least safe, because the larger group's smaller variance dominates the pooled estimate and understates the standard error. Welch's test is the default regardless: it costs almost nothing when variances are equal and protects the error rate when they are not.

Note what was NOT done: a preliminary F test of equal variances, used to decide whether to pool. That two-stage procedure distorts the error rate of the test that follows, because the second test is chosen using the same data it is applied to.

(c) What the result establishes. p > 0.05 , so the data do not distinguish the hypothesis of no difference from the alternatives nearby. It does not establish that the groups are the same. The observed difference is 4.5 with a standard error of 2.9 , so the compatible range runs from roughly − 1.6 to 10.6 , comfortably including zero, and also including differences large enough to matter. The correct report is that this study was not big enough to settle the question. Supporting a claim of no difference would require an equivalence test against a margin stated before the data were seen.

A complete answer does each of these:

  • assembles statistic
  • interprets non rejection
  • justifies variance handling

Recognition · Interpretation

A trial of 40 patients reports a difference in means of 1.8 with p = 0.34 . Which statement does the result support?

2 hints available, least help first.

Hint 1: Retrieval cue

What does a p-value condition on?

Hint 2: Concept cue

Ask which effect sizes this trial could have detected at all.

Direct application

Two independent groups give sample means differing by x ¯ 1 − x ¯ 2 = 5.0 . The standard error of that difference is 2.0 , computed without pooling the variances.

Under the null hypothesis of no difference in population means, report the value of the test statistic.

Enter the value. It is checked against the answer and the precision this task asks for.

Error diagnosis · Explanation · Evaluation

A team reports:

We compared conversion rates between the two variants. B converted at 4.1% versus 3.8% for A, with p = 0.31 . We conclude the variants perform equivalently and recommend choosing on development cost.

Identify the error and say what should be reported instead.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

State what a large p-value licenses, in one sentence.

Hint 2: Concept cue

Ask which other hypotheses these data are equally compatible with.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The error. Failure to reject is being converted into a positive claim of equivalence. A p-value of 0.31 says the data would not be surprising if the variants performed identically. It does not say they do.

What the result is also compatible with. A real difference the study was too small to detect. Data consistent with 'no difference' are equally consistent with 'B is 0.5 points better' and with 'B is 0.4 points worse', if the sample cannot separate those. The test asks about compatibility with one hypothesis, and compatibility does not single out the null.

Why this matters commercially. If a 0.5-percentage-point gain would change the decision, and the data cannot rule it in or out, then the study has not answered the question the recommendation depends on.

What should be reported. The estimated difference of 0.3 percentage points with a confidence interval, which shows the range of differences the data are compatible with. If that interval runs from about − 0.3 to + 0.9 points, it should be stated plainly that a commercially meaningful difference remains possible in either direction.

If equivalence is the actual claim. It requires an equivalence test against a pre-specified margin of practical indifference. A different procedure that must be planned in advance, not inferred from a non-significant result.

A further note on the standard error. Nothing in the report says how the two proportions' variability was combined. With unequal group sizes or unequal variances, a pooled standard error can understate uncertainty; the default should be the unpooled form, chosen in advance rather than after inspecting the variances.

A complete answer does each of these:

  • assembles statistic
  • interprets non rejection
  • justifies variance handling

Transfer · Evaluation · Explanation

A logistics firm tests two routing algorithms. Each of 18 depots runs algorithm X for a fortnight and algorithm Y for the next fortnight. The analyst compares the 18 X-fortnight averages against the 18 Y-fortnight averages with a two-sample t -test, d f = 34 .

Assess the analysis, say what it should be, and name a design concern the test cannot fix.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Count the independent units the design produced, not the measurements.

Hint 2: Strategy cue

Ask what every depot has in common between its two fortnights, and what differs besides the algorithm.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The structure. Each depot is measured under both algorithms, so these are paired measurements with the depot as the unit. This is a crossover design, and the vocabulary is the only thing unfamiliar about it.

Why the proposed test is wrong. A two-sample test computes its standard error from the spread across all 36 depot-fortnights, which includes the large differences between depots, size, geography, traffic, staffing. Those differences are identical for a depot's two fortnights and cancel completely in a within-depot comparison. Including them inflates the standard error and will usually understate what the study established.

The correct analysis. Form D i = Y ¯ i , X − Y ¯ i , Y for each of the 18 depots, then a one-sample t -test on those differences: T = D ¯ / ( s D / 18 ) with d f = 17 , not 34. The degrees of freedom follow from 18 independent units, not 36 measurements.

The design concern no test can fix. Every depot ran X first and Y second, so the algorithm is completely confounded with the time period. Any seasonal change in demand, any learning by depot staff, any unrelated operational change between the fortnights enters the estimate as though it were an algorithm effect. The remedy is design, not analysis: randomize the order so that half the depots receive Y first, which turns period into something that can be estimated and separated rather than something inseparable from the treatment.

Why the variance question does not arise here. Once the analysis is paired, each depot contributes one difference and there is a single sample of 18 differences with one variance. The Welch-versus-pooled question belongs to two independent samples; pairing removes it entirely, which is another reason the unit of analysis has to be settled before any variance comparison is contemplated.

A complete answer does each of these:

  • assembles statistic
  • interprets non rejection
  • justifies variance handling

Method selection · Classification

For each study, name the test the design supports and say what the unit of analysis is. Do not compute anything.

(a) Sixty patients each have blood pressure measured before and after a drug.

(b) Sixty patients are randomized, thirty to a drug and thirty to placebo, with blood pressure measured once.

(c) Twelve clinics are randomized, six to a new protocol; forty patients are measured at each clinic.

(d) One sample of sixty patients is compared against a published population mean of 140.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

For each design, ask what was randomized or what varies.

Hint 2: Concept cue

Two readings on one person are not two independent observations.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) Paired. The unit of analysis is the PATIENT, and each supplies one difference. A paired t test on the sixty differences, with 59 degrees of freedom. Treating the two sets of sixty readings as independent samples would ignore that each pair shares a person, and would usually overstate the standard error.

(b) Two independent samples. The unit is the patient, and each patient appears in one arm only. Welch's two-sample t test. The superficial resemblance to (a), sixty patients, two sets of readings, is exactly what misleads: what differs is whether a reading in one group is tied to a particular reading in the other.

(c) Clustered. The treatment was assigned to CLINICS, so the unit of analysis is the clinic and the effective sample size is twelve, not four hundred and eighty. A t test on the twelve clinic means, or a model with cluster-robust standard errors at the clinic level. Running a two-sample test on 480 patients would treat pupils within a clinic as independent evidence about a treatment that never varied within a clinic.

(d) One sample against a fixed value. The unit is the patient; the comparison value is a constant, not an estimate. A one-sample t test with 59 degrees of freedom.

The through-line. In every case the question is the same: what could have come out differently under the design? That is the unit of analysis, and it fixes the standard error before any arithmetic begins.

A complete answer does each of these:

  • matches test to design
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.