Practice: Confidence Intervals for Experimental Research
Question
Recognition · Interpretation
A study reports a 95% confidence interval of
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what is still random once the data are in hand.
Hint 2: Concept cue
The 95% attaches to something. Is it the interval, or the method that produced it?
Direct application · Interpretation · Explanation
A randomized trial of 64 staff, 32 per arm, measures minutes saved per shift. Treated: mean 14.2,
Build the 95% interval for the effect and write the sentence you would put in the report.
Finally: without computing a test statistic, say whether a two-sided test of no difference would reject at
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Write the template, then fill in estimate, critical value and standard error.
Hint 2: Strategy cue
After computing the endpoints, ask what a reader should conclude from the width.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
Estimate.
Standard error. Independent groups, Welch form:
Interval.
The sentence. "The estimated saving is 4.4 minutes per shift (95% CI: −0.1 to 8.9 minutes). The interval includes zero, so the difference is not significant at the 5% level; it is also compatible with savings of up to about 9 minutes, so this trial does not resolve whether the tool helps."
Why worded that way. Reporting only 'not significant' would suggest the tool was shown not to work. The interval's width is the actual finding: the study was too small to distinguish no effect from a substantial one. A useful next step states the smallest saving that would change the decision and sizes a study to detect it, noting that halving this width needs roughly four times the staff.
By the duality. A null value lying outside a 95% interval is rejected by the corresponding two-sided test at
A complete answer does each of these:
- constructs interval
- states coverage correctly
- reads magnitude and precision
- relates to testing
Comparison · Evaluation
Two studies both report
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what each interval says about plausible effect sizes.
Hint 2: Concept cue
What did the p-value discard that the interval retains?
Direct application · Construction
A sample of
2 hints available, least help first.
Hint 1: Retrieval cue
estimate
Hint 2: Concept cue
The standard error of a mean is
Error diagnosis · Explanation · Evaluation
A report states:
The 95% confidence interval for the mean reduction is 2.1 to 6.7 points. There is therefore a 95% probability that the true mean reduction lies between 2.1 and 6.7, and only a 5% chance it lies outside.
Identify the error and state what the interval does support.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask which quantities are random and which are fixed once the data exist.
Hint 2: Concept cue
The confidence level attaches to a procedure. What does that permit you to say about one interval?
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The error. A probability is being assigned to a fixed unknown constant. Once the data are observed, 2.1 and 6.7 are fixed numbers and the true mean reduction is a fixed value; either it lies between them or it does not. Nothing random remains for a probability to describe.
Where the 95% belongs. To the procedure, before the data arrive. Across repeated samples, 95% of the intervals this method produces would cover the parameter. You have one such interval and cannot know whether yours is among the 95% that cover or the 5% that miss.
A picture that makes it concrete. Imagine running the study 100 times and computing 100 intervals. About 95 would contain the parameter. The 95% characterises that collection, not any single member of it.
What the interval does support. That the values between 2.1 and 6.7 are those not rejected by the corresponding two-sided test at the 5% level. The effect sizes compatible with these data at that level. It also supports the substantive reading: the data indicate a reduction, and plausibly one between about 2 and 7 points, which is a useful statement about magnitude and precision.
If a probability about the parameter is genuinely wanted, that requires a Bayesian credible interval, which needs a prior and answers a different question.
What the interval does license. By the duality, it is the set of values a two-sided test at the matching level would not reject. That is a statement about which hypotheses are compatible with these data, not a probability distribution over the parameter, which is what the report claimed.
A complete answer does each of these:
- constructs interval
- states coverage correctly
- reads magnitude and precision
- relates to testing
Transfer · Evaluation · Explanation
A report states:
Group A averaged 71.2 (95% CI: 68.1–74.3) and group B averaged 76.8 (95% CI: 73.4–80.2). Since the intervals overlap slightly, the groups are not significantly different.
Assess the reasoning, compute what should have been computed, and state the conclusion.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what quantity the question is about, then build an interval for that quantity.
Hint 2: Strategy cue
Recover each standard error from its interval's width, then combine them.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The error. Comparing two intervals for separate means is not a test of their difference. Overlapping intervals are entirely compatible with a significant difference, because the standard error of a difference is smaller than the sum of the two individual half-widths.
Why.
What should be computed. An interval for the difference. Each half-width is about
The estimated difference is
The conclusion. The interval for the difference excludes zero, so the groups do differ significantly at the 5% level. The opposite of what was reported. The corresponding statistic is
What to report. The difference of 5.6 points with its interval of roughly 1 to 10, which states both that a difference is supported and how large it plausibly is.
Why overlap is the wrong comparison. The duality applies to an interval and the test of the value it was built for. Two separate intervals are not an interval for the difference: the relevant standard error is
A complete answer does each of these:
- constructs interval
- states coverage correctly
- reads magnitude and precision
- relates to testing
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.