Practice: Confidence Intervals for Experimental Research

Recognition · Interpretation

A study reports a 95% confidence interval of [ 4.1 ,   5.3 ] for a population mean. Which statement is correct?

2 hints available, least help first.

Hint 1: Retrieval cue

Ask what is still random once the data are in hand.

Hint 2: Concept cue

The 95% attaches to something. Is it the interval, or the method that produced it?

Direct application · Interpretation · Explanation

A randomized trial of 64 staff, 32 per arm, measures minutes saved per shift. Treated: mean 14.2, s = 9.6 . Control: mean 9.8, s = 8.4 . Use t ∗ ≈ 2.00 .

Build the 95% interval for the effect and write the sentence you would put in the report.

Finally: without computing a test statistic, say whether a two-sided test of no difference would reject at α = 0.05 , and why you can tell.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Write the template, then fill in estimate, critical value and standard error.

Hint 2: Strategy cue

After computing the endpoints, ask what a reader should conclude from the width.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

Estimate. τ ^ = 14.2 − 9.8 = 4.4 minutes.

Standard error. Independent groups, Welch form:

S E = 9.6 2 32 + 8.4 2 32 = 2.88 + 2.205 = 5.085 ≈ 2.255 .

Interval.

4.4 ± 2.00 × 2.255 = 4.4 ± 4.51 = [ − 0.11 ,   8.91 ] .

The sentence. "The estimated saving is 4.4 minutes per shift (95% CI: −0.1 to 8.9 minutes). The interval includes zero, so the difference is not significant at the 5% level; it is also compatible with savings of up to about 9 minutes, so this trial does not resolve whether the tool helps."

Why worded that way. Reporting only 'not significant' would suggest the tool was shown not to work. The interval's width is the actual finding: the study was too small to distinguish no effect from a substantial one. A useful next step states the smallest saving that would change the decision and sizes a study to detect it, noting that halving this width needs roughly four times the staff.

By the duality. A null value lying outside a 95% interval is rejected by the corresponding two-sided test at α = 0.05 , and one lying inside is not. So the interval settles the test without a separate calculation, provided both use the same standard error and degrees of freedom. Read that way the interval carries more than the test does: it names every value the data do not reject, not merely whether one particular value survived.

A complete answer does each of these:

  • constructs interval
  • states coverage correctly
  • reads magnitude and precision
  • relates to testing

Comparison · Evaluation

Two studies both report p = 0.04 . Study A: effect 12.0, 95% CI [ 0.5 ,   23.5 ] . Study B: effect 1.2, 95% CI [ 0.05 ,   2.35 ] . What distinguishes them?

2 hints available, least help first.

Hint 1: Retrieval cue

Ask what each interval says about plausible effect sizes.

Hint 2: Concept cue

What did the p-value discard that the interval retains?

Direct application · Construction

A sample of n = 25 has x ¯ = 50 and s = 10 . Using t 24 , 0.975 = 2.064 , what is the 95% confidence interval for the mean?

2 hints available, least help first.

Hint 1: Retrieval cue

estimate ± critical value × standard error.

Hint 2: Concept cue

The standard error of a mean is s / n .

Error diagnosis · Explanation · Evaluation

A report states:

The 95% confidence interval for the mean reduction is 2.1 to 6.7 points. There is therefore a 95% probability that the true mean reduction lies between 2.1 and 6.7, and only a 5% chance it lies outside.

Identify the error and state what the interval does support.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask which quantities are random and which are fixed once the data exist.

Hint 2: Concept cue

The confidence level attaches to a procedure. What does that permit you to say about one interval?

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The error. A probability is being assigned to a fixed unknown constant. Once the data are observed, 2.1 and 6.7 are fixed numbers and the true mean reduction is a fixed value; either it lies between them or it does not. Nothing random remains for a probability to describe.

Where the 95% belongs. To the procedure, before the data arrive. Across repeated samples, 95% of the intervals this method produces would cover the parameter. You have one such interval and cannot know whether yours is among the 95% that cover or the 5% that miss.

A picture that makes it concrete. Imagine running the study 100 times and computing 100 intervals. About 95 would contain the parameter. The 95% characterises that collection, not any single member of it.

What the interval does support. That the values between 2.1 and 6.7 are those not rejected by the corresponding two-sided test at the 5% level. The effect sizes compatible with these data at that level. It also supports the substantive reading: the data indicate a reduction, and plausibly one between about 2 and 7 points, which is a useful statement about magnitude and precision.

If a probability about the parameter is genuinely wanted, that requires a Bayesian credible interval, which needs a prior and answers a different question.

What the interval does license. By the duality, it is the set of values a two-sided test at the matching level would not reject. That is a statement about which hypotheses are compatible with these data, not a probability distribution over the parameter, which is what the report claimed.

A complete answer does each of these:

  • constructs interval
  • states coverage correctly
  • reads magnitude and precision
  • relates to testing

Transfer · Evaluation · Explanation

A report states:

Group A averaged 71.2 (95% CI: 68.1–74.3) and group B averaged 76.8 (95% CI: 73.4–80.2). Since the intervals overlap slightly, the groups are not significantly different.

Assess the reasoning, compute what should have been computed, and state the conclusion.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask what quantity the question is about, then build an interval for that quantity.

Hint 2: Strategy cue

Recover each standard error from its interval's width, then combine them.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The error. Comparing two intervals for separate means is not a test of their difference. Overlapping intervals are entirely compatible with a significant difference, because the standard error of a difference is smaller than the sum of the two individual half-widths.

Why. S E diff = S E A 2 + S E B 2 , and a sum of squares under a root is less than the sum of its parts. Judging by eye whether two intervals touch applies the wrong yardstick.

What should be computed. An interval for the difference. Each half-width is about 2 × S E , so S E A ≈ ( 74.3 − 68.1 ) / 4 ≈ 1.55 and S E B ≈ ( 80.2 − 73.4 ) / 4 ≈ 1.70 . Then

S E diff = 1.55 2 + 1.70 2 = 2.40 + 2.89 = 5.29 ≈ 2.30 .

The estimated difference is 76.8 − 71.2 = 5.6 , giving

5.6 ± 2.00 × 2.30 = 5.6 ± 4.60 = [ 1.0 ,   10.2 ] .

The conclusion. The interval for the difference excludes zero, so the groups do differ significantly at the 5% level. The opposite of what was reported. The corresponding statistic is T = 5.6 / 2.30 ≈ 2.43 , beyond the critical value of about 2.

What to report. The difference of 5.6 points with its interval of roughly 1 to 10, which states both that a difference is supported and how large it plausibly is.

Why overlap is the wrong comparison. The duality applies to an interval and the test of the value it was built for. Two separate intervals are not an interval for the difference: the relevant standard error is S E A 2 + S E B 2 , which is smaller than the sum of the two half-widths. So intervals can overlap while a two-sided test of equal means rejects. Build the interval for the difference and apply the duality to that.

A complete answer does each of these:

  • constructs interval
  • states coverage correctly
  • reads magnitude and precision
  • relates to testing
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.