Confidence Intervals for Experimental Research

What you will be able to do

Given an estimate and its standard error, the learner can construct the appropriate interval and state what the confidence level does and does not claim.

Orientation

An estimate without uncertainty answers only half the question. The interval says which other values the data are compatible with.

Constructing the interval takes one line. Interpreting it is harder, and misreading it as a probability statement about the parameter is among the most common errors in applied statistics.

Intuition

From a point estimate to an interval estimate

The canonical text gives the construction and the trade between coverage and width. The remaining difficulty is entirely in the interpretation, and it is not pedantry.

The probability belongs to the procedure, before the data arrive. Repeat the study many times and 95% of the intervals produced will contain the fixed parameter. That is a statement about the method. Once this interval is computed it either contains the parameter or it does not, and nothing about the arithmetic reveals which.

Why that matters in practice. Read as “a 95% chance the parameter is in here”, an interval starts supporting claims the procedure never made: that values near the centre are more likely than values near the edge, or that a parameter just outside has been ruled out. Neither follows. The interval is the set of null values a two-sided test at the matching level would not reject, which is a statement about compatibility rather than about probability.

The temptation is strongest when the result matters most. A borderline interval invites the Bayesian reading precisely because the frequentist one refuses to answer the question the reader wants answered. Answering it properly needs a prior and a credible interval, which is a different calculation reported under a different name.

Definition

The template and its instances

The canonical statement gives the forms. Three things decide whether an interval computed from them means what it is reported to mean.

The confidence level describes the procedure, not the interval in front of you. A 95% procedure produces intervals that cover the fixed parameter in 95% of repetitions. Once computed, this interval either contains the parameter or does not; the probability is not 0.95, it is 0 or 1 and unknown which. An interval read as “a 95% probability the parameter lies inside” is being asked for a Bayesian credible interval, which needs a prior and answers a different question.

Which critical value depends on what was estimated. z applies when the standard deviation is known, which is rare; t n − 1 applies when it was estimated from the same data, and the extra width is the price of that estimation. Welch's degrees of freedom for a difference are not n 1 + n 2 − 2 unless the variances are pooled, and pooling is an assumption rather than a simplification.

The Wald proportion interval fails exactly where it is most used. p ^ ± z p ^ ( 1 − p ^ ) / n has coverage well below its nominal level for small n or p ^ near 0 or 1, and at p ^ = 0 it returns an interval of zero width, which asserts certainty from no positive observations. Wilson or Agresti-Coull intervals are the usual repair.

Example

What an interval reports that a p-value does not

Two studies test the same intervention against a null of no effect. Both return p = 0.04 .

Study A. Estimated effect 12.0 points, 95% interval [ 0.5 ,   23.5 ] .

Study B. Estimated effect 1.2 points, 95% interval [ 0.05 ,   2.35 ] .

Identical p-values, and the two results say entirely different things.

Study A is compatible with an effect anywhere from negligible to very large. It establishes that something is probably happening and leaves the magnitude wide open. Study B pins the effect down tightly, and to a value that may be too small to matter.

What the p-value discarded. It reported only that both intervals exclude zero. Everything about magnitude and precision, which is what a decision needs, is in the interval and absent from the p-value.

Reading a null result the same way. Suppose a third study gives an estimate of 0.4 with interval [ − 4.8 ,   5.6 ] and p = 0.88 . Reporting "no significant effect" suggests the intervention does nothing. The interval says the study is compatible with a harm of nearly 5 points and a benefit of over 5. It did not resolve the question at all. That is a different finding from a tight interval around zero, and only the interval distinguishes them.

Worked example

Building and reporting an interval

Problem. A randomized trial of a scheduling tool measures minutes saved per shift. 64 staff are randomized, 32 per arm. The treated arm averages 14.2 minutes saved with s 1 = 9.6 ; the control arm averages 9.8 with s 0 = 8.4 .

Build a 95% interval for the effect and report it properly.

Goal. An interval with the right standard error, and a correctly worded conclusion.

Relevant principle. estimate ± critical value × S E , with S E matching the design, here two independent groups.

Step 1: the estimate.

τ ^ = 14.2 − 9.8 = 4.4  minutes .

Step 2: the standard error. Independent groups with no equal-variance assumption:

S E = 9.6 2 32 + 8.4 2 32 = 92.16 32 + 70.56 32 = 2.88 + 2.205 = 5.085 ≈ 2.255 .

Reason: the Welch form, the safer default; nothing here justifies assuming equal variances.

Step 3: the critical value. With roughly 61 degrees of freedom by Welch–Satterthwaite, t 0.025 , 61 ≈ 2.00 .

Reason: the standard error was estimated from the data, so the reference is t rather than z ; at this many degrees of freedom the two nearly coincide.

Step 4: assemble.

4.4 ± 2.00 × 2.255 = 4.4 ± 4.51 = [ − 0.11 ,   8.91 ] .

Step 5: report it correctly. "The estimated saving is 4.4 minutes per shift (95% CI: −0.1 to 8.9). The interval includes zero, so the difference is not significant at the 5% level, but it is also compatible with savings up to about 9 minutes. This study does not resolve whether the tool helps."

Reason: the interval's width is the finding here. Reporting only "not significant" would suggest the tool was shown not to work, which this does not show.

Result. τ ^ = 4.4 minutes, 95% CI [ − 0.1 ,   8.9 ] , inconclusive.

Check. Does the duality hold? Zero lies just inside the interval, so a two-sided test at α = 0.05 should just fail to reject: T = 4.4 / 2.255 ≈ 1.95 against a critical value of 2.00. Consistent.

Interpretation. The trial was too small. A useful follow-up states the smallest saving that would change the decision and sizes the next study to distinguish it, noting that halving this interval's width needs roughly four times the staff.

Non-example

Statements an interval does not support

"There is a 95% probability the true mean lies between 4.1 and 5.3." After computation, both endpoints are fixed numbers and the parameter is a fixed constant. Nothing random remains for the probability to describe.

"95% of the data fall in this interval." The interval describes a parameter, not the observations. The spread of the data is governed by s , not by s / n , and at n = 400 these differ twentyfold.

"95% of future samples will produce a mean inside this interval." That is a prediction interval for a future statistic, which is a different and wider construction.

"The intervals overlap, so the groups do not differ." Overlapping intervals for two group means do not imply a non-significant difference. The comparison needs an interval for the difference, built from the standard error of the difference.

"The effect is not significant, so there is no effect." A large p-value and an interval straddling zero are compatible with substantial effects in either direction, as the width shows directly.

A Wald interval for a proportion with few events. With 2 events in 40 trials, p ^ ± z p ^ ( 1 − p ^ ) / n performs badly and can produce a lower endpoint below zero. Score or exact intervals are appropriate there.

Contrast

What the confidence level describes

The procedure, before dataThe computed interval, after data
What is randomWhich sample or assignment occursNothing — the endpoints are fixed numbers
Correct statement"This procedure covers the parameter 95% of the time""Either this interval contains the parameter or it does not"
Probability appliesYes, to the procedureNo, not to this interval
What 95% refers toLong-run success rate of the method—

Why the error is so natural. The interval is right there and the parameter is not, so it feels like a statement about where the parameter probably sits. The frequentist framework simply does not supply that: it assigns probabilities to procedures, never to fixed unknown constants.

A picture that helps. Imagine running the study 100 times, each time computing an interval. Roughly 95 of those intervals would cover the parameter and 5 would miss. You have one of the hundred and cannot know which kind it is. The 95% describes the collection, not your draw.

What you may legitimately say. That the interval contains the values not rejected by the corresponding two-sided test. The values compatible with your data at that level. This is the duality, and it is often the most useful reading in practice.

If you want a probability about the parameter, you need a Bayesian credible interval, which requires a prior and answers a different question.

Exercise

1: fully structured. A sample of 36 has mean 82 and s = 12 .

(a) Compute the 95% interval for the mean, using t 0.025 , 35 ≈ 2.03 . (b) State what the 95% refers to. (c) How would a 99% interval differ?

Check: (a) S E = 12 / 6 = 2 , so 82 ± 2.03 × 2 = 82 ± 4.06 = [ 77.9 ,   86.1 ] ; (b) the long-run coverage of the procedure, across repetitions, 95% of intervals so constructed contain the parameter; it is not a probability about these endpoints; (c) wider, since a higher confidence level requires a larger critical value at the same standard error.

2: partly structured. A trial reports an estimated effect of 2.1 with 95% interval [ − 0.4 ,   4.6 ] .

(a) Would a two-sided test at α = 0.05 reject a null of zero? (b) Is this evidence the treatment does nothing? (c) What single sentence would you write?

Check: (a) no, zero lies inside the interval, and by the duality it is not rejected; (b) no, the interval is also compatible with a benefit of 4.6, so the study did not resolve the question; (c) something like: "The estimated effect is 2.1 (95% CI: −0.4 to 4.6); the data are compatible with anything from a small harm to a substantial benefit, so this trial is inconclusive."

3: unstructured. A report states: "Group A averaged 71.2 (95% CI: 68.1–74.3) and group B averaged 76.8 (95% CI: 73.4–80.2). Since the intervals overlap slightly, the groups are not significantly different."

Assess the reasoning and say what should be computed instead.

Check: the inference is invalid. Comparing two intervals for separate means is not a test of their difference; overlapping intervals can accompany a significant difference, because the standard error of the difference is S E A 2 + S E B 2 , smaller than the sum of the two half-widths. Here the estimated difference is 5.6, and with roughly S E A ≈ 1.58 and S E B ≈ 1.73 , S E diff ≈ 2.50 + 2.99 ≈ 2.34 , giving T ≈ 2.39 , likely significant at the 5% level despite the overlap. What should be computed is an interval for the difference itself, roughly 5.6 ± 2 × 2.34 = [ 0.9 ,   10.3 ] , which excludes zero and also reports the magnitude the comparison supports.

What to carry forward

The template. estimate ± critical value × S E , with the standard error matching the design.

One mean. X ¯ ± t α / 2 , n − 1 s n when σ is estimated.

Difference in means. ( X ¯ 1 − X ¯ 2 ) ± t ∗ s 1 2 / n 1 + s 2 2 / n 2 ; paired data use D ¯ ± t ∗ s D / n .

Proportions. The Wald interval uses p ^ ; with few events or p ^ near 0 or 1, prefer score or exact intervals.

Coverage. A property of the procedure across repetitions. Once computed, the endpoints are fixed and no probability attaches to the particular interval.

Duality. A null value outside a 100 ( 1 − α ) % interval is rejected by the corresponding two-sided test, so the interval is the set of values compatible with the data at that level.

Width is information. A wide interval around zero means the study was uninformative, not that the effect is absent.

Do not compare two intervals to test a difference. Build an interval for the difference, using the standard error of the difference.

The recurring error. "There is a 95% probability the parameter lies in this interval."

Next step

Practice Confidence Intervals for Experimental Research

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.