Confidence Intervals for Experimental Research

An interval estimate is the same three ingredients as a test, rearranged: an estimate, a standard error, and a critical value that sets the coverage. What the confidence statement describes is the procedure's long-run behaviour, not the probability that one computed interval contains the parameter, and keeping that straight is what separates a reportable interval from a misreported one.

Definition

Every interval here has the form estimate ± critical value × standard error , where z α / 2 and t α / 2 , d f denote positive upper-tail critical values. One mean: X ¯ ± z α / 2 σ n with σ known, or X ¯ ± t α / 2 , n − 1 s n when it is estimated. Difference in means: ( X ¯ 1 − X ¯ 2 ) ± t ∗ S E , with the Welch standard error S E = s 1 2 / n 1 + s 2 2 / n 2 ; for paired data, D ¯ ± t α / 2 , n − 1 s D n . One proportion (Wald): p ^ ± z α / 2 p ^ ( 1 − p ^ ) / n . Experimental effect: τ ^ ± z α / 2 S E ( τ ^ ) using the design-based standard error. A 100 ( 1 − α ) % confidence procedure is one whose intervals capture the fixed parameter in at least that proportion of repeated samples or assignments when its assumptions hold; where they hold only approximately, the realised coverage is approximate too.

Formal statement

estimate ± critical value × S E ; X ¯ ± t α / 2 , n − 1 s n ; ( X ¯ 1 − X ¯ 2 ) ± t ∗ s 1 2 / n 1 + s 2 2 / n 2 ; τ ^ ± z α / 2 S E ( τ ^ ) .

Assumptions and scope

  • The confidence level describes the procedure across repetitions. Once computed, the endpoints are fixed and the parameter either lies within them or does not; no probability attaches to the particular interval.

  • A 100 ( 1 − α ) % interval and a two-sided level- α test using the same assumptions and standard error agree: a null value outside the interval is rejected.

  • The Wald interval for a proportion behaves poorly when counts are small or p ^ is near 0 or 1: its coverage falls well below the nominal level, and because it is centred on p ^ with a symmetric width it can place an endpoint outside [ 0 , 1 ] . At p ^ = 0 it returns zero width, asserting certainty from no positive observations. Wilson (score) and Clopper-Pearson (exact) intervals are the usual repairs, and neither leaves [ 0 , 1 ] .

  • The standard error must match the design. Paired data use s D / n ; independent groups use the Welch form; a randomized experiment uses the design-based standard error.

  • Coverage rests on the sampling distribution being approximately as assumed, which for a mean comes from the central limit theorem and can fail for very small samples from skewed populations.

  • An interval reports magnitude and precision together, which a p-value does not. Reporting only whether an interval excludes the null discards most of what it was computed to convey.

  • A stated confidence level is NOMINAL. Coverage is exact only for the procedures whose distributional assumptions hold exactly: the normal-model mean interval with σ known, the t interval under normality, and Clopper-Pearson by construction (which is conservative, covering at least the nominal level). The Welch and large-sample proportion intervals are approximate, their coverage approaching the nominal level as n grows and departing from it in small or skewed samples.

Worked material

Example

What an interval reports that a p-value does not

Two studies test the same intervention against a null of no effect. Both return p = 0.04 .

Study A. Estimated effect 12.0 points, 95% interval [ 0.5 ,   23.5 ] .

Study B. Estimated effect 1.2 points, 95% interval [ 0.05 ,   2.35 ] .

Identical p-values, and the two results say entirely different things.

Study A is compatible with an effect anywhere from negligible to very large. It establishes that something is probably happening and leaves the magnitude wide open. Study B pins the effect down tightly, and to a value that may be too small to matter.

What the p-value discarded. It reported only that both intervals exclude zero. Everything about magnitude and precision, which is what a decision needs, is in the interval and absent from the p-value.

Reading a null result the same way. Suppose a third study gives an estimate of 0.4 with interval [ − 4.8 ,   5.6 ] and p = 0.88 . Reporting "no significant effect" suggests the intervention does nothing. The interval says the study is compatible with a harm of nearly 5 points and a benefit of over 5. It did not resolve the question at all. That is a different finding from a tight interval around zero, and only the interval distinguishes them.

Non-example

Statements an interval does not support

"There is a 95% probability the true mean lies between 4.1 and 5.3." After computation, both endpoints are fixed numbers and the parameter is a fixed constant. Nothing random remains for the probability to describe.

"95% of the data fall in this interval." The interval describes a parameter, not the observations. The spread of the data is governed by s , not by s / n , and at n = 400 these differ twentyfold.

"95% of future samples will produce a mean inside this interval." That is a prediction interval for a future statistic, which is a different and wider construction.

"The intervals overlap, so the groups do not differ." Overlapping intervals for two group means do not imply a non-significant difference. The comparison needs an interval for the difference, built from the standard error of the difference.

"The effect is not significant, so there is no effect." A large p-value and an interval straddling zero are compatible with substantial effects in either direction, as the width shows directly.

A Wald interval for a proportion with few events. With 2 events in 40 trials, p ^ ± z p ^ ( 1 − p ^ ) / n performs badly and can produce a lower endpoint below zero. Score or exact intervals are appropriate there.

Contrast

What the confidence level describes

The procedure, before dataThe computed interval, after data
What is randomWhich sample or assignment occursNothing — the endpoints are fixed numbers
Correct statement"This procedure covers the parameter 95% of the time""Either this interval contains the parameter or it does not"
Probability appliesYes, to the procedureNo, not to this interval
What 95% refers toLong-run success rate of the method—

Why the error is so natural. The interval is right there and the parameter is not, so it feels like a statement about where the parameter probably sits. The frequentist framework simply does not supply that: it assigns probabilities to procedures, never to fixed unknown constants.

A picture that helps. Imagine running the study 100 times, each time computing an interval. Roughly 95 of those intervals would cover the parameter and 5 would miss. You have one of the hundred and cannot know which kind it is. The 95% describes the collection, not your draw.

What you may legitimately say. That the interval contains the values not rejected by the corresponding two-sided test. The values compatible with your data at that level. This is the duality, and it is often the most useful reading in practice.

If you want a probability about the parameter, you need a Bayesian credible interval, which requires a prior and answers a different question.

Common errors

Common misconception

A 95% confidence interval has a 95% probability of containing the true parameter, so one can say the parameter is 95% likely to lie between the computed endpoints.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.