Confidence Intervals for Experimental Research
An interval estimate is the same three ingredients as a test, rearranged: an estimate, a standard error, and a critical value that sets the coverage. What the confidence statement describes is the procedure's long-run behaviour, not the probability that one computed interval contains the parameter, and keeping that straight is what separates a reportable interval from a misreported one.
Definition
Every interval here has the form
Formal statement
Assumptions and scope
The confidence level describes the procedure across repetitions. Once computed, the endpoints are fixed and the parameter either lies within them or does not; no probability attaches to the particular interval.
A
interval and a two-sided level-test using the same assumptions and standard error agree: a null value outside the interval is rejected. The Wald interval for a proportion behaves poorly when counts are small or
is near 0 or 1: its coverage falls well below the nominal level, and because it is centred on with a symmetric width it can place an endpoint outside . At it returns zero width, asserting certainty from no positive observations. Wilson (score) and Clopper-Pearson (exact) intervals are the usual repairs, and neither leaves. The standard error must match the design. Paired data use
; independent groups use the Welch form; a randomized experiment uses the design-based standard error. Coverage rests on the sampling distribution being approximately as assumed, which for a mean comes from the central limit theorem and can fail for very small samples from skewed populations.
An interval reports magnitude and precision together, which a p-value does not. Reporting only whether an interval excludes the null discards most of what it was computed to convey.
A stated confidence level is NOMINAL. Coverage is exact only for the procedures whose distributional assumptions hold exactly: the normal-model mean interval with
known, the interval under normality, and Clopper-Pearson by construction (which is conservative, covering at least the nominal level). The Welch and large-sample proportion intervals are approximate, their coverage approaching the nominal level as grows and departing from it in small or skewed samples.
Worked material
Example
What an interval reports that a p-value does not
Two studies test the same intervention against a null of no effect. Both return
Study A. Estimated effect 12.0 points, 95% interval
Study B. Estimated effect 1.2 points, 95% interval
Identical p-values, and the two results say entirely different things.
Study A is compatible with an effect anywhere from negligible to very large. It establishes that something is probably happening and leaves the magnitude wide open. Study B pins the effect down tightly, and to a value that may be too small to matter.
What the p-value discarded. It reported only that both intervals exclude zero. Everything about magnitude and precision, which is what a decision needs, is in the interval and absent from the p-value.
Reading a null result the same way. Suppose a third study gives an estimate of 0.4 with interval
Non-example
Statements an interval does not support
"There is a 95% probability the true mean lies between 4.1 and 5.3." After computation, both endpoints are fixed numbers and the parameter is a fixed constant. Nothing random remains for the probability to describe.
"95% of the data fall in this interval." The interval describes a parameter, not the observations. The spread of the data is governed by
"95% of future samples will produce a mean inside this interval." That is a prediction interval for a future statistic, which is a different and wider construction.
"The intervals overlap, so the groups do not differ." Overlapping intervals for two group means do not imply a non-significant difference. The comparison needs an interval for the difference, built from the standard error of the difference.
"The effect is not significant, so there is no effect." A large p-value and an interval straddling zero are compatible with substantial effects in either direction, as the width shows directly.
A Wald interval for a proportion with few events. With 2 events in 40 trials,
Contrast
What the confidence level describes
| The procedure, before data | The computed interval, after data | |
|---|---|---|
| What is random | Which sample or assignment occurs | Nothing — the endpoints are fixed numbers |
| Correct statement | "This procedure covers the parameter 95% of the time" | "Either this interval contains the parameter or it does not" |
| Probability applies | Yes, to the procedure | No, not to this interval |
| What 95% refers to | Long-run success rate of the method | — |
Why the error is so natural. The interval is right there and the parameter is not, so it feels like a statement about where the parameter probably sits. The frequentist framework simply does not supply that: it assigns probabilities to procedures, never to fixed unknown constants.
A picture that helps. Imagine running the study 100 times, each time computing an interval. Roughly 95 of those intervals would cover the parameter and 5 would miss. You have one of the hundred and cannot know which kind it is. The 95% describes the collection, not your draw.
What you may legitimately say. That the interval contains the values not rejected by the corresponding two-sided test. The values compatible with your data at that level. This is the duality, and it is often the most useful reading in practice.
If you want a probability about the parameter, you need a Bayesian credible interval, which requires a prior and answers a different question.
Common errors
Common misconception
A 95% confidence interval has a 95% probability of containing the true parameter, so one can say the parameter is 95% likely to lie between the computed endpoints.
Related units
Requires
Connected
- Hypothesis Tests for Experimental Research (contrasts with)