Neyman Repeated-Sampling Inference
Hold the potential outcomes fixed and let the assignment vary: that is the frame in which a difference in means has a variance at all. The exact design variance contains a term built from both potential outcomes per unit, which no study observes, so the estimator used in practice deliberately drops it. The result is conservative rather than exact, and knowing which it is changes what an interval claims.
Definition
For a completely randomized two-arm experiment with
Formal statement
Assumptions and scope
The potential outcomes are treated as fixed constants and the assignment as the only source of randomness. This is a different frame from classical sampling inference, where the units are a random draw from a population; the arithmetic often coincides and the claims do not.
is non-negative, so omitting it can only inflate the variance. The estimator is conservative in a precise sense: it does not understate uncertainty, and it is exact when every unit shares the same treatment effect. Constant treatment effects are a much stronger condition than a zero average effect.
for allmakes ; an average effect of zero with individual effects cancelling does not.The large-sample interval uses a normal critical value. With small arms a
approximation with an appropriate degrees-of-freedom rule is usual, and neither is exact for a finite population. The variance formula is written for complete randomization. Bernoulli assignment, blocking and pairing each induce a different randomization distribution, so this expression does not carry over unchanged.
Worked material
Example
When the unidentifiable term vanishes
When the dropped term is zero. Suppose the treatment adds exactly 5 to every unit's outcome:
When it is not. Suppose the treatment helps half the units by 10 and does nothing for the other half. The average effect is 5, exactly as before, and a study reporting only
Why this is not a defect. Both situations produce the same observable data in expectation. Nothing in the study distinguishes them, so an estimator that assumed the first would understate uncertainty whenever the second held. Erring the other way costs width and never costs coverage.
A zero average effect is not the same as constant effects. If half the units gain 10 and half lose 10, then
Non-example
Four procedures that are not this one
These resemble Neyman inference and answer different questions.
A classical two-sample
A test of the sharp null. Fisher's approach asks whether treatment changed any unit's outcome, and derives its reference distribution by re-randomizing under that null. Neyman estimates an average and quantifies how much it would move. A sharp null can be rejected when the average effect is zero, and an average effect can be non-zero while a randomization test does not reject.
A variance computed as though the arms were independent samples. Treating
A standard error from a design other than complete randomization. Under blocking, pairing, or Bernoulli assignment, the randomization distribution of
Contrast
Exact, conservative, and the difference between them
Exact against conservative.
| Exact design variance | What is reported | |
|---|---|---|
| Formula | ||
| Third term | Present, subtracted | Omitted |
| Computable from data | No | Yes |
| Relationship | — | Never smaller |
| Equal when | — |
Conservative is a direction, not an error. An estimator is conservative when it errs toward claiming less than the evidence supports. This one does: intervals no narrower than warranted, tests rejecting no more often than their stated level. That is a property a reader can rely on.
What it is not. It is not an approximation awaiting a better method, and not a simplification that more data would remove. The missing term needs a quantity no study of any size produces. A learner who reads "conservative" as "imprecise, and fixable" will go looking for a sharper formula and find that the obstacle is the same one the whole subject is built around.
Common errors
Common misconception
The standard error used in a randomized experiment omits a term, so it is an approximation or an error; with better data or a better method the exact variance could be computed and a narrower, more accurate interval reported.
Related units
Requires
Connected
- Blocked and Paired Randomized Experiments (suggested next)