Module 3 of 4 · Lesson 1 of 4
Neyman Repeated-Sampling Inference
A design-based variance and interval, and the unidentifiable term the estimator omits.
What you will be able to do
Given summary data from a completely randomized experiment, the learner can compute the difference in means, compute its conservative variance and standard error, form a confidence interval, and explain which term of the exact variance has been omitted and why that makes the result conservative rather than incorrect.
Orientation
A completely randomized trial has a variance you can write down exactly. One of its three terms involves a quantity no experiment can ever observe: the variation in individual treatment effects, which would require seeing both potential outcomes for the same unit.
So the standard error everyone reports is deliberately too large. Understanding why, and why erring in that direction is the defensible choice rather than a shortcut, is what this unit is for.
Intuition
Variation across the assignments the design permits
Ask what could have come out differently.
Not the people: the sixty in the trial are the sixty, and each carries two fixed potential outcomes written before anyone did anything. Not the outcomes: those are properties of the units. The only thing that could have gone another way is who got treated.
So the spread of the estimate is the spread across the allocations the design permitted. That is a finite, known list, which is why a variance exists at all without assuming the units were sampled from anywhere.
Then the awkward part. The exact variance contains a term measuring how much the treatment effect varies from unit to unit, and computing it needs both potential outcomes for every unit. Every study has one of each. The term is unavailable, not hard to estimate, unavailable, so it is dropped. Because it is subtracted and never negative, dropping it makes the answer bigger. Intervals come out a little wide, and tests a little cautious. That is the trade: certainty about direction obtained by giving up precision.
Definition
Three variances, and which are estimable
The definition above introduces three variances. They look alike and they are not, and the difference decides everything that follows.
Why unbiasedness is a statement about assignments.
What to carry forward. Two of the three quantities have observable counterparts and one does not. That asymmetry is not a gap in the method; it is the identification problem of causal inference appearing in a variance formula, and the derivation that follows turns it into a deliberate, stated choice.
Derivation
Deriving the design variance and its unidentifiable term
Taking the variance of
Where each term comes from. The first two are what a difference of two arm means contributes: each arm's potential-outcome variance, divided by that arm's size.
Why the third is subtracted rather than added. The arms are not two independent samples. They partition one fixed set of
Why the third term cannot be computed.
This is not a small-sample difficulty. Doubling
What is done about it. Replace the first two terms by their observable counterparts and drop the third:
with
The direction of the resulting error. Since
Example
When the unidentifiable term vanishes
When the dropped term is zero. Suppose the treatment adds exactly 5 to every unit's outcome:
When it is not. Suppose the treatment helps half the units by 10 and does nothing for the other half. The average effect is 5, exactly as before, and a study reporting only
Why this is not a defect. Both situations produce the same observable data in expectation. Nothing in the study distinguishes them, so an estimator that assumed the first would understate uncertainty whenever the second held. Erring the other way costs width and never costs coverage.
A zero average effect is not the same as constant effects. If half the units gain 10 and half lose 10, then
Worked example
An interval and what it omits
Problem. A completely randomized trial assigns 50 units to each arm. The observed mean difference is
Report a 95% interval and state what it omits.
Goal. Produce
Relevant principle. The design variance has three terms; two are estimable from one arm each, and the third needs both potential outcomes per unit and is therefore dropped.
Step 1: the conservative variance.
Reason: each arm's sample variance estimates that arm's finite-population variance, divided by that arm's size.
Step 2: the standard error.
Step 3: the interval.
Reason: the large-sample normal approximation; with arms this size it is reasonable, though a
Step 4: name what was left out. The exact variance is
Reason: each unit reveals one potential outcome, so
Result.
Check. Is the interval too narrow anywhere? No: the omitted term is subtracted in the exact formula, so leaving it out can only make the reported variance larger. Whatever
Interpretation. The trial supports a positive average effect. It does not support a claim that every unit benefited. The same
Non-example
Four procedures that are not this one
These resemble Neyman inference and answer different questions.
A classical two-sample
A test of the sharp null. Fisher's approach asks whether treatment changed any unit's outcome, and derives its reference distribution by re-randomizing under that null. Neyman estimates an average and quantifies how much it would move. A sharp null can be rejected when the average effect is zero, and an average effect can be non-zero while a randomization test does not reject.
A variance computed as though the arms were independent samples. Treating
A standard error from a design other than complete randomization. Under blocking, pairing, or Bernoulli assignment, the randomization distribution of
Contrast
Exact, conservative, and the difference between them
Exact against conservative.
| Exact design variance | What is reported | |
|---|---|---|
| Formula | ||
| Third term | Present, subtracted | Omitted |
| Computable from data | No | Yes |
| Relationship | — | Never smaller |
| Equal when | — |
Conservative is a direction, not an error. An estimator is conservative when it errs toward claiming less than the evidence supports. This one does: intervals no narrower than warranted, tests rejecting no more often than their stated level. That is a property a reader can rely on.
What it is not. It is not an approximation awaiting a better method, and not a simplification that more data would remove. The missing term needs a quantity no study of any size produces. A learner who reads "conservative" as "imprecise, and fixable" will go looking for a sharper formula and find that the obstacle is the same one the whole subject is built around.
Exercise
Work in order; the structure thins out as you go.
1: everything supplied. A trial has
(a) Compute
Check:
2: partly supplied. The same trial is re-analysed by a colleague who reports
Is the colleague's interval narrower or wider than yours? Is their assumption checkable from the data? What would you report?
Check: narrower, since assuming
3: unsupplied. A trial reports a 95% interval of
Say what is wrong, what the interval does cover, and what would have to be known to say anything about the spread of individual effects.
Check: the interval is about the average effect, not about individual units. It is compatible with every unit gaining 4, and equally with half gaining 8 and half gaining nothing. Saying anything about the spread requires
What to carry forward
The frame. Potential outcomes fixed, assignment random. The variance of
Unbiasedness.
Exact variance.
What is reported.
Why.
What that supplies. Conservative, not wrong: the reported variance is never too small, so intervals are never too narrow. Exact when every unit's effect is identical.
Constant effects ≠ zero average effect.
The recurring error. Reading an interval for the average effect as a statement about individual units, or reading "conservative" as a flaw to be engineered away.