Neyman Repeated-Sampling Inference

What you will be able to do

Given summary data from a completely randomized experiment, the learner can compute the difference in means, compute its conservative variance and standard error, form a confidence interval, and explain which term of the exact variance has been omitted and why that makes the result conservative rather than incorrect.

Orientation

A completely randomized trial has a variance you can write down exactly. One of its three terms involves a quantity no experiment can ever observe: the variation in individual treatment effects, which would require seeing both potential outcomes for the same unit.

So the standard error everyone reports is deliberately too large. Understanding why, and why erring in that direction is the defensible choice rather than a shortcut, is what this unit is for.

Intuition

Variation across the assignments the design permits

Ask what could have come out differently.

Not the people: the sixty in the trial are the sixty, and each carries two fixed potential outcomes written before anyone did anything. Not the outcomes: those are properties of the units. The only thing that could have gone another way is who got treated.

So the spread of the estimate is the spread across the allocations the design permitted. That is a finite, known list, which is why a variance exists at all without assuming the units were sampled from anywhere.

Then the awkward part. The exact variance contains a term measuring how much the treatment effect varies from unit to unit, and computing it needs both potential outcomes for every unit. Every study has one of each. The term is unavailable, not hard to estimate, unavailable, so it is dropped. Because it is subtracted and never negative, dropping it makes the answer bigger. Intervals come out a little wide, and tests a little cautious. That is the trade: certainty about direction obtained by giving up precision.

Definition

Three variances, and which are estimable

The definition above introduces three variances. They look alike and they are not, and the difference decides everything that follows.

S 1 2 and S 0 2 are ordinary. Each is the variance of one potential outcome across the units. The spread of Y i ( 1 ) over everybody, and of Y i ( 0 ) over everybody. No unit reveals both, but each arm reveals one of them for the units assigned to it, so each has an observable counterpart: the within-arm sample variance s w 2 .

S τ 2 is not. It is the variance of the individual effects τ i = Y i ( 1 ) − Y i ( 0 ) , and every single summand needs one unit's two potential outcomes at once. No arm supplies that, no arrangement of arms supplies it, and no sample size supplies it. The quantity is well defined and permanently invisible.

Why unbiasedness is a statement about assignments. E [ τ ^ ] = τ averages over the allocations the design permitted, holding the potential outcomes fixed. It does not say the trial you ran landed near τ ; it says the procedure has no systematic tilt across the randomizations that could have happened. The variance below is the spread of the same distribution, over assignments, not over samples of people.

What to carry forward. Two of the three quantities have observable counterparts and one does not. That asymmetry is not a gap in the method; it is the identification problem of causal inference appearing in a variance formula, and the derivation that follows turns it into a deliberate, stated choice.

Derivation

Deriving the design variance and its unidentifiable term

Taking the variance of τ ^ over the assignments the design permits gives

Var ⁡ ( τ ^ ) = S 1 2 N 1 + S 0 2 N 0 − S τ 2 N .

Where each term comes from. The first two are what a difference of two arm means contributes: each arm's potential-outcome variance, divided by that arm's size.

Why the third is subtracted rather than added. The arms are not two independent samples. They partition one fixed set of N units, so a unit landing in the treated arm is a unit not in the control arm. That negative dependence removes variance, and the amount it removes is governed by how much the individual effects τ i vary. When every unit responds identically the partition costs nothing, and the term vanishes.

Why the third term cannot be computed. S τ 2 is a variance of the quantities τ i = Y i ( 1 ) − Y i ( 0 ) . Each one needs both of a unit's potential outcomes. A study observes exactly one per unit, the other is the outcome under the assignment that did not happen, so no summand is available, for any i , in any experiment.

This is not a small-sample difficulty. Doubling N supplies twice as many units and still one potential outcome each. The term is not identified: unavailable in principle, rather than hard to estimate.

What is done about it. Replace the first two terms by their observable counterparts and drop the third:

Var ^ ( τ ^ ) = s 1 2 N 1 + s 0 2 N 0 , S E ( τ ^ ) = s 1 2 N 1 + s 0 2 N 0 ,

with s w 2 the observed within-arm sample variances, and a large-sample interval τ ^ ± z α / 2 S E ( τ ^ ) .

The direction of the resulting error. Since S τ 2 ≥ 0 , the dropped term was being subtracted, so omitting it can only make the reported variance larger than the true one. The standard error is never too small and the interval never too narrow. The estimator is conservative, and exactly correct when every unit's effect is the same.

Example

When the unidentifiable term vanishes

When the dropped term is zero. Suppose the treatment adds exactly 5 to every unit's outcome: τ i = 5 for all i . Then every τ i equals τ , so S τ 2 = 0 and the conservative estimator is not conservative at all. It targets the exact design variance. Constant effects make the omission free.

When it is not. Suppose the treatment helps half the units by 10 and does nothing for the other half. The average effect is 5, exactly as before, and a study reporting only τ ^ cannot tell the two situations apart. But now τ i varies, S τ 2 > 0 , and the true variance is strictly smaller than what gets reported. The interval is wider than the design warranted.

Why this is not a defect. Both situations produce the same observable data in expectation. Nothing in the study distinguishes them, so an estimator that assumed the first would understate uncertainty whenever the second held. Erring the other way costs width and never costs coverage.

A zero average effect is not the same as constant effects. If half the units gain 10 and half lose 10, then τ = 0 while S τ 2 is large. A study could report a tight interval around zero and be describing a treatment that substantially helps and substantially harms.

Worked example

An interval and what it omits

Problem. A completely randomized trial assigns 50 units to each arm. The observed mean difference is τ ^ = 4.8 , with within-arm sample variances s 1 2 = 100 and s 0 2 = 144 .

Report a 95% interval and state what it omits.

Goal. Produce S E ( τ ^ ) , the interval, and an account of the exact variance.

Relevant principle. The design variance has three terms; two are estimable from one arm each, and the third needs both potential outcomes per unit and is therefore dropped.

Step 1: the conservative variance.

Var ^ ( τ ^ ) = s 1 2 N 1 + s 0 2 N 0 = 100 50 + 144 50 = 2 + 2.88 = 4.88 .

Reason: each arm's sample variance estimates that arm's finite-population variance, divided by that arm's size.

Step 2: the standard error.

S E ( τ ^ ) = 4.88 ≈ 2.21 .

Step 3: the interval.

4.8 ± 1.96 × 2.21 = ( 0.47 , 9.13 ) .

Reason: the large-sample normal approximation; with arms this size it is reasonable, though a t critical value is common in practice.

Step 4: name what was left out. The exact variance is S 1 2 50 + S 0 2 50 − S τ 2 100 . The final term measures variation in the individual treatment effects, and computing it requires Y i ( 1 ) and Y i ( 0 ) for the same unit.
Reason: each unit reveals one potential outcome, so τ i is never observed for any unit and its variance cannot be formed.

Result. τ ^ = 4.8 , S E ≈ 2.21 , 95% interval ( 0.47 , 9.13 ) , which excludes zero.

Check. Is the interval too narrow anywhere? No: the omitted term is subtracted in the exact formula, so leaving it out can only make the reported variance larger. Whatever S τ 2 turns out to be, the true interval is no wider than this one.

Interpretation. The trial supports a positive average effect. It does not support a claim that every unit benefited. The same τ ^ and the same interval would arise if the treatment helped some units and harmed others, and the very term that would reveal that is the one no study can compute.

Non-example

Four procedures that are not this one

These resemble Neyman inference and answer different questions.

A classical two-sample t -test. The arithmetic can coincide, Welch's standard error is s 1 2 / n 1 + s 0 2 / n 0 , the same expression, but the frame is different. There the units are a random sample from a population and the inference is about a population parameter. Here the units are fixed and the randomness is the assignment. Getting the same number from two different arguments is a coincidence, not evidence that the arguments are the same.

A test of the sharp null. Fisher's approach asks whether treatment changed any unit's outcome, and derives its reference distribution by re-randomizing under that null. Neyman estimates an average and quantifies how much it would move. A sharp null can be rejected when the average effect is zero, and an average effect can be non-zero while a randomization test does not reject.

A variance computed as though the arms were independent samples. Treating τ ^ as a difference of two independent sample means omits the finite-population correction implicit in the exact formula. The two arms are a partition of one fixed set, not two draws.

A standard error from a design other than complete randomization. Under blocking, pairing, or Bernoulli assignment, the randomization distribution of τ ^ differs, and so does its variance. The expression here is written for one mechanism, and carrying it to another is the same error as analysing a paired trial as if independently assigned.

Contrast

Exact, conservative, and the difference between them

Exact against conservative.

Exact design varianceWhat is reported
Formula S 1 2 N 1 + S 0 2 N 0 − S τ 2 N s 1 2 N 1 + s 0 2 N 0
Third termPresent, subtractedOmitted
Computable from dataNoYes
Relationship—Never smaller
Equal when— τ i constant across units

Conservative is a direction, not an error. An estimator is conservative when it errs toward claiming less than the evidence supports. This one does: intervals no narrower than warranted, tests rejecting no more often than their stated level. That is a property a reader can rely on.

What it is not. It is not an approximation awaiting a better method, and not a simplification that more data would remove. The missing term needs a quantity no study of any size produces. A learner who reads "conservative" as "imprecise, and fixable" will go looking for a sharper formula and find that the obstacle is the same one the whole subject is built around.

Exercise

Work in order; the structure thins out as you go.

1: everything supplied. A trial has N 1 = 40 , N 0 = 60 , τ ^ = 3.2 , s 1 2 = 64 , s 0 2 = 90 .

(a) Compute Var ^ ( τ ^ ) . (b) Compute S E . (c) Form a 95% interval. (d) Write the exact variance symbolically and circle the term you did not use.

Check: 64 / 40 + 90 / 60 = 1.6 + 1.5 = 3.1 ; S E ≈ 1.76 ; 3.2 ± 1.96 ( 1.76 ) = ( − 0.25 , 6.65 ) ; omitted term S τ 2 / 100 .

2: partly supplied. The same trial is re-analysed by a colleague who reports S E = 1.31 , having assumed the treatment effect is the same for every unit and estimated S τ 2 as zero.

Is the colleague's interval narrower or wider than yours? Is their assumption checkable from the data? What would you report?

Check: narrower, since assuming S τ 2 = 0 removes nothing but licenses a smaller variance only if the assumption holds. It is not checkable: constant effects concern both potential outcomes per unit. Report the conservative interval and state the assumption a narrower one would require.

3: unsupplied. A trial reports a 95% interval of ( 1.1 , 6.9 ) for an average effect, and a reader concludes that most units gained between 1.1 and 6.9.

Say what is wrong, what the interval does cover, and what would have to be known to say anything about the spread of individual effects.

Check: the interval is about the average effect, not about individual units. It is compatible with every unit gaining 4, and equally with half gaining 8 and half gaining nothing. Saying anything about the spread requires S τ 2 , which needs both potential outcomes per unit and is exactly what the design cannot supply.

What to carry forward

The frame. Potential outcomes fixed, assignment random. The variance of τ ^ is its spread across the allocations the design permitted.

Unbiasedness. E [ τ ^ ] = τ over assignments. A statement about averaging over allocations, not about the one that happened.

Exact variance. S 1 2 N 1 + S 0 2 N 0 − S τ 2 N .

What is reported. Var ^ ( τ ^ ) = s 1 2 N 1 + s 0 2 N 0 , and S E = Var ^ .

Why. S τ 2 is a variance of individual effects; each needs both potential outcomes for one unit, and no study observes both.

What that supplies. Conservative, not wrong: the reported variance is never too small, so intervals are never too narrow. Exact when every unit's effect is identical.

Constant effects ≠ zero average effect. τ i = τ for all i makes S τ 2 = 0 . Effects of + 10 and − 10 cancelling do not.

The recurring error. Reading an interval for the average effect as a statement about individual units, or reading "conservative" as a flaw to be engineered away.

Next step

Practice Neyman Repeated-Sampling Inference

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.