Neyman Repeated-Sampling Inference

Hold the potential outcomes fixed and let the assignment vary: that is the frame in which a difference in means has a variance at all. The exact design variance contains a term built from both potential outcomes per unit, which no study observes, so the estimator used in practice deliberately drops it. The result is conservative rather than exact, and knowing which it is changes what an interval claims.

Definition

For a completely randomized two-arm experiment with τ ^ = Y ¯ 1 − Y ¯ 0 , define the finite-population variances S w 2 = 1 N − 1 ∑ i ( Y i ( w ) − Y ¯ ( w ) ) 2 for w ∈ { 0 , 1 } and S τ 2 = 1 N − 1 ∑ i ( τ i − τ ) 2 . The exact randomization variance is Var ⁡ ( τ ^ ) = S 1 2 N 1 + S 0 2 N 0 − S τ 2 N . Since S τ 2 is not identified, the estimator used is Var ^ ( τ ^ ) = s 1 2 N 1 + s 0 2 N 0 , with s w 2 the observed within-arm sample variances.

Formal statement

Var ⁡ ( τ ^ ) = S 1 2 N 1 + S 0 2 N 0 − S τ 2 N ; Var ^ ( τ ^ ) = s 1 2 N 1 + s 0 2 N 0 ; S E ( τ ^ ) = Var ^ ( τ ^ ) ; interval τ ^ ± z α / 2 S E ( τ ^ ) .

Assumptions and scope

  • The potential outcomes are treated as fixed constants and the assignment as the only source of randomness. This is a different frame from classical sampling inference, where the units are a random draw from a population; the arithmetic often coincides and the claims do not.

  • S τ 2 is non-negative, so omitting it can only inflate the variance. The estimator is conservative in a precise sense: it does not understate uncertainty, and it is exact when every unit shares the same treatment effect.

  • Constant treatment effects are a much stronger condition than a zero average effect. τ i = τ for all i makes S τ 2 = 0 ; an average effect of zero with individual effects cancelling does not.

  • The large-sample interval uses a normal critical value. With small arms a t approximation with an appropriate degrees-of-freedom rule is usual, and neither is exact for a finite population.

  • The variance formula is written for complete randomization. Bernoulli assignment, blocking and pairing each induce a different randomization distribution, so this expression does not carry over unchanged.

Worked material

Example

When the unidentifiable term vanishes

When the dropped term is zero. Suppose the treatment adds exactly 5 to every unit's outcome: τ i = 5 for all i . Then every τ i equals τ , so S τ 2 = 0 and the conservative estimator is not conservative at all. It targets the exact design variance. Constant effects make the omission free.

When it is not. Suppose the treatment helps half the units by 10 and does nothing for the other half. The average effect is 5, exactly as before, and a study reporting only τ ^ cannot tell the two situations apart. But now τ i varies, S τ 2 > 0 , and the true variance is strictly smaller than what gets reported. The interval is wider than the design warranted.

Why this is not a defect. Both situations produce the same observable data in expectation. Nothing in the study distinguishes them, so an estimator that assumed the first would understate uncertainty whenever the second held. Erring the other way costs width and never costs coverage.

A zero average effect is not the same as constant effects. If half the units gain 10 and half lose 10, then τ = 0 while S τ 2 is large. A study could report a tight interval around zero and be describing a treatment that substantially helps and substantially harms.

Non-example

Four procedures that are not this one

These resemble Neyman inference and answer different questions.

A classical two-sample t -test. The arithmetic can coincide, Welch's standard error is s 1 2 / n 1 + s 0 2 / n 0 , the same expression, but the frame is different. There the units are a random sample from a population and the inference is about a population parameter. Here the units are fixed and the randomness is the assignment. Getting the same number from two different arguments is a coincidence, not evidence that the arguments are the same.

A test of the sharp null. Fisher's approach asks whether treatment changed any unit's outcome, and derives its reference distribution by re-randomizing under that null. Neyman estimates an average and quantifies how much it would move. A sharp null can be rejected when the average effect is zero, and an average effect can be non-zero while a randomization test does not reject.

A variance computed as though the arms were independent samples. Treating τ ^ as a difference of two independent sample means omits the finite-population correction implicit in the exact formula. The two arms are a partition of one fixed set, not two draws.

A standard error from a design other than complete randomization. Under blocking, pairing, or Bernoulli assignment, the randomization distribution of τ ^ differs, and so does its variance. The expression here is written for one mechanism, and carrying it to another is the same error as analysing a paired trial as if independently assigned.

Contrast

Exact, conservative, and the difference between them

Exact against conservative.

Exact design varianceWhat is reported
Formula S 1 2 N 1 + S 0 2 N 0 − S τ 2 N s 1 2 N 1 + s 0 2 N 0
Third termPresent, subtractedOmitted
Computable from dataNoYes
Relationship—Never smaller
Equal when— τ i constant across units

Conservative is a direction, not an error. An estimator is conservative when it errs toward claiming less than the evidence supports. This one does: intervals no narrower than warranted, tests rejecting no more often than their stated level. That is a property a reader can rely on.

What it is not. It is not an approximation awaiting a better method, and not a simplification that more data would remove. The missing term needs a quantity no study of any size produces. A learner who reads "conservative" as "imprecise, and fixable" will go looking for a sharper formula and find that the obstacle is the same one the whole subject is built around.

Common errors

Common misconception

The standard error used in a randomized experiment omits a term, so it is an approximation or an error; with better data or a better method the exact variance could be computed and a narrower, more accurate interval reported.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.