Fisher Randomization Inference

Under the hypothesis that treatment changed nothing for anyone, every missing potential outcome is known: it equals the observed one. That makes the whole experiment recomputable under every allocation the design permitted, and the observed statistic can be placed in the list of values it could have taken. The result is an exact p-value that assumes no model, no distribution and no large sample, only the assignment mechanism.

Definition

The sharp null hypothesis is H 0 : Y i ( 1 ) = Y i ( 0 ) for every unit i . It is sharp because it determines every missing potential outcome: under it, each unit's outcome is what was observed regardless of assignment. Choosing a statistic T ( W , Y ) , commonly Y ¯ 1 − Y ¯ 0 , the randomization distribution is the set of values T takes as W ranges over the allocations the mechanism permits, with the observed outcomes held fixed. The two-sided randomization p-value is p = P ( | T ( W , Y ) | ≥ | T obs | ) , the probability taken over the assignment mechanism.

Formal statement

H 0 : Y i ( 1 ) = Y i ( 0 )   ∀ i ; p = P ( | T ( W , Y ) | ≥ | T obs | ∣ H 0 ) ; by Monte Carlo over B draws, p ^ = 1 + # { b : | T b | ≥ | T obs | } B + 1 .

Assumptions and scope

  • The sharp null is stronger than the hypothesis of zero average effect. It asserts no effect for any unit; an average effect of zero permits individual effects that cancel, and does not fill in the missing potential outcomes.

  • The enumeration must be over the allocations the actual mechanism permits. A test that enumerates all splits of the units when the design was blocked, paired, or Bernoulli uses the wrong reference set and its p-value does not mean what it claims.

  • The statistic must be chosen before the treatment effects are inspected. Selecting a statistic that happens to make the observed data extreme converts an exact test into a search.

  • The Monte Carlo estimate adds one to numerator and denominator, which includes the observed assignment among the draws. This keeps the estimate from reporting zero, which no finite simulation can justify.

  • Rejecting the sharp null establishes that treatment affected somebody, not that it affected everyone or that the average effect is large. Failing to reject is not evidence that the treatment does nothing.

Forms this is expressed in

The same content in several forms. Each makes something visible that the others leave implicit, so moving between them is part of understanding the topic rather than a presentation choice.

tabular

The randomization distribution for a four-unit experiment with two treated. Observed outcomes are 8, 6, 5, 3 for units A, B, C, D; the realised assignment treated A and B, giving T obs = 7 − 4 = 3 .

Under the sharp null each unit's outcome is fixed at what was observed, so every allocation is recomputed from the same four numbers.

Treated pair Y ¯ 1 Y ¯ 0 T | T | ≥ 3
A, B7.04.03.0yes
A, C6.54.52.0no
A, D5.55.50.0no
B, C5.55.50.0no
B, D4.56.5−2.0no
C, D4.07.0−3.0yes

Two of the six equally likely allocations are at least as extreme as the observed one, so the two-sided randomization p-value is 2 / 6 ≈ 0.33 . With four units nothing can reach conventional significance: the smallest attainable two-sided p-value is 2 / 6 , which is a property of the design rather than of the data.

Worked material

Example

Fisher's tea taster

A participant claims she can tell whether milk or tea went into the cup first. She is given eight cups, told that exactly four are milk-first, and asked to identify which four.

Under the sharp null she has no ability, so which cups she picks has nothing to do with how they were made. There are

( 8 4 ) = 70

sets of four she could name, all equally likely under the null. Exactly one is entirely correct.

If she identifies all four, the one-sided p-value is 1 / 70 ≈ 0.014 .

Three features of the design matter here. The number 70 comes from the design, eight cups, four of each, and she was told the split, not from any assumption about her or about tea. The test is exact rather than approximate. And the design fixes the best attainable evidence in advance: had she been given six cups with three of each, the smallest possible p-value would be 1 / 20 = 0.05 , and no performance however perfect could do better.

Non-example

Five procedures that are not this test

These resemble a randomization test and are not one.

Permuting when the design did not permute freely. A paired experiment randomizes one unit within each of ten pairs, giving 2 10 = 1024 allocations. Enumerating all ( 20 10 ) = 184,756 ways to split twenty units uses a reference set the design never allowed, and the resulting p-value describes an experiment that was not run.

Choosing the statistic after seeing the effects. Trying the difference in means, then a rank statistic, then a trimmed mean, and reporting whichever is most extreme. Each test is individually exact; selecting among them is a search, and the reported p-value no longer has its stated meaning.

A bootstrap. Resampling units with replacement models sampling from a population. The randomization test resamples assignments from a known mechanism with the units held fixed. Different source of randomness, different claim.

Concluding no effect from a large p-value. A non-significant randomization test says the data are consistent with no effect for anyone. With four units, as above, that is guaranteed regardless of the truth.

Reporting p = 0 from a simulation. Drawing 10,000 allocations and finding none as extreme does not establish p = 0 ; it establishes p is small. The ( 1 + ⋅ ) / ( B + 1 ) form exists precisely to avoid claiming otherwise.

Contrast

Fisher and Neyman answer different questions

Fisher against Neyman. Both take their randomness from the assignment mechanism. Almost everything else differs.

FisherNeyman
What is nulled or targeted Y i ( 1 ) = Y i ( 0 ) for every unitThe average effect τ
Missing potential outcomesFilled in exactly, under the nullLeft missing; their variance contribution bounded
OutputA p-valueAn estimate, standard error, interval
ExactnessExact for any N Large-sample approximation
Assumes a distributionNoNormal approximation for the interval
Rejection meansTreatment affected at least one unit—

Why one does not determine the other. The sharp null is strictly stronger than a zero average effect. A treatment that helps half the units by 10 and harms the other half by 10 has τ = 0 , so Neyman's interval should cover zero, while the sharp null is flatly false, and a well-chosen statistic may reject it.

The converse also occurs: with few units, a real and uniform effect can leave the randomization test unable to reject anything, because the design admits too few allocations to place the observed value in a tail.

The practical reading. A randomization p-value answers "did this treatment do anything to anybody?". An interval answers "how big is the average effect, and how sure are we?". Reporting one as though it answered the other is the error this contrast exists to prevent.

Common errors

Common misconception

Rejecting the sharp null shows the average treatment effect is non-zero, and failing to reject it shows the treatment has no effect; the randomization test and a confidence interval for the average effect are two ways of asking the same question.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.