Fisher Randomization Inference
Under the hypothesis that treatment changed nothing for anyone, every missing potential outcome is known: it equals the observed one. That makes the whole experiment recomputable under every allocation the design permitted, and the observed statistic can be placed in the list of values it could have taken. The result is an exact p-value that assumes no model, no distribution and no large sample, only the assignment mechanism.
Definition
The sharp null hypothesis is
Formal statement
Assumptions and scope
The sharp null is stronger than the hypothesis of zero average effect. It asserts no effect for any unit; an average effect of zero permits individual effects that cancel, and does not fill in the missing potential outcomes.
The enumeration must be over the allocations the actual mechanism permits. A test that enumerates all splits of the units when the design was blocked, paired, or Bernoulli uses the wrong reference set and its p-value does not mean what it claims.
The statistic must be chosen before the treatment effects are inspected. Selecting a statistic that happens to make the observed data extreme converts an exact test into a search.
The Monte Carlo estimate adds one to numerator and denominator, which includes the observed assignment among the draws. This keeps the estimate from reporting zero, which no finite simulation can justify.
Rejecting the sharp null establishes that treatment affected somebody, not that it affected everyone or that the average effect is large. Failing to reject is not evidence that the treatment does nothing.
Forms this is expressed in
The same content in several forms. Each makes something visible that the others leave implicit, so moving between them is part of understanding the topic rather than a presentation choice.
tabular
The randomization distribution for a four-unit experiment with two treated. Observed outcomes are 8, 6, 5, 3 for units A, B, C, D; the realised assignment treated A and B, giving
Under the sharp null each unit's outcome is fixed at what was observed, so every allocation is recomputed from the same four numbers.
| Treated pair | ||||
|---|---|---|---|---|
| A, B | 7.0 | 4.0 | 3.0 | yes |
| A, C | 6.5 | 4.5 | 2.0 | no |
| A, D | 5.5 | 5.5 | 0.0 | no |
| B, C | 5.5 | 5.5 | 0.0 | no |
| B, D | 4.5 | 6.5 | −2.0 | no |
| C, D | 4.0 | 7.0 | −3.0 | yes |
Two of the six equally likely allocations are at least as extreme as the observed one, so the two-sided randomization p-value is
Worked material
Example
Fisher's tea taster
A participant claims she can tell whether milk or tea went into the cup first. She is given eight cups, told that exactly four are milk-first, and asked to identify which four.
Under the sharp null she has no ability, so which cups she picks has nothing to do with how they were made. There are
sets of four she could name, all equally likely under the null. Exactly one is entirely correct.
If she identifies all four, the one-sided p-value is
Three features of the design matter here. The number 70 comes from the design, eight cups, four of each, and she was told the split, not from any assumption about her or about tea. The test is exact rather than approximate. And the design fixes the best attainable evidence in advance: had she been given six cups with three of each, the smallest possible p-value would be
Non-example
Five procedures that are not this test
These resemble a randomization test and are not one.
Permuting when the design did not permute freely. A paired experiment randomizes one unit within each of ten pairs, giving
Choosing the statistic after seeing the effects. Trying the difference in means, then a rank statistic, then a trimmed mean, and reporting whichever is most extreme. Each test is individually exact; selecting among them is a search, and the reported p-value no longer has its stated meaning.
A bootstrap. Resampling units with replacement models sampling from a population. The randomization test resamples assignments from a known mechanism with the units held fixed. Different source of randomness, different claim.
Concluding no effect from a large p-value. A non-significant randomization test says the data are consistent with no effect for anyone. With four units, as above, that is guaranteed regardless of the truth.
Reporting
Contrast
Fisher and Neyman answer different questions
Fisher against Neyman. Both take their randomness from the assignment mechanism. Almost everything else differs.
| Fisher | Neyman | |
|---|---|---|
| What is nulled or targeted | The average effect | |
| Missing potential outcomes | Filled in exactly, under the null | Left missing; their variance contribution bounded |
| Output | A p-value | An estimate, standard error, interval |
| Exactness | Exact for any | Large-sample approximation |
| Assumes a distribution | No | Normal approximation for the interval |
| Rejection means | Treatment affected at least one unit | — |
Why one does not determine the other. The sharp null is strictly stronger than a zero average effect. A treatment that helps half the units by 10 and harms the other half by 10 has
The converse also occurs: with few units, a real and uniform effect can leave the randomization test unable to reject anything, because the design admits too few allocations to place the observed value in a tail.
The practical reading. A randomization p-value answers "did this treatment do anything to anybody?". An interval answers "how big is the average effect, and how sure are we?". Reporting one as though it answered the other is the error this contrast exists to prevent.
Common errors
Common misconception
Rejecting the sharp null shows the average treatment effect is non-zero, and failing to reject it shows the treatment has no effect; the randomization test and a confidence interval for the average effect are two ways of asking the same question.
Related units
Requires
Connected
- Neyman Repeated-Sampling Inference (contrasts with)