Module 3 of 4 · Lesson 2 of 4
Fisher Randomization Inference
A randomization test built by relabelling assignments, and what rejecting the sharp null establishes.
What you will be able to do
Given a small experiment and its assignment mechanism, the learner can state the sharp null, enumerate or sample the permitted allocations, compute the randomization p-value, and explain what rejecting or failing to reject establishes, distinguishing it from a claim about the average effect.
Orientation
With eight cups of tea and no distributional assumptions at all, you can still compute exactly how surprising a result is, by asking what else the randomisation could have produced.
The test is exact. Its reference distribution comes from the experiment's own design, which is knowable because somebody chose it.
Intuition
The randomization distribution of the statistic
Every test needs an answer to: what values could this statistic have taken if nothing were going on?
Ordinarily that answer comes from an assumption, outcomes are normal, the sample is large, a
Suppose the treatment changed nobody's outcome at all. Then the number each unit produced is the number it would have produced either way. The outcome column is nailed down. The only thing that could have been different is the labels, who was called treated.
So take the same outcomes, relabel them according to another allocation the design allowed, and recompute the statistic. That is a value it could have produced. Do it for all of them and you have the complete list, with no approximation anywhere.
Then look at where the observed value sits. Near the middle: the data look like what no-effect produces. Out in the tail: either something unlikely happened, or treatment did something to somebody.
Definition
The sharp null and the randomization distribution
The canonical statement defines the sharp null, the statistic and the p-value. Three properties follow that decide what the test can be asked for.
Sharpness is what makes the experiment recomputable.
The reference distribution comes from the design, not from a model. No normality is assumed and no large sample is needed: the randomness being integrated over is the assignment mechanism, which is known exactly because somebody chose it. A test valid in a study of twelve units is valid for that reason.
The statistic must be chosen before the effects are seen. Any
Example
Fisher's tea taster
A participant claims she can tell whether milk or tea went into the cup first. She is given eight cups, told that exactly four are milk-first, and asked to identify which four.
Under the sharp null she has no ability, so which cups she picks has nothing to do with how they were made. There are
sets of four she could name, all equally likely under the null. Exactly one is entirely correct.
If she identifies all four, the one-sided p-value is
Three features of the design matter here. The number 70 comes from the design, eight cups, four of each, and she was told the split, not from any assumption about her or about tea. The test is exact rather than approximate. And the design fixes the best attainable evidence in advance: had she been given six cups with three of each, the smallest possible p-value would be
Worked example
Four patients, six allocations
Problem. Four patients, two randomized to treatment by complete randomization. Observed outcomes: A = 8, B = 6, C = 5, D = 3. The realised assignment treated A and B.
Test the sharp null with
Goal. A p-value derived from the design, and a statement of what it establishes.
Relevant principle. Under the sharp null every outcome is fixed at its observed value, so each permitted allocation is a relabelling of the same four numbers.
Step 1: the observed statistic.
Treated A, B:
Step 2: enumerate the permitted allocations.
Complete randomization with
Reason: the mechanism fixes the treated count, so only the six two-element subsets are possible.
Step 3: recompute
| Treated | |||
|---|---|---|---|
| A, B | 7.0 | 4.0 | 3.0 |
| A, C | 6.5 | 4.5 | 2.0 |
| A, D | 5.5 | 5.5 | 0.0 |
| B, C | 5.5 | 5.5 | 0.0 |
| B, D | 4.5 | 6.5 | −2.0 |
| C, D | 4.0 | 7.0 | −3.0 |
Reason: the sharp null says A scores 8 whether treated or not, so the outcome column never changes, only who is labelled treated.
Step 4: count the extreme allocations.
Result. The randomization p-value is about 0.33. The sharp null is not rejected.
Check. Could this design have produced a small p-value at all? The most extreme possible outcome is one tail of six allocations, giving a two-sided minimum of
Interpretation. Report that the experiment is too small to detect anything at conventional levels, not that the treatment had no effect. The observed difference of 3 is the largest the design permits, and it still yields the largest possible p-value. A fact about
Non-example
Five procedures that are not this test
These resemble a randomization test and are not one.
Permuting when the design did not permute freely. A paired experiment randomizes one unit within each of ten pairs, giving
Choosing the statistic after seeing the effects. Trying the difference in means, then a rank statistic, then a trimmed mean, and reporting whichever is most extreme. Each test is individually exact; selecting among them is a search, and the reported p-value no longer has its stated meaning.
A bootstrap. Resampling units with replacement models sampling from a population. The randomization test resamples assignments from a known mechanism with the units held fixed. Different source of randomness, different claim.
Concluding no effect from a large p-value. A non-significant randomization test says the data are consistent with no effect for anyone. With four units, as above, that is guaranteed regardless of the truth.
Reporting
Contrast
Fisher and Neyman answer different questions
Fisher against Neyman. Both take their randomness from the assignment mechanism. Almost everything else differs.
| Fisher | Neyman | |
|---|---|---|
| What is nulled or targeted | The average effect | |
| Missing potential outcomes | Filled in exactly, under the null | Left missing; their variance contribution bounded |
| Output | A p-value | An estimate, standard error, interval |
| Exactness | Exact for any | Large-sample approximation |
| Assumes a distribution | No | Normal approximation for the interval |
| Rejection means | Treatment affected at least one unit | — |
Why one does not determine the other. The sharp null is strictly stronger than a zero average effect. A treatment that helps half the units by 10 and harms the other half by 10 has
The converse also occurs: with few units, a real and uniform effect can leave the randomization test unable to reject anything, because the design admits too few allocations to place the observed value in a tail.
The practical reading. A randomization p-value answers "did this treatment do anything to anybody?". An interval answers "how big is the average effect, and how sure are we?". Reporting one as though it answered the other is the error this contrast exists to prevent.
Exercise
Work in order.
1: fully structured. Six units, three treated by complete randomization. Outcomes: 12, 10, 9 (treated) and 7, 6, 4 (control).
(a) Compute
Check:
2: partly structured. A trial of 40 units reports a randomization p-value of 0.002, computed by drawing 5,000 allocations, with 9 at least as extreme as observed.
(a) Verify the reported figure. (b) Would reporting
Check:
3: unstructured. A researcher runs a matched-pair experiment with 12 pairs, then tests the sharp null by enumerating every way of splitting the 24 units into two groups of 12 and computing the difference in means for each.
Explain what is wrong, what should have been enumerated, and whether the resulting p-value is likely too large or too small.
Check: the design permitted
What to carry forward
The null.
Why that matters. With outcomes fixed, each permitted allocation is a relabelling, so the statistic can be recomputed exactly.
The p-value.
By simulation.
Exactness. No distributional assumption, no large-sample argument. Valid for four units.
What rejection means. Treatment affected at least one unit. Not that the average effect is non-zero, and not that it affected everyone.
What non-rejection means. The data are consistent with no effect for anyone, which a small design guarantees regardless of the truth. Check the smallest attainable p-value before interpreting.
Against Neyman. Same randomness, different question: "did anything happen?" against "how big is the average, and how sure are we?".
The recurring errors. Enumerating allocations the design did not permit; choosing the statistic after seeing the effects; reading a large p-value as evidence of no effect.