Module 3 of 4 · Lesson 2 of 4

Fisher Randomization Inference

A randomization test built by relabelling assignments, and what rejecting the sharp null establishes.

What you will be able to do

Given a small experiment and its assignment mechanism, the learner can state the sharp null, enumerate or sample the permitted allocations, compute the randomization p-value, and explain what rejecting or failing to reject establishes, distinguishing it from a claim about the average effect.

Orientation

With eight cups of tea and no distributional assumptions at all, you can still compute exactly how surprising a result is, by asking what else the randomisation could have produced.

The test is exact. Its reference distribution comes from the experiment's own design, which is knowable because somebody chose it.

Intuition

The randomization distribution of the statistic

Every test needs an answer to: what values could this statistic have taken if nothing were going on?

Ordinarily that answer comes from an assumption, outcomes are normal, the sample is large, a t distribution applies. Fisher's answer comes from the randomization itself.

Suppose the treatment changed nobody's outcome at all. Then the number each unit produced is the number it would have produced either way. The outcome column is nailed down. The only thing that could have been different is the labels, who was called treated.

So take the same outcomes, relabel them according to another allocation the design allowed, and recompute the statistic. That is a value it could have produced. Do it for all of them and you have the complete list, with no approximation anywhere.

Then look at where the observed value sits. Near the middle: the data look like what no-effect produces. Out in the tail: either something unlikely happened, or treatment did something to somebody.

Definition

The sharp null and the randomization distribution

The canonical statement defines the sharp null, the statistic and the p-value. Three properties follow that decide what the test can be asked for.

Sharpness is what makes the experiment recomputable. H 0 : Y i ( 1 ) = Y i ( 0 ) fixes every missing potential outcome, so for any reallocation W the whole outcome vector is known and T can be evaluated exactly. Neyman's null, that the average effect is zero, fixes nothing unit by unit and admits no such recomputation. The two nulls are different hypotheses, and rejecting the sharp one does not establish that the average effect is nonzero.

The reference distribution comes from the design, not from a model. No normality is assumed and no large sample is needed: the randomness being integrated over is the assignment mechanism, which is known exactly because somebody chose it. A test valid in a study of twelve units is valid for that reason.

The statistic must be chosen before the effects are seen. Any T gives a valid p-value under the sharp null, so the choice cannot make the test invalid; it decides what the test is powerful against. Selecting T after inspecting the outcomes converts that freedom into a way of manufacturing significance, which is why the choice is recorded in advance.

Example

Fisher's tea taster

A participant claims she can tell whether milk or tea went into the cup first. She is given eight cups, told that exactly four are milk-first, and asked to identify which four.

Under the sharp null she has no ability, so which cups she picks has nothing to do with how they were made. There are

( 8 4 ) = 70

sets of four she could name, all equally likely under the null. Exactly one is entirely correct.

If she identifies all four, the one-sided p-value is 1 / 70 ≈ 0.014 .

Three features of the design matter here. The number 70 comes from the design, eight cups, four of each, and she was told the split, not from any assumption about her or about tea. The test is exact rather than approximate. And the design fixes the best attainable evidence in advance: had she been given six cups with three of each, the smallest possible p-value would be 1 / 20 = 0.05 , and no performance however perfect could do better.

Worked example

Four patients, six allocations

Problem. Four patients, two randomized to treatment by complete randomization. Observed outcomes: A = 8, B = 6, C = 5, D = 3. The realised assignment treated A and B.

Test the sharp null with T = Y ¯ 1 − Y ¯ 0 , two-sided.

Goal. A p-value derived from the design, and a statement of what it establishes.

Relevant principle. Under the sharp null every outcome is fixed at its observed value, so each permitted allocation is a relabelling of the same four numbers.

Step 1: the observed statistic.
Treated A, B: Y ¯ 1 = ( 8 + 6 ) / 2 = 7 . Control C, D: Y ¯ 0 = ( 5 + 3 ) / 2 = 4 .

T obs = 7 − 4 = 3.

Step 2: enumerate the permitted allocations.
Complete randomization with N = 4 , N 1 = 2 permits ( 4 2 ) = 6 allocations, each with probability 1 / 6 .
Reason: the mechanism fixes the treated count, so only the six two-element subsets are possible.

Step 3: recompute T for each, holding outcomes fixed.

Treated Y ¯ 1 Y ¯ 0 T
A, B7.04.03.0
A, C6.54.52.0
A, D5.55.50.0
B, C5.55.50.0
B, D4.56.5−2.0
C, D4.07.0−3.0

Reason: the sharp null says A scores 8 whether treated or not, so the outcome column never changes, only who is labelled treated.

Step 4: count the extreme allocations.
| T | ≥ 3 holds for { A , B } and { C , D } : two of six.

p = 2 6 ≈ 0.33 .

Result. The randomization p-value is about 0.33. The sharp null is not rejected.

Check. Could this design have produced a small p-value at all? The most extreme possible outcome is one tail of six allocations, giving a two-sided minimum of 2 / 6 ≈ 0.33 . So no result from four units could have reached 0.05. The non-rejection reflects the design's resolution, not the treatment.

Interpretation. Report that the experiment is too small to detect anything at conventional levels, not that the treatment had no effect. The observed difference of 3 is the largest the design permits, and it still yields the largest possible p-value. A fact about ( 4 2 ) , knowable before any patient was recruited.

Non-example

Five procedures that are not this test

These resemble a randomization test and are not one.

Permuting when the design did not permute freely. A paired experiment randomizes one unit within each of ten pairs, giving 2 10 = 1024 allocations. Enumerating all ( 20 10 ) = 184,756 ways to split twenty units uses a reference set the design never allowed, and the resulting p-value describes an experiment that was not run.

Choosing the statistic after seeing the effects. Trying the difference in means, then a rank statistic, then a trimmed mean, and reporting whichever is most extreme. Each test is individually exact; selecting among them is a search, and the reported p-value no longer has its stated meaning.

A bootstrap. Resampling units with replacement models sampling from a population. The randomization test resamples assignments from a known mechanism with the units held fixed. Different source of randomness, different claim.

Concluding no effect from a large p-value. A non-significant randomization test says the data are consistent with no effect for anyone. With four units, as above, that is guaranteed regardless of the truth.

Reporting p = 0 from a simulation. Drawing 10,000 allocations and finding none as extreme does not establish p = 0 ; it establishes p is small. The ( 1 + ⋅ ) / ( B + 1 ) form exists precisely to avoid claiming otherwise.

Contrast

Fisher and Neyman answer different questions

Fisher against Neyman. Both take their randomness from the assignment mechanism. Almost everything else differs.

FisherNeyman
What is nulled or targeted Y i ( 1 ) = Y i ( 0 ) for every unitThe average effect τ
Missing potential outcomesFilled in exactly, under the nullLeft missing; their variance contribution bounded
OutputA p-valueAn estimate, standard error, interval
ExactnessExact for any N Large-sample approximation
Assumes a distributionNoNormal approximation for the interval
Rejection meansTreatment affected at least one unit—

Why one does not determine the other. The sharp null is strictly stronger than a zero average effect. A treatment that helps half the units by 10 and harms the other half by 10 has τ = 0 , so Neyman's interval should cover zero, while the sharp null is flatly false, and a well-chosen statistic may reject it.

The converse also occurs: with few units, a real and uniform effect can leave the randomization test unable to reject anything, because the design admits too few allocations to place the observed value in a tail.

The practical reading. A randomization p-value answers "did this treatment do anything to anybody?". An interval answers "how big is the average effect, and how sure are we?". Reporting one as though it answered the other is the error this contrast exists to prevent.

Exercise

Work in order.

1: fully structured. Six units, three treated by complete randomization. Outcomes: 12, 10, 9 (treated) and 7, 6, 4 (control).

(a) Compute T obs . (b) How many allocations does the design permit? (c) What is the smallest two-sided p-value this design could produce?

Check: Y ¯ 1 = 31 / 3 ≈ 10.33 , Y ¯ 0 = 17 / 3 ≈ 5.67 , so T obs ≈ 4.67 ; ( 6 3 ) = 20 allocations; the observed split is the most extreme possible, and with its mirror image that is 2 / 20 = 0.10 , so 0.05 is unreachable here.

2: partly structured. A trial of 40 units reports a randomization p-value of 0.002, computed by drawing 5,000 allocations, with 9 at least as extreme as observed.

(a) Verify the reported figure. (b) Would reporting p = 0 have been defensible had none been as extreme? (c) What does this p-value establish about the average effect?

Check: ( 1 + 9 ) / ( 5001 ) ≈ 0.0020 ; no, with zero extreme draws the estimate is 1 / 5001 ≈ 0.0002 , small but not zero; it establishes that the treatment affected at least one unit, and says nothing directly about the size of the average effect.

3: unstructured. A researcher runs a matched-pair experiment with 12 pairs, then tests the sharp null by enumerating every way of splitting the 24 units into two groups of 12 and computing the difference in means for each.

Explain what is wrong, what should have been enumerated, and whether the resulting p-value is likely too large or too small.

Check: the design permitted 2 12 = 4096 allocations, one choice per pair, not ( 24 12 ) = 2,704,156 . The enumeration includes allocations the design could never produce, for example both members of a pair treated. The reference distribution is therefore wider than the true one, since pairing removes between-pair variation, so the observed statistic looks less extreme and the p-value comes out too large.

What to carry forward

The null. Y i ( 1 ) = Y i ( 0 ) for every unit. No effect on anyone. Sharp because it fills in every missing potential outcome.

Why that matters. With outcomes fixed, each permitted allocation is a relabelling, so the statistic can be recomputed exactly.

The p-value. p = P ( | T | ≥ | T obs | ) over the allocations the mechanism permits.

By simulation. p ^ = 1 + # { b : | T b | ≥ | T obs | } B + 1 . The added one counts the observed assignment and prevents reporting zero.

Exactness. No distributional assumption, no large-sample argument. Valid for four units.

What rejection means. Treatment affected at least one unit. Not that the average effect is non-zero, and not that it affected everyone.

What non-rejection means. The data are consistent with no effect for anyone, which a small design guarantees regardless of the truth. Check the smallest attainable p-value before interpreting.

Against Neyman. Same randomness, different question: "did anything happen?" against "how big is the average, and how sure are we?".

The recurring errors. Enumerating allocations the design did not permit; choosing the statistic after seeing the effects; reading a large p-value as evidence of no effect.

Next step

Practice Fisher Randomization Inference

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Blocked and Paired Randomized Experiments

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.