Practice: Fisher Randomization Inference
Question
Recognition · Interpretation
Under the sharp null
2 hints available, least help first.
Hint 1: Retrieval cue
Write what
Hint 2: Concept cue
Ask why the null is described as sharp rather than merely as no average effect.
Direct application · Interpretation · Explanation
Five units are completely randomized, two to treatment. Observed outcomes: unit A = 9, B = 7 (treated); C = 6, D = 4, E = 2 (control).
(a) State the sharp null and compute
(b) How many allocations does the design permit, and what is the probability of each?
(c) List the values of
(d) State what your result establishes, and what the smallest attainable p-value is for this design.
(e) A colleague says the large p-value shows the average treatment effect is zero. Say whether the test supports that, and give a case that separates the two hypotheses.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Under the null, the five outcomes never change. Only which two are labelled treated changes.
Hint 2: Strategy cue
Before computing all ten, ask which allocations put the largest outcomes together, those give the extreme values.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a)
(b) Complete randomization with
(c) Holding the five outcomes fixed and relabelling: treating
(d) The sharp null is not rejected at conventional levels: these data are consistent with treatment having done nothing to anyone. The smallest attainable two-sided p-value is
(d) It does not. The randomization test nulls
A complete answer does each of these:
- states sharp null
- enumerates permitted allocations
- interprets rejection
- distinguishes from average null
Comparison · Method selection · Interpretation
Four experiments each have 16 units and each is to be tested against the sharp null. State how many allocations the randomization distribution contains in each case, and say what goes wrong if all
(i) Complete randomization, 8 treated.
(ii) Eight matched pairs, one treated per pair.
(iii) Bernoulli assignment with
(iv) Blocking into two strata of 8, with 4 treated in each stratum.
Finally: in which of these designs would rejecting the sharp null also establish that the average treatment effect is nonzero? Explain.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
For each design, ask what free choices the mechanism actually makes.
Hint 2: Concept cue
Three of these constrain the allocation more tightly than a free split; one constrains it less.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(i)
(ii)
(iii) All
(iv)
The common principle: the reference set is whatever the mechanism could have produced. Enumerating more than that answers a question about an experiment nobody ran.
Sharp null against average effect. In all four. Rejecting the sharp null says treatment changed something for somebody, in whichever design produced the reference set. It does not establish a nonzero average effect in any of them: effects that cancel across units leave
A complete answer does each of these:
- states sharp null
- enumerates permitted allocations
- interprets rejection
- distinguishes from average null
Error diagnosis · Explanation · Evaluation
A report on a completely randomized trial of 200 units states:
The randomization test gave
, so we accept the null that the average treatment effect is zero. Since Fisher's test and Neyman's interval both use the assignment mechanism, the confidence interval would tell us the same thing, and we omit it.
There are three errors. Identify each, state what the test did establish, and say what the report should contain instead.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Write down the hypothesis the randomization test actually nulls.
Hint 2: Concept cue
Consider separately: which null, what failing to reject licenses, and whether the two methods answer one question.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
First error: the wrong null. The randomization test nulls
Second error: accepting a null. Failing to reject means the data are consistent with no effect for anyone; it is not evidence that no effect exists. With 200 units the design has resolution, so this is less severe than in a tiny experiment, but the inference is still 'not rejected', not 'true'.
Third error: treating the two methods as interchangeable. Both draw randomness from the assignment mechanism, and there the similarity ends. Fisher returns a p-value about a per-unit null; Neyman returns an estimate, a standard error and an interval for the average effect. Neither determines the other, and the interval carries information the p-value does not, in particular the magnitude and precision of the estimated average.
What the report should contain. The randomization p-value with its null stated as a per-unit hypothesis, the estimated average effect with a conservative standard error and interval, and, if a null result is to be interpreted, some indication of what size of effect the design could have detected.
A complete answer does each of these:
- states sharp null
- enumerates permitted allocations
- interprets rejection
- distinguishes from average null
Transfer · Construction · Evaluation
An engineering team runs an A/B test on a website. Rather than randomizing users, their system assigns by hashing the user id and routing even hashes to variant B. The team wants an exact test of whether variant B changed anything for anyone, and proposes a randomization test that permutes the A/B labels freely across the 5,000 users.
Assess the proposal. Address whether an exact randomization test is available here at all, what the reference set would need to be, and what you would recommend.
The team adds that they only care whether B changed the average. Say whether that changes which test they should be asking for.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what set of allocations this assignment rule could have produced.
Hint 2: Strategy cue
Separate whether a mechanism exists from what would be enumerated if one did.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
Is an exact test available? Only if there is a mechanism with known probabilities. Hashing a user id is deterministic: given the id, the assignment is fixed, and it could not have come out otherwise. There is no set of allocations the design might have produced, so there is no reference distribution and no exact randomization test. The procedure is closer to assignment by an arbitrary but fixed characteristic than to randomization.
What the proposed reference set would be. Permuting labels freely assumes every split of 5,000 users into the observed group sizes was possible. Under the hash rule none of them was possible except the one observed. The resulting p-value would describe a randomized experiment that was never run.
Is it salvageable? Sometimes. If the hash includes a seed chosen at random per experiment, then re-drawing the seed is a genuine mechanism, and the reference set is the allocations generated by the seeds that could have been drawn, not free permutations. That is a real randomization test, but it must enumerate seeds rather than labels.
Recommendation. Either introduce an explicit random component to the assignment and enumerate over it, or drop the claim of exactness and analyse the comparison as observational, stating the assumption that hash parity is unrelated to the outcome. That assumption is plausible and is still an assumption, which is the distinction the team's proposal obscures.
Which question they are asking. It changes the test entirely. An exact randomization test addresses the sharp null, did B change anything for anyone, and answers with a p-value. A claim about the average is Neyman's question, answered with an estimate, a standard error and an interval. The two can disagree: gains for some users offset by losses for others leave the average near zero while the sharp null is plainly false. Since the hashing gives them no mechanism, neither test is available until assignment is actually randomized.
A complete answer does each of these:
- states sharp null
- enumerates permitted allocations
- interprets rejection
- distinguishes from average null
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.