Blocked and Paired Randomized Experiments
What you will be able to do
Given a description of units and a design, the learner can decide whether blocking or pairing is warranted, compute the estimator the design requires, and identify analyses that ignore the design structure.
Orientation
If you know something about your units before assigning treatment, randomising within groups uses that knowledge. It also changes what the analysis must do.
The estimator is a weighted average or a mean of differences. The competence is noticing when an analysis has quietly thrown the design away.
Intuition
Using known unit characteristics before assignment
Complete randomization treats every unit as interchangeable. If you genuinely know nothing about them, that is the best you can do.
But you usually know something. Outcomes differ by site, by baseline severity, by school. Leaving that to the coin means accepting whatever imbalance the draw produces, and then carrying that imbalance as noise in the comparison.
Blocking uses the knowledge instead. Sort units into groups that are alike on the variable that matters, and randomize inside each group. Within a group the variable is held fixed by construction, so it cannot contribute to the difference you observe.
Pairing is the limiting case: groups of two. Each pair produces one number, the treated minus the control, and everything the two shared cancels out of it. Those differences vary less than raw outcomes do, which is what reduces the variance.
The price is a commitment made before any data arrive. The design has restricted which allocations were possible, and every later calculation has to respect that restriction.
Definition
Blocked and paired estimators
The canonical statement gives the two estimators. What it does not say is why the weights are what they are, or what the pair difference costs.
The weights
A pair difference removes everything the pair shares.
The analysis is committed to the blocks thereafter. Having randomized within strata, an analysis that pools across them is estimating something the design did not produce. The degrees of freedom follow the pairs, not the units, which is the arithmetic consequence of that commitment:
Example
Weights that answer different questions
A trial blocks by site. Site A holds 40% of the sample with an estimated effect of 8; site B holds 60% with an estimated effect of 2.
Size-weighted.
This estimates the average effect over the units in the trial, which is usually what is wanted.
Equally weighted.
This estimates the average of the two site effects, giving a small site and a large one the same influence.
Neither is wrong. They answer different questions, and 4.4 against 5 is not a rounding difference. It is a choice of estimand. Reporting one while describing the other is the error.
When they coincide. Equal strata sizes, or equal stratum effects. Neither is common enough to assume.
Worked example
Twelve pairs
Problem. A matched-pair experiment uses 12 pairs, matched on baseline score. The treated-minus-control differences have mean
A colleague proposes analysing the 24 units as two independent groups instead.
Goal. Produce the paired estimate and standard error, and say what the colleague's approach discards.
Relevant principle. The unit of analysis is the pair. The design permitted one choice per pair, so the analysis works on differences.
Step 1: the estimate.
Reason: the average of within-pair differences estimates the average treatment effect directly, since each difference holds the pair's shared characteristics fixed.
Step 2: the standard error.
Reason:
Step 3: the test statistic.
Step 4: what the alternative discards. Treating the 24 units as independently assigned computes a standard error from the spread of raw outcomes across all units. That spread includes the between-pair variation the matching was designed to remove, differences in baseline score that cancel exactly within a pair.
Reason: if pairs were well matched, the raw outcomes vary much more than the differences do, so the unpaired standard error is larger and the result looks weaker than the design earned.
Result.
Check. Does the pairing help here?
Interpretation. Report the paired analysis. The unpaired one has lower power, and more importantly its standard error refers to a randomization distribution the experiment never had, since allocations treating both members of a pair were impossible by design.
Non-example
Groupings that are not blocking
Grouping after assignment. Splitting the analysis by a variable recorded after treatment began, adherence, completion, side effects, is not blocking. Those variables can be affected by treatment, so the groups are defined partly by what the treatment did, and the comparison within them is no longer protected by randomization.
Matching in an observational study. Pairing treated and untreated units on covariates after the fact resembles a paired design and is not one. In a matched-pair experiment the pairs are formed first and treatment is randomized within them; in observational matching, nobody randomized anything, and comparability rests on an assumption rather than on a mechanism.
Post-hoc subgroup analysis. Reporting the effect among older participants, having chosen 'older' after seeing the results, is not a blocked analysis. The subgroup was not a design feature, and the multiplicity it introduces is unaccounted for.
Re-randomizing until balanced. Drawing repeatedly and keeping the allocation that looks most balanced restricts the mechanism, but not in a way any standard analysis accounts for. If balance on a variable matters, block on it in advance; that is the version whose analysis is known.
A before-and-after comparison. Measuring each unit before and after treatment produces paired numbers and a paired
Contrast
What the design permits, and what the analysis assumes
| Complete randomization | Blocked | Paired | |
|---|---|---|---|
| Allocations permitted | |||
| Estimator | |||
| Standard error | Combined across strata | ||
| Unit of analysis | Unit | Unit, within stratum | Pair |
| Balance on the blocking variable | By chance | By construction | By construction |
The failure this table exists to prevent. Every row of the paired column differs from the completely randomized column. An analysis that uses the first column's estimator, standard error and reference distribution on data from the third is not approximately right. It describes a different experiment.
Which direction the error runs. With well-chosen blocks, ignoring them inflates the standard error, so the result looks weaker than the design earned. With badly chosen blocks the loss is small. Neither case makes the unpaired analysis correct; it makes the cost variable.
Balance, twice. Complete randomization leaves balance to the draw and gives valid inference anyway. Blocking removes the question for the variable blocked on. The temptation in between, re-randomizing after seeing imbalance, takes the commitment of blocking without its known analysis.
Exercise
1: fully structured. A trial blocks by clinic. Clinic A: 150 units,
(a) Compute the size-weighted estimate. (b) Compute the equally weighted estimate. (c) Which targets the average effect over the 200 units, and what does the other target?
Check:
2: partly structured. A paired trial with 9 pairs reports
(a) Compute
Check:
3: unstructured. A team runs a stepped-wedge trial: twelve clinics all eventually receive the intervention, but the order in which they switch is randomized. The analyst compares all clinic-months under the intervention with all clinic-months before it, as two independent groups.
Say what the design constrains that this analysis ignores, and what the comparison risks confusing with a treatment effect.
Check: the randomization was over the order of switching, not over clinic-months, so the permitted allocations are orderings and the analysis must respect clinic and time. Pooling all post-intervention periods against all pre-intervention ones confounds the treatment with time: later periods are all post-intervention by construction, so any secular trend, seasonality, a policy change, improving practice, enters the estimate as though it were an effect.
What to carry forward
Blocking. Partition on a pretreatment variable, randomize within each stratum. Balance on that variable by construction rather than by luck.
Blocked estimator.
Pairing. Blocking with two units per block.
Unit of analysis.
Allocations permitted.
The commitment. Restricting the mechanism obliges the analysis to match. Ignoring the structure usually inflates the standard error and always misstates the reference distribution.
Before, not after. Block on what you expect to matter, in advance. Re-randomizing after seeing imbalance takes the cost without the known analysis.
The recurring error. Analysing a paired or blocked experiment as if every unit had been independently assigned.