Blocked and Paired Randomized Experiments

What you will be able to do

Given a description of units and a design, the learner can decide whether blocking or pairing is warranted, compute the estimator the design requires, and identify analyses that ignore the design structure.

Orientation

If you know something about your units before assigning treatment, randomising within groups uses that knowledge. It also changes what the analysis must do.

The estimator is a weighted average or a mean of differences. The competence is noticing when an analysis has quietly thrown the design away.

Intuition

Using known unit characteristics before assignment

Complete randomization treats every unit as interchangeable. If you genuinely know nothing about them, that is the best you can do.

But you usually know something. Outcomes differ by site, by baseline severity, by school. Leaving that to the coin means accepting whatever imbalance the draw produces, and then carrying that imbalance as noise in the comparison.

Blocking uses the knowledge instead. Sort units into groups that are alike on the variable that matters, and randomize inside each group. Within a group the variable is held fixed by construction, so it cannot contribute to the difference you observe.

Pairing is the limiting case: groups of two. Each pair produces one number, the treated minus the control, and everything the two shared cancels out of it. Those differences vary less than raw outcomes do, which is what reduces the variance.

The price is a commitment made before any data arrive. The design has restricted which allocations were possible, and every later calculation has to respect that restriction.

Definition

Blocked and paired estimators

The canonical statement gives the two estimators. What it does not say is why the weights are what they are, or what the pair difference costs.

The weights N k / N name the estimand, not a convenience. Weighting stratum effects by stratum size targets the sample-average effect over all N units. Equal weights 1 / K target the average of the stratum effects, which is a different quantity whenever strata differ in size: a stratum of 12 and a stratum of 300 count equally under 1 / K and 25 times apart under N k / N . Neither is wrong; reporting one while describing the other is.

A pair difference removes everything the pair shares. D j subtracts the two units' common level, so whatever made them a pair — site, cohort, baseline severity — cancels before the estimate is formed. That is the precision the design secures in advance, and it is why S E ( τ ^ ) = s D / J is computed from J differences rather than from 2 J outcomes.

The analysis is committed to the blocks thereafter. Having randomized within strata, an analysis that pools across them is estimating something the design did not produce. The degrees of freedom follow the pairs, not the units, which is the arithmetic consequence of that commitment: J pairs give J − 1 , however many units they contain.

Example

Weights that answer different questions

A trial blocks by site. Site A holds 40% of the sample with an estimated effect of 8; site B holds 60% with an estimated effect of 2.

Size-weighted.

τ ^ = 0.4 ( 8 ) + 0.6 ( 2 ) = 4.4 .

This estimates the average effect over the units in the trial, which is usually what is wanted.

Equally weighted.

8 + 2 2 = 5.

This estimates the average of the two site effects, giving a small site and a large one the same influence.

Neither is wrong. They answer different questions, and 4.4 against 5 is not a rounding difference. It is a choice of estimand. Reporting one while describing the other is the error.

When they coincide. Equal strata sizes, or equal stratum effects. Neither is common enough to assume.

Worked example

Twelve pairs

Problem. A matched-pair experiment uses 12 pairs, matched on baseline score. The treated-minus-control differences have mean D ¯ = 1.8 and standard deviation s D = 2.4 .

A colleague proposes analysing the 24 units as two independent groups instead.

Goal. Produce the paired estimate and standard error, and say what the colleague's approach discards.

Relevant principle. The unit of analysis is the pair. The design permitted one choice per pair, so the analysis works on differences.

Step 1: the estimate.

τ ^ = D ¯ = 1.8 .

Reason: the average of within-pair differences estimates the average treatment effect directly, since each difference holds the pair's shared characteristics fixed.

Step 2: the standard error.

S E ( τ ^ ) = s D J = 2.4 12 ≈ 0.69 .

Reason: J = 12 independent differences, not 24 independent observations. The denominator is 12 , not 24 .

Step 3: the test statistic.

t = 1.8 0.69 ≈ 2.60 on  J − 1 = 11  degrees of freedom .

Step 4: what the alternative discards. Treating the 24 units as independently assigned computes a standard error from the spread of raw outcomes across all units. That spread includes the between-pair variation the matching was designed to remove, differences in baseline score that cancel exactly within a pair.

Reason: if pairs were well matched, the raw outcomes vary much more than the differences do, so the unpaired standard error is larger and the result looks weaker than the design earned.

Result. τ ^ = 1.8 , S E ≈ 0.69 , t ≈ 2.60 on 11 degrees of freedom.

Check. Does the pairing help here? s D = 2.4 is the spread of differences. If raw outcomes had a standard deviation of, say, 6, the unpaired standard error would be roughly 6 2 / 12 ≈ 2.4 , more than three times larger. The design secured precision, and only the paired analysis collects it.

Interpretation. Report the paired analysis. The unpaired one has lower power, and more importantly its standard error refers to a randomization distribution the experiment never had, since allocations treating both members of a pair were impossible by design.

Non-example

Groupings that are not blocking

Grouping after assignment. Splitting the analysis by a variable recorded after treatment began, adherence, completion, side effects, is not blocking. Those variables can be affected by treatment, so the groups are defined partly by what the treatment did, and the comparison within them is no longer protected by randomization.

Matching in an observational study. Pairing treated and untreated units on covariates after the fact resembles a paired design and is not one. In a matched-pair experiment the pairs are formed first and treatment is randomized within them; in observational matching, nobody randomized anything, and comparability rests on an assumption rather than on a mechanism.

Post-hoc subgroup analysis. Reporting the effect among older participants, having chosen 'older' after seeing the results, is not a blocked analysis. The subgroup was not a design feature, and the multiplicity it introduces is unaccounted for.

Re-randomizing until balanced. Drawing repeatedly and keeping the allocation that looks most balanced restricts the mechanism, but not in a way any standard analysis accounts for. If balance on a variable matters, block on it in advance; that is the version whose analysis is known.

A before-and-after comparison. Measuring each unit before and after treatment produces paired numbers and a paired t -test computes fine. But the second measurement is not a control outcome. It is the same unit later, with everything else that changed in the meantime. The algebra transfers; the causal interpretation does not.

Contrast

What the design permits, and what the analysis assumes

Complete randomizationBlockedPaired
Allocations permitted ( N N 1 ) ∏ k ( N k N 1 k ) 2 J
Estimator Y ¯ 1 − Y ¯ 0 ∑ k N k N τ ^ k D ¯
Standard error s 1 2 / N 1 + s 0 2 / N 0 Combined across strata s D / J
Unit of analysisUnitUnit, within stratumPair
Balance on the blocking variableBy chanceBy constructionBy construction

The failure this table exists to prevent. Every row of the paired column differs from the completely randomized column. An analysis that uses the first column's estimator, standard error and reference distribution on data from the third is not approximately right. It describes a different experiment.

Which direction the error runs. With well-chosen blocks, ignoring them inflates the standard error, so the result looks weaker than the design earned. With badly chosen blocks the loss is small. Neither case makes the unpaired analysis correct; it makes the cost variable.

Balance, twice. Complete randomization leaves balance to the draw and gives valid inference anyway. Blocking removes the question for the variable blocked on. The temptation in between, re-randomizing after seeing imbalance, takes the commitment of blocking without its known analysis.

Exercise

1: fully structured. A trial blocks by clinic. Clinic A: 150 units, τ ^ A = 6 . Clinic B: 50 units, τ ^ B = − 2 .

(a) Compute the size-weighted estimate. (b) Compute the equally weighted estimate. (c) Which targets the average effect over the 200 units, and what does the other target?

Check: ( 150 / 200 ) ( 6 ) + ( 50 / 200 ) ( − 2 ) = 4.5 − 0.5 = 4.0 ; equal weighting gives ( 6 − 2 ) / 2 = 2.0 ; the size-weighted figure targets the sample-average effect, the other the average of clinic effects. A difference of 2.0, not a rounding artefact.

2: partly structured. A paired trial with 9 pairs reports D ¯ = 3.1 and s D = 1.8 .

(a) Compute S E ( τ ^ ) and the t statistic with its degrees of freedom. (b) A reviewer asks why the denominator is not 18 . Answer them.

Check: S E = 1.8 / 3 = 0.6 , t = 3.1 / 0.6 ≈ 5.17 on 8 degrees of freedom; the design produced 9 independent differences, one per pair. The two members of a pair were not independently assigned, so there are 9 units of analysis, not 18.

3: unstructured. A team runs a stepped-wedge trial: twelve clinics all eventually receive the intervention, but the order in which they switch is randomized. The analyst compares all clinic-months under the intervention with all clinic-months before it, as two independent groups.

Say what the design constrains that this analysis ignores, and what the comparison risks confusing with a treatment effect.

Check: the randomization was over the order of switching, not over clinic-months, so the permitted allocations are orderings and the analysis must respect clinic and time. Pooling all post-intervention periods against all pre-intervention ones confounds the treatment with time: later periods are all post-intervention by construction, so any secular trend, seasonality, a policy change, improving practice, enters the estimate as though it were an effect.

What to carry forward

Blocking. Partition on a pretreatment variable, randomize within each stratum. Balance on that variable by construction rather than by luck.

Blocked estimator. τ ^ = ∑ k N k N τ ^ k . Size weights target the sample-average effect; equal weights target the average stratum effect.

Pairing. Blocking with two units per block. D j = Y j , T − Y j , C , τ ^ = D ¯ , S E = s D / J .

Unit of analysis. J pairs, not 2 J observations. The denominator is J .

Allocations permitted. 2 J for pairs, ∏ k ( N k N 1 k ) for blocks, not ( N N 1 ) . Randomization tests must enumerate the right set.

The commitment. Restricting the mechanism obliges the analysis to match. Ignoring the structure usually inflates the standard error and always misstates the reference distribution.

Before, not after. Block on what you expect to matter, in advance. Re-randomizing after seeing imbalance takes the cost without the known analysis.

The recurring error. Analysing a paired or blocked experiment as if every unit had been independently assigned.

Next step

Practice Blocked and Paired Randomized Experiments

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.