Module 2 of 4 · Lesson 1 of 1

Randomized Assignment

The assignment mechanism, and why it rather than observed balance is what licenses a causal comparison.

What you will be able to do

Given a description of how treatment was assigned, the learner can identify the mechanism, state whether it is known and independent of the potential outcomes, say what causal claim it supports, and explain why observed imbalance in a single realisation is not evidence against it.

Orientation

Randomisation does not balance the groups. It makes the imbalance a known quantity, which is a different and more useful guarantee.

That is a different skill from computing a difference in means, and it is the one that decides whether the difference means anything. The arithmetic is identical whether treatment was assigned by a coin, by a waiting list, or by a clinician's judgement about who would benefit. Only the first licenses a causal reading.

Intuition

What randomization guarantees, and what it does not

Randomization is often described as making the groups the same. It does not. With sixty units and a coin, one group will be a little older, or a little sicker, or contain more of whatever you did not measure.

What randomization does is make the rule independent of the units. The coin has no access to anyone's potential outcomes, so it cannot preferentially place high responders in treatment. Over the assignments the rule could have produced, the two groups match on everything at once, including the variables nobody recorded, which is the part no amount of statistical adjustment can replicate.

That is why the mechanism has to be written down before outcomes are seen, and why "we compared the people who took it with the people who did not" is not a design. The value lies in the rule, not in the split it happened to produce.

Definition

Assignment mechanisms

An assignment mechanism gives the probability of each possible treatment vector W = ( W 1 , … , W N ) .

Bernoulli assignment. Each unit is treated independently with probability p :

P ( W = w ) = ∏ i = 1 N p w i ( 1 − p ) 1 − w i .

The treated count N 1 = ∑ i W i is random. Every vector in { 0 , 1 } N has positive probability, including all-treated and all-control.

Complete randomization. Exactly N 1 of the N units are treated. There are ( N N 1 ) such vectors, each with probability

P ( W = w ) = ( N N 1 ) − 1 ,

and every other vector has probability zero. The treated count is fixed by design; individual assignments are therefore dependent, since knowing N 1 − 1 of them constrains the last.

The estimator. With Y ¯ 1 and Y ¯ 0 the observed treated and control means,

τ ^ = Y ¯ 1 − Y ¯ 0 .

Under complete randomization E [ τ ^ ] = τ S , the finite-sample average treatment effect. The expectation is taken over assignments, with the potential outcomes held fixed.

Example

Three procedures, one of which is not randomization

A clinic flips a fair coin for each of 40 patients. Bernoulli assignment with p = 0.5 . The treated count is random: it could be 17, or 23, and with probability 2 × 0.5 40 it is 0 or 40.

A clinic writes 40 slips, 20 marked "treat", shuffles, and deals one per patient. Complete randomization with N 1 = 20 . Exactly twenty are treated in every possible realisation; there are ( 40 20 ) ≈ 1.4 × 10 11 equally likely allocations.

A clinic offers the treatment and records who accepts. Not an assignment mechanism at all. The probability of each vector is unknown and depends on the patients, who plausibly decide on grounds related to how much they expect to benefit, exactly the dependence on potential outcomes that randomization exists to exclude.

The first two support a causal reading of τ ^ . The third supports a description of who accepted.

Worked example

The groups came out unbalanced

Problem. Six patients are completely randomized, three to treatment. The realised assignment and observed outcomes are:

PatientAge W i Y i obs
171114
244012
368111
439015
547011
666117

Mean age is 68.3 in treatment and 43.3 in control. A reviewer objects that the randomization has failed and asks for it to be run again.

Goal. Decide what the imbalance shows and what, if anything, should be done.

Relevant principle. Randomization makes the assignment rule independent of the units; it does not make any single realisation balanced.

Step 1: count the allocations. With N = 6 and N 1 = 3 there are ( 6 3 ) = 20 equally likely assignments.
Reason: complete randomization fixes the treated count and spreads probability uniformly over the vectors that satisfy it.

Step 2: ask how unusual this one is. The three oldest patients are 71, 68 and 66. Exactly one of the 20 allocations places all three in treatment, so this outcome had probability 1 / 20 = 0.05 under the design.
Reason: the probability is computed from the mechanism, which is known, rather than from a model of the data.

Step 3: ask what re-randomizing would do. Drawing again until the ages look similar replaces the stated mechanism with a different one: "uniform over allocations whose age difference is small". That rule is still independent of the outcomes, but it is no longer the rule used to compute any subsequent standard error or randomization test.
Reason: every inference in the design-based approach refers to the distribution of assignments the mechanism actually allows.

Result. The imbalance is not evidence of failure. It is one of twenty allocations the design permits, and a one-in-twenty event is not unusual. Nothing needs correcting, and re-drawing after seeing the imbalance would silently change the reference distribution.

Check. Does the estimate look distorted? Y ¯ 1 = 14 , Y ¯ 0 ≈ 12.67 , so τ ^ ≈ 1.33 . If age depresses the outcome, this estimate is pulled downward, which is exactly the randomization error the design accounts for by averaging over assignments, not by fixing this one.

Interpretation. If age is expected to matter, the time to act is before assignment: block on age and randomize within blocks. That is a design decision, made in advance, and it changes the mechanism deliberately rather than in response to a realisation.

Non-example

Five procedures described as randomized that are not

Alternation. Patients are assigned treatment, control, treatment, control by order of arrival. The rule is deterministic: knowing a patient's position determines their assignment, and whoever schedules arrivals controls the allocation. It is also predictable, so a clinician who prefers a patient to receive treatment can delay them by one slot.

Assignment by date. Everyone admitted on an even-numbered day is treated. Admission date is not a chance device. It correlates with staffing, referral patterns, and the severity mix of who presents when.

Haphazard is not random. "We just split them however" has no stated probabilities. Without a mechanism there is nothing to average over, so no standard error or randomization test has a defined reference distribution, even if the split happens to look balanced.

Assignment after seeing outcomes. A pilot treats the first ten patients, observes poor responses, and moves later patients into control. The rule now depends on the outcomes, which is precisely the dependence randomization excludes. No amount of subsequent adjustment restores it.

Volunteering. Participants choose their arm. The probability of each vector is unknown and is plausibly related to expected benefit.

Contrast

Bernoulli, complete, and not random at all

Bernoulli against complete randomization.

BernoulliComplete randomization
Treated countRandomFixed at N 1
Unit assignmentsIndependentDependent, since the total is fixed
Allowed vectorsAll 2 N Only those with exactly N 1 treated
Probability of a vector p n 1 ( 1 − p ) N − n 1 ( N N 1 ) − 1
Degenerate allocationsPossibleExcluded by construction

Both are valid mechanisms and both are independent of the potential outcomes. They differ in what they hold fixed, and that difference propagates: the randomization distribution of τ ^ is not the same under the two, so a test or standard error computed for one does not apply to the other.

The distinction that matters more. Either of these against a rule that consults the units. Bernoulli and complete randomization are two ways of being random; alternation, self-selection and clinician choice are not near-misses but a different kind of thing, and no analysis converts one into the other.

Exercise

Work these in order. The first states the mechanism; the last requires you to find it.

1: stated mechanism. Twelve units, completely randomized with N 1 = 4 .

(a) How many allocations does the design permit? (b) What is the probability of any particular one? (c) A colleague notes that the four treated units are the four with the highest baseline score and asks whether this is possible under the design.

Check: ( 12 4 ) = 495 ; each has probability 1 / 495 ; it is possible and has probability 1 / 495 ≈ 0.002 , unlikely, but permitted, and observing it does not mean the mechanism was violated.

2: mechanism to be named. A study reports: "Each participant was assigned to the intervention with probability one half, independently of the others. Of the 50 participants, 28 received the intervention."

(a) Name the mechanism. (b) Is 28 out of 50 surprising? (c) What would have been impossible under complete randomization with N 1 = 25 ?

Check: Bernoulli with p = 0.5 ; 28 is unremarkable since the count is random with mean 25 and standard deviation 50 × 0.25 ≈ 3.5 ; under complete randomization with N 1 = 25 a treated count of 28 could not occur at all.

3: mechanism to be found. A regional programme reports: "Places were limited, so the 200 applicants were ranked by the date their application arrived and the first 100 were enrolled. Enrolled and non-enrolled applicants were then compared on employment twelve months later."

State whether this is an assignment mechanism in the sense of this unit, what quantity the comparison estimates, and what would have to be assumed for it to estimate a causal effect.

Check: it is not a randomized mechanism, arrival order is deterministic and plausibly related to motivation, information and circumstance, all of which bear on employment. The comparison estimates the difference in mean employment between earlier and later applicants. A causal reading requires assuming that, absent the programme, early and late applicants would have had the same employment outcomes: an assumption about the applicants, not a consequence of the procedure.

What to carry forward

What a mechanism is. The probability of each treatment vector W , fixed by the design before outcomes are seen.

Bernoulli. Independent per unit with probability p ; treated count random; all 2 N vectors possible.

Complete randomization. Exactly N 1 treated; ( N N 1 ) equally likely vectors; assignments dependent.

Why it supplies anything. The rule is independent of the potential outcomes, so it cannot sort units by how they would respond. Groups match in expectation on everything, measured and unmeasured.

The estimator. τ ^ = Y ¯ 1 − Y ¯ 0 , unbiased for τ S over repeated assignments, not correct on any single one.

Imbalance. Ordinary. It is randomization error, not failure, and re-drawing after seeing it changes the mechanism that every later inference refers to. Balance is secured in advance by blocking, not afterwards by redrawing.

The recurring error. Treating any procedure that produced two groups as a design. Alternation, arrival order, admission date, volunteering and clinician judgement are not randomization, and the difference in means they produce estimates a descriptive contrast.

Next step

Practice Randomized Assignment

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Neyman Repeated-Sampling Inference

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.