Module 4 of 4 · Lesson 4 of 4

Matching for Causal Inference

Matching units into comparable pairs, and how the estimand changes with the matched sample.

What you will be able to do

Given a described matching procedure, the learner can say which estimand it targets, how its design choices change the population described, and what it does not establish.

Orientation

Find someone similar and compare. The idea is intuitive; what it estimates, and for whom, depends entirely on choices made while matching.

Matching is the most intuitive method in this part of the subject, which is why its two traps catch so many studies: the estimand quietly changes, and the result gets described as though a coin had been tossed.

Intuition

How each matching choice changes the estimand

The canonical intuition states that matching assembles comparable pairs after treatment occurred, and that this differs from a matched-pair experiment. This block traces what each implementation choice does to the quantity being estimated.

Suppose 500 treated units and 5,000 controls, and consider four specifications.

One-to-one, no caliper, all treated units retained. Every treated unit is matched, so the estimand is the average effect on the treated over all 500. Some matches will be poor, because a nearest available control is not necessarily a similar one.

One-to-one with a caliper of 0.1 standard deviations. Treated units with no control inside the caliper are dropped. If 60 are dropped, the estimand is the average effect on the remaining 440, and those 440 are the treated units that resemble some control. The dropped units are typically the ones most strongly indicated for treatment, so the estimate no longer describes them.

Matching with replacement. A single control may serve several treated units. Match quality improves, and the effective sample shrinks: if 440 treated units draw on 160 distinct controls, the variance reflects 160 rather than 440, and standard errors computed as though the controls were distinct are too small.

Many-to-one, five controls per treated unit. Variance falls because each comparison averages five outcomes, and bias rises because the second through fifth nearest controls are progressively less similar.

Each specification answers a different question. Reporting the estimand, the number of treated units retained, and the number of distinct controls used is what allows a reader to tell which question was answered.

Definition

Matches, estimands, and design choices

Nearest neighbour. For unit i , under covariate distance d ,

j ( i ) = arg ⁡ min j : W j ≠ W i d ( X i , X j ) ,

and using propensity-score distance,

j ( i ) = arg ⁡ min j : W j ≠ W i | e ^ ( X i ) − e ^ ( X j ) | .

The opposite-treatment restriction W j ≠ W i is essential: a treated unit must be compared with a control, and vice versa.

Matched estimator for the treated. With each treated unit matched to one control,

τ ^ A T T = 1 N T ∑ i : W i = 1 ( Y i obs − Y j ( i ) obs ) ,

which targets the average treatment effect on the treated, not automatically the ATE. The target changes with the matching scheme, the units being matched, the replacement rule and any weights.

Design choices.

  • With replacement. A strong control may be reused; usually reduces bias, but induces dependence between matched sets and lowers the effective sample size.
  • Without replacement. Each control used once; results can depend on the order in which matches are formed.
  • Caliper, rejects matches farther apart than a chosen distance, improving comparability and discarding units.
  • Exact or coarsened exact, forces equality on selected covariates or categories.
  • Mahalanobis, accounts for the scale and covariance of the covariates.

Assessment. Quality is judged by covariate balance and overlap after matching, standardized differences, variance ratios, covariate distributions, how many units went unmatched, and sensitivity to the design choices, never by whether the outcome comparison became favourable.

Example

The same data, three estimands

A study has 500 treated units and 5,000 controls.

Scheme 1 — each treated unit matched to its nearest control, without replacement. All 500 treated units are retained. The estimate targets the effect on the treated: what the treatment did for the units that received it.

Scheme 2: the same, with a caliper of 0.05 on the propensity score. Eighty treated units have no control within the caliper and are dropped. The remaining 420 are better matched. The estimate now targets the effect on those 420. A subpopulation defined by having a comparable control available, which typically means the less extreme treated units.

Scheme 3: each control matched to its nearest treated unit. The roles reverse, and the estimate targets the effect on the controls: what the treatment would have done for units that did not receive it.

Three numbers, three questions. These can differ substantially whenever effects vary with the characteristics that drove treatment, which is the ordinary case, not an exotic one.

What a report must say. Which scheme was used, how many units were discarded, and which population the estimate describes. "We matched on propensity score and found an effect of 3.2" leaves the reader unable to tell which of these three quantities the 3.2 refers to.

On the 80 discarded units. They are not a rounding error. They are the treated units least like any control, often the most intensively treated, and excluding them may remove exactly the cases of most interest.

Worked example

Reading a matched design

Problem. A study estimates the effect of a mentoring programme on graduation. 1,200 students enrolled; 9,000 did not. The authors match each enrolled student to the non-enrolled student with the closest propensity score, with replacement and a caliper of 0.1. After matching, 1,050 enrolled students are retained, drawing on 640 distinct non-enrolled students. All standardized differences fall below 0.05. They report "the effect of the programme" as +6.2 percentage points.

Goal. Say what was estimated and what the report must add.

Relevant principle. The estimand follows from the matching scheme, and discarding or reusing units changes the population described and the information available.

Step 1: identify the estimand. Each enrolled student is matched to a non-enrolled one, so this targets the effect on the treated, what the programme did for students who enrolled.

Reason: the matched set is constructed around the treated units, so the average is over them.

Step 2: account for the 150 dropped students. The caliper excluded 150 enrolled students with no sufficiently close match. The estimate describes the 1,050 retained, not all 1,200.

Reason: a unit with no acceptable match contributes nothing, so it is outside the population the estimate covers, and such students are typically those with the most distinctive profiles.

Step 3: read the replacement figure. 1,050 matches drew on only 640 distinct controls, so some controls serve several treated students.

Reason: reuse improves match quality and means the control side rests on 640 distinct outcomes, not 1,050. The matched sets are also no longer independent, which the standard error must reflect.

Step 4: judge the balance figure correctly. Standardized differences below 0.05 show the matching achieved comparability on the measured covariates. That is evidence about the procedure.

Reason: it says nothing about whether the measured covariates suffice, which is unconfoundedness and is not assessable from these data.

Step 5: name the plausible confounder. Enrolment was voluntary, so motivation, family support and prior engagement plausibly drive both enrolment and graduation, and are rarely measured.

Result. The study estimates the effect on the 1,050 matched enrolled students, resting on 640 distinct controls with reuse, under an unverifiable assumption.

Check. Would a different scheme give a different number? Almost certainly. Without replacement, match quality would fall and the dependence would vanish; a tighter caliper would drop more students and narrow the population further. Reporting sensitivity to these choices is part of reporting the result.

Interpretation. Rewrite the claim as: "Among the 1,050 enrolled students with a comparable non-enrolled match, the programme is associated with a 6.2-point increase in graduation, under the assumption that the measured covariates capture what drove enrolment." That is longer, and it is what was estimated.

Non-example

Things matching does not deliver

It is not a randomized paired design. Pairs assembled after treatment occurred carry no mechanism. The comparability is assumed, not produced.

It does not address unmeasured confounding. Matching operates on the covariates supplied to it. Perfect balance on those says nothing about the ones nobody measured.

Balance is not proof of sufficiency. Standardized differences below a threshold show the procedure worked on the variables it was given. The same limit that applies to weighting and to regression adjustment.

A matched estimate is not automatically the ATE. Matching treated to control targets the effect on the treated, and calipers narrow it further.

Matching on a post-treatment variable is not matching. As everywhere in this subject, only pretreatment covariates are admissible; matching on a consequence of treatment removes part of the effect.

Choosing the scheme by the answer is not a design decision. Trying several matching specifications and reporting the one with the most favourable outcome makes the reported uncertainty meaningless. Quality is judged by balance and overlap, before the outcomes are examined.

Contrast

Matched pairs, randomized and observational

Randomized matched pairsObservational matching
When pairs are formedBefore assignmentAfter treatment occurred
What decides treatment within a pairA known random mechanismThe unit's own circumstances
Source of comparabilityThe mechanismAn assumption about measured covariates
Unmeasured differences within a pairBalanced in expectation by the coinUnaddressed
What the analysis assumesNothing beyond the designUnconfoundedness given X
Estimand by defaultThe sample average effectThe effect on the treated
Permitted allocations 2 J , knownNone — assignment is not a mechanism

Why the confusion is so natural. Both produce a table of pairs, and both can be analysed by differencing within pairs. The arithmetic is genuinely similar.

What differs entirely. In the randomized design, the reason two paired units received different treatments is a coin, and the coin knows nothing about them. In the observational design, the reason is whatever led one to be treated and the other not, and that reason may well relate to how each would respond.

The consequence for reporting. A matched observational study should not be described as "as good as randomized". It is a study in which comparability on measured covariates has been improved and the central assumption is unchanged. Saying so is not excessive caution; it is the difference between what was done and what the reader will otherwise assume.

Exercise

1: fully structured. A study matches each of 300 treated units to one control, without replacement, from a pool of 2,000.

(a) Which estimand does this target? (b) If a caliper drops 40 treated units, what changes? (c) What should the report state?

Check: (a) the effect on the treated; (b) the estimate now describes the 260 retained treated units, those with a comparable control, which are typically the less extreme cases; (c) the scheme, the number discarded, the population the estimate describes, post-matching balance on the original covariates, and that unconfoundedness is assumed.

2: partly structured. An analyst matches with replacement and finds that 500 treated units drew on 180 distinct controls.

(a) What does the reuse achieve? (b) What does it cost? (c) What must the standard error account for?

Check: (a) better matches, since the closest available control can serve whoever needs it, usually reducing bias; (b) the control side rests on 180 distinct outcomes rather than 500, so the effective sample size is far below the apparent count; (c) the dependence between matched sets that share a control, treating the 500 differences as independent would understate the uncertainty.

3: unstructured. A hospital reports: "We matched each of 400 patients receiving the new protocol to a historical patient on the old protocol, using age, sex and admission severity. Post-matching balance was excellent. The matched-pair design means we can interpret the 9% mortality reduction causally, as in a paired trial."

Assess the claim and say what should be reported instead.

Check: the design is observational, and calling it a paired trial is the error, pairs were assembled after treatment, and no mechanism assigned protocol within a pair, so comparability on the three covariates is achieved while unconfoundedness remains assumed. The historical control makes it worse: the old-protocol patients were treated at an earlier time, so the protocol is confounded with everything else that changed, staffing, other treatments, admission criteria, diagnostic practice. Balance on age, sex and severity does not touch that. What should be reported: the estimand as the effect on patients receiving the new protocol who had a comparable historical match; how many were unmatched; balance on the original covariates; the temporal confounding stated plainly; and a sensitivity analysis for unmeasured differences, with a concurrent rather than historical comparison group if one can be constructed.

What to carry forward

The match. j ( i ) = arg ⁡ min j : W j ≠ W i d ( X i , X j ) , nearest under a covariate or propensity-score distance, always to a unit under the opposite treatment.

The default estimand. τ ^ A T T = 1 N T ∑ i : W i = 1 ( Y i − Y j ( i ) ) . The effect on the treated, not the ATE.

Design choices change the question. Replacement, calipers, exact or Mahalanobis distance, and which side is matched all alter the estimand, the population and the effective sample size.

Discarded units narrow the population. State how many went unmatched and which units the estimate describes.

Replacement trades bias for dependence. Better matches, fewer distinct controls, and matched sets that are no longer independent.

Judge by balance and overlap. Standardized differences, variance ratios, unmatched counts and sensitivity to the scheme, assessed before the outcomes are examined, never by whether the comparison became favourable.

It is not a paired experiment. Pairs assembled after treatment carry no mechanism, and unconfoundedness is assumed exactly as in any other observational analysis.

The recurring error. Describing a matched sample as "as good as randomized".

Next step

Practice Matching for Causal Inference

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lesson

This is the last lesson in Causal Inference from Experiments and Observational Data. Review the course map to see what is left.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.