Matching for Causal Inference
Matching builds a comparison by pairing each unit with a similar unit under the opposite treatment, then comparing outcomes. Its design choices, the distance, the caliper, replacement, which units are matched, decide both what is estimated and which population the estimate describes. Its resemblance to a paired experiment is superficial: the pairs are assembled after treatment occurred, and unconfoundedness is assumed exactly as before.
Definition
For unit
targets the average treatment effect on the treated, not automatically the ATE. Design choices include matching with replacement (a strong control may be reused, usually reducing bias but increasing dependence and reducing effective sample size) or without (each control used once, with results possibly depending on match order); a caliper rejecting matches farther apart than a chosen distance; exact or coarsened exact matching forcing equality on selected covariates; and Mahalanobis matching accounting for covariate scale and covariance.
Formal statement
Assumptions and scope
The opposite-treatment restriction is essential: a treated unit must be matched to a control and vice versa, or nothing causal is being compared.
Matching each treated unit to a control targets the effect on the treated, not the average treatment effect. The estimand follows from the scheme and must be stated.
Discarding unmatched units changes the population the estimate describes. The report must say which units remain and what estimand the matched sample supports.
Quality is judged by covariate balance and overlap after matching, never by whether the outcome comparison became favourable.
Matching with replacement usually reduces bias but induces dependence between matched sets and lowers the effective sample size; matching without replacement can make results depend on the order in which matches were formed.
A matched observational comparison still requires unconfoundedness. Balance on measured covariates is not evidence that the covariate set suffices.
Matching is not a randomized paired design. The pairs are assembled after treatment occurred, so no mechanism supplies the comparability that randomization would have.
Worked material
Example
The same data, three estimands
A study has 500 treated units and 5,000 controls.
Scheme 1 — each treated unit matched to its nearest control, without replacement. All 500 treated units are retained. The estimate targets the effect on the treated: what the treatment did for the units that received it.
Scheme 2: the same, with a caliper of 0.05 on the propensity score. Eighty treated units have no control within the caliper and are dropped. The remaining 420 are better matched. The estimate now targets the effect on those 420. A subpopulation defined by having a comparable control available, which typically means the less extreme treated units.
Scheme 3: each control matched to its nearest treated unit. The roles reverse, and the estimate targets the effect on the controls: what the treatment would have done for units that did not receive it.
Three numbers, three questions. These can differ substantially whenever effects vary with the characteristics that drove treatment, which is the ordinary case, not an exotic one.
What a report must say. Which scheme was used, how many units were discarded, and which population the estimate describes. "We matched on propensity score and found an effect of 3.2" leaves the reader unable to tell which of these three quantities the 3.2 refers to.
On the 80 discarded units. They are not a rounding error. They are the treated units least like any control, often the most intensively treated, and excluding them may remove exactly the cases of most interest.
Non-example
Things matching does not deliver
It is not a randomized paired design. Pairs assembled after treatment occurred carry no mechanism. The comparability is assumed, not produced.
It does not address unmeasured confounding. Matching operates on the covariates supplied to it. Perfect balance on those says nothing about the ones nobody measured.
Balance is not proof of sufficiency. Standardized differences below a threshold show the procedure worked on the variables it was given. The same limit that applies to weighting and to regression adjustment.
A matched estimate is not automatically the ATE. Matching treated to control targets the effect on the treated, and calipers narrow it further.
Matching on a post-treatment variable is not matching. As everywhere in this subject, only pretreatment covariates are admissible; matching on a consequence of treatment removes part of the effect.
Choosing the scheme by the answer is not a design decision. Trying several matching specifications and reporting the one with the most favourable outcome makes the reported uncertainty meaningless. Quality is judged by balance and overlap, before the outcomes are examined.
Contrast
Matched pairs, randomized and observational
| Randomized matched pairs | Observational matching | |
|---|---|---|
| When pairs are formed | Before assignment | After treatment occurred |
| What decides treatment within a pair | A known random mechanism | The unit's own circumstances |
| Source of comparability | The mechanism | An assumption about measured covariates |
| Unmeasured differences within a pair | Balanced in expectation by the coin | Unaddressed |
| What the analysis assumes | Nothing beyond the design | Unconfoundedness given |
| Estimand by default | The sample average effect | The effect on the treated |
| Permitted allocations | None — assignment is not a mechanism |
Why the confusion is so natural. Both produce a table of pairs, and both can be analysed by differencing within pairs. The arithmetic is genuinely similar.
What differs entirely. In the randomized design, the reason two paired units received different treatments is a coin, and the coin knows nothing about them. In the observational design, the reason is whatever led one to be treated and the other not, and that reason may well relate to how each would respond.
The consequence for reporting. A matched observational study should not be described as "as good as randomized". It is a study in which comparability on measured covariates has been improved and the central assumption is unchanged. Saying so is not excessive caution; it is the difference between what was done and what the reader will otherwise assume.
Common errors
Common misconception
Matching treated units to similar controls creates matched pairs, so the study can be analysed and interpreted like a paired randomized experiment.
Related units
Requires
Connected
- Blocked and Paired Randomized Experiments (contrasts with)