Practice: Matching for Causal Inference
Question
Recognition · Interpretation
Each treated unit is matched to one control and the differences are averaged. Which quantity does this target?
2 hints available, least help first.
Hint 1: Retrieval cue
Ask which units the average is taken over.
Hint 2: Concept cue
One difference per treated unit. Whose effect is being averaged?
Method selection · Direct application · Evaluation
A study matches each of 1,200 enrolled students to the closest non-enrolled student by propensity score, with replacement and a caliper of 0.1. After matching, 1,050 enrolled students are retained, drawing on 640 distinct non-enrolled students. All standardized differences fall below 0.05. The authors report "the effect of the programme" as +6.2 percentage points.
Say what was estimated, and what the report must add.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask which units the average runs over, and how many were discarded.
Hint 2: Strategy cue
Take the three numbers, 1,200, 1,050, 640, and say what each one changes.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The estimand. Each enrolled student is matched to a non-enrolled one, so this targets the effect on the treated, what the programme did for students who enrolled, not for students generally.
The 150 dropped students. The caliper excluded enrolled students with no sufficiently close match, so the estimate describes the 1,050 retained rather than all 1,200. Those excluded are typically the most distinctive profiles, which may be the cases of most interest.
The replacement figure. 1,050 matches drew on 640 distinct controls, so some controls serve several treated students. This improves match quality and means the control side rests on 640 distinct outcomes, not 1,050. The matched sets are no longer independent, and the standard error must account for that.
What the balance figure shows. That the matching achieved comparability on the measured covariates, evidence about the procedure. It says nothing about whether those covariates suffice, which is unconfoundedness and is not assessable from these data. Enrolment was voluntary, so motivation, family support and prior engagement plausibly drive both enrolment and graduation, and are rarely measured.
How the claim should read. "Among the 1,050 enrolled students with a comparable non-enrolled match, the programme is associated with a 6.2-point increase in graduation, assuming the measured covariates capture what drove enrolment." Longer, and it is what was estimated.
What settles quality here. Post-match standardized differences and the overlap of the score distributions, not the estimate. Balance on the matched covariates is what the caliper and the matching were for; it is also all they can deliver, since nothing in the procedure addresses covariates nobody measured.
A complete answer does each of these:
- names the estimand
- tracks population changes
- separates from randomized pairs
- judges by balance
Comparison · Evaluation
What distinguishes a randomized matched-pair experiment from an observational propensity-matched study?
2 hints available, least help first.
Hint 1: Retrieval cue
Ask when the pairs were formed relative to treatment.
Hint 2: Concept cue
In each design, what decided which member of a pair was treated?
Error diagnosis · Explanation · Evaluation
A hospital reports:
We matched each of 400 patients receiving the new protocol to a historical patient on the old protocol, using age, sex and admission severity. Post-matching balance was excellent. The matched-pair design means we can interpret the 9% mortality reduction causally, as in a paired trial.
Identify the errors and say what should be reported instead.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what decided which protocol each patient received.
Hint 2: Concept cue
The two groups differ in more than protocol. What else separates them?
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
First error: calling it a paired trial. The pairs were assembled after treatment occurred, and nothing assigned protocol within a pair. In a randomized matched-pair experiment a coin decides who is treated within each pair, which balances unmeasured characteristics in expectation. Here the reason one patient got the new protocol and the other the old is simply that they were admitted at different times, and that reason may relate to how each would fare.
Second error: historical controls. The old-protocol patients were treated earlier, so the protocol is confounded with everything else that changed, staffing, adjunct treatments, admission criteria, diagnostic practice, even case mix. Matching on age, sex and severity does not touch any of it. This is the more serious problem, because the confounding is structural rather than incidental.
What the balance figure establishes. That the matching worked on three covariates. Unconfoundedness would require those three to capture everything driving both protocol and mortality, which is implausible when the protocols are separated in time.
The estimand. The effect on patients receiving the new protocol who had a comparable historical match, not the effect of the protocol generally. The report should say how many of the 400 went unmatched.
What should be reported. The estimand as stated above; the number unmatched; balance on the original covariates; the temporal confounding named plainly as a limitation; a sensitivity analysis for unmeasured differences; and, if at all possible, a concurrent comparison group rather than a historical one, since that is a design remedy and no analysis substitutes for it.
How the match should have been judged. By post-match balance and overlap on the matching variables, inspected before any outcome was compared, standardized differences, and whether the historical pool actually contained comparable patients. A match evaluated by the effect it produced has no check at all: the comparison is the thing in question.
A complete answer does each of these:
- names the estimand
- tracks population changes
- separates from randomized pairs
- judges by balance
Transfer · Evaluation · Explanation
An economist evaluates a regional enterprise grant. She writes:
For each of the 210 firms that received a grant, we selected a comparison firm from a neighbouring region with the closest sector, size and prior growth. 186 grant firms found an acceptable comparison. We then compared employment growth within each pair. Because every grant firm has its own control, this is effectively a paired design.
Assess the claim, name the estimand, and say what you would report.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what decided which firms received grants, and whether the three covariates capture it.
Hint 2: Strategy cue
The comparison firms come from elsewhere. What does that introduce besides grant status?
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
This is observational matching. The vocabulary is a comparison cohort and the structure is nearest-neighbour matching on three covariates. Calling it a paired design imports a guarantee it does not have: no mechanism assigned grant status within a pair, so comparability rests on an assumption rather than on randomization.
The estimand. The effect on the treated, specifically on the 186 grant firms that found an acceptable comparison, not on all 210 and not on firms generally. The 24 unmatched firms should be described, since firms with no comparable counterpart are often the most distinctive and may be where grants were most consequential.
The specific confounding risk. Grants were awarded for reasons, plausibly including management quality, growth plans, or an application capability, none of which is captured by sector, size and prior growth. Using a neighbouring region adds a second problem: regional economic conditions, local labour markets and other regional policies differ, so region is confounded with grant status in the same way the historical comparison confounds time with treatment.
What I would report. The estimand as the effect on the 186 matched grant firms; the number unmatched and how they differ; balance on the original covariates after matching; whether controls were reused and, if so, that the matched pairs are not independent; the regional confounding named as a limitation; and a sensitivity analysis for unmeasured confounders such as management quality. If firms just below and just above an award threshold exist, a design exploiting that discontinuity would rest on far weaker assumptions than matching does.
The missing check. Nothing in the note reports post-match balance or overlap. Without standardized differences after matching there is no evidence the comparison firms resemble the grant recipients on the variables used, and judging the scheme by the estimated grant effect inverts the order. The estimate is what the balance was supposed to license.
A complete answer does each of these:
- names the estimand
- tracks population changes
- separates from randomized pairs
- judges by balance
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.