Module 4 of 4 · Lesson 1 of 4
Unconfoundedness and Overlap
Unconfoundedness and overlap, and why only overlap can be checked from the data.
What you will be able to do
Given an observational study, the learner can state the assumptions adjustment requires, judge which of them the data can speak to, and say what the comparison would estimate if an assumption failed.
Orientation
Two assumptions stand between an observational comparison and a causal claim. One is checkable from the data; the other is not, and no amount of data will change that.
Most published effect estimates come from studies like this. The competence is not scepticism; it is knowing which question to ask.
Intuition
Reading a covariate table for the two assumptions
Both assumptions are usually decided by looking at a table of covariates. Here is what to look at, and what the table cannot tell you.
The table. A study of an after-school programme, 800 pupils, 240 enrolled.
| Covariate | Enrolled | Not enrolled | Range overlap |
|---|---|---|---|
| Prior test score | 42.1 | 55.8 | full |
| Family income (£000) | 19.4 | 31.2 | full |
| Parent contacted school | 71% | 28% | full |
| Already in remedial support | 94% | 6% | partial |
What the first three rows support. Large differences, full overlap. Differences are what adjustment is for; the arms differing on a covariate is the ordinary case, not a defect. Each row is also a prompt to ask whether that variable plausibly drove enrollment and affects the outcome, because those are the ones that have to be in the model.
What the fourth row does. 94% against 6% with partial range overlap means almost every enrolled pupil was already in remedial support and almost no unenrolled pupil was. For pupils in remedial support there are few untreated comparisons; for pupils outside it there are few treated ones. An estimate for "all pupils" here is largely model extrapolation, and the defensible response is to restrict the estimand to the region where both arms are populated and report what that region contains.
The asymmetry, made concrete. Everything above was read off the table. Now the other question: did anything unrecorded drive enrollment? Teacher judgement about which pupils would benefit, parental persistence, a pupil's own willingness. None appear in any column, and each could affect outcomes directly. No row of this table, and no statistic computed from it, bears on that. A perfect overlap diagnostic and a perfectly balanced adjustment are both compatible with a decisive omitted variable.
So the two assumptions fail differently. Overlap failure announces itself in the data and can be handled by changing what is claimed. Unconfoundedness failure is silent, and the only responses available are design-based: find a variable that was recorded, argue from how the decision was actually made, or seek a source of variation that was not chosen by anyone.
Definition
The two assumptions, and the identification they give
The canonical statement gives both assumptions and the identification they jointly deliver. The asymmetry between them is what governs practice.
Only one of the two can be checked. Overlap is a statement about the observed covariate distribution: estimate
The covariate set has to be pretreatment. A variable measured after assignment may sit on the causal path, and conditioning on it removes part of the effect being estimated. A variable associated with treatment but not with the outcome adds variance without removing bias. Deciding membership of
A failure of overlap changes the estimand rather than the estimate. Where support is thin, an estimator either extrapolates the outcome model into a region with no data or, if the sample is trimmed, reports an effect for a population that no longer matches the one named. Both are defensible; reporting either as the ATE for the original population is not.
Example
Where each assumption fails
Unconfoundedness fails. A study compares patients who received a new surgical technique with those who received the standard one, adjusting for age, sex and comorbidity count. Surgeons chose the technique, and they chose partly on operative fitness. A judgement recorded nowhere in the data. Fit patients did better regardless of technique. Within every level of the measured covariates, assignment still depends on the potential outcomes, so the adjusted estimate mixes the technique's effect with the surgeons' selection.
Nothing in the data reveals this. The balance table looks fine, because the variable driving it was never measured.
Overlap fails. A study of an intensive tutoring programme finds that every student below the 20th percentile was enrolled. For those students
Both hold, plausibly. A workplace randomly audited by a regulator on a published rota, with the rota depending only on recorded sector and size. Assignment depends on
Worked example
Reading a balance table
Problem. A study estimates the effect of a job-training programme on earnings using propensity-score weighting. The report states: after weighting, all standardized mean differences on the twelve measured covariates are below 0.05, and the propensity model achieves an AUC of 0.88. The authors conclude the groups are comparable and the adjusted difference is causal.
Goal. Say what has been established and what has not.
Relevant principle. Balance on measured covariates is evidence about the adjustment procedure, not about the sufficiency of the covariate set.
Step 1: what the balance table shows. Standardized differences below 0.05 indicate the weighting succeeded at what it was asked to do: the weighted treated and control groups have similar distributions on those twelve variables.
Reason: that is the definition of the diagnostic. It compares the two groups on variables that were measured.
Step 2: what it cannot show. Unconfoundedness concerns whether the twelve are enough. Prior motivation, health, family support and local labour demand may all drive both enrolment and earnings. If any is unmeasured, the groups can be perfectly balanced on the twelve and still differ systematically.
Reason: the assumption is about the joint distribution of assignment and the potential outcomes given
Step 3: read the AUC correctly. An AUC of 0.88 says treatment is well predicted by the covariates. That is not reassurance; it is a warning about overlap. Strong prediction means many units have propensity scores near zero or one, which are the cases where comparable units are scarce and weights become extreme.
Reason: the propensity model's job is balance, not classification. A model that predicts assignment perfectly would leave no comparable units at all.
Step 4: what should be reported. The overlap diagnostics: the distribution of
Result. The study has demonstrated successful balancing on twelve covariates. It has not demonstrated that those twelve suffice. The high AUC does not by itself establish an overlap failure, but it is a reason to look: it is consistent with covariate separation sharp enough to leave parts of the population without comparable units, which the score distributions would show directly.
Check. Is there a diagnostic that would establish unconfoundedness? No. This is not a gap in the reporting but a property of the assumption: it concerns unobserved quantities. What can be done is a sensitivity analysis, asking how strong an unmeasured confounder would need to be to overturn the result.
Interpretation. Report the estimate with its assumptions stated, the overlap diagnostics shown, and a sensitivity analysis. Describing the study as 'as good as randomized' asserts precisely the thing that cannot be checked.
Non-example
Things that do not establish unconfoundedness
A balance table. Balance on measured covariates shows the adjustment worked on those covariates. The assumption is about whether they are sufficient, which the same data cannot address.
A large sample. Confounding is bias, not noise. A million observations give a precise estimate of a confounded quantity.
Adjusting for everything available. Including post-treatment variables can block part of the effect or open a non-causal path by conditioning on a collider. The rule is pretreatment common causes, not maximum coverage.
A good propensity model. High classification accuracy means treatment is predictable from covariates, which strains overlap. The score's purpose is balance, not prediction.
A statistical test for confounding. No test of the observed data can detect an unmeasured confounder, because the data carry no trace of a variable nobody recorded.
Similar point estimates from several methods. Matching, weighting and regression that agree have agreed about the same covariate set under the same assumption. They can be jointly wrong in the same direction, and usually would be.
Contrast
Which assumption the data can check
| Overlap | Unconfoundedness | |
|---|---|---|
| Statement | ||
| About | Measured covariates and assignment | Assignment and unobserved outcomes |
| Checkable from data | Yes | No |
| How | Propensity distributions, weights, effective sample size | Not available |
| Failure looks like | Extreme weights, extrapolation, estimates driven by few units | Nothing at all |
| Remedy | Trim, restrict the estimand, redesign | Measure more, or a sensitivity analysis |
Why the asymmetry matters. Analysts check what is checkable and then report as though both assumptions had been verified. The checkable one is the less consequential of the two.
Against randomization. A randomized experiment gets unconfoundedness from the mechanism and overlap by construction. Every unit had a genuine chance of either arm. An observational study must assume the first and demonstrate the second. That is the cost of not having assigned treatment, and it is why design occupies the first half of this subject.
Failure is not visible in the output. Poor overlap is detectable in diagnostics. Unmeasured confounding produces a clean table, a tight interval and a wrong answer.
Exercise
1: fully structured. A study compares patients on drug A with those on drug B, adjusting for age, sex, baseline severity and comorbidity.
(a) State unconfoundedness in terms of these variables. (b) State overlap. (c) Which could you check with the data in hand?
Check: (a) among patients with the same age, sex, baseline severity and comorbidity, which drug was prescribed is independent of how each would fare on either; (b) at every combination of those four, both drugs must have a positive probability of being prescribed; (c) overlap, by inspecting the distribution of estimated prescribing probability in each arm, and post-adjustment balance. Not unconfoundedness.
2: partly structured. In the same study, 8% of patients have an estimated propensity below 0.02.
(a) What does that indicate? (b) The analyst trims them. What changes? (c) What must the report now say?
Check: those patients had almost no chance of drug A, so there are few comparable treated units and their weights would be enormous; trimming removes them, stabilising the estimate; the estimand is no longer the ATE over all patients but over the retained subpopulation, and the report must say which patients the estimate describes.
3: unstructured. A city reports that neighbourhoods which installed CCTV saw a 12% larger fall in burglary than those which did not, adjusting for prior burglary rate, population density and median income.
Assess the causal claim. Name a plausible unmeasured confounder, say what it would do to the estimate, and state which assumption your concern is about.
Check: installation was chosen, plausibly where residents organised or where councillors were responsive, community cohesion and local political attention are unmeasured and plausibly reduce burglary independently. That is a violation of unconfoundedness, not overlap, and it would inflate the apparent effect by attributing the consequences of an engaged community to the cameras. No balance table on the three measured covariates can address it; a sensitivity analysis can bound how strong such a confounder would need to be.
What to carry forward
Unconfoundedness.
Overlap.
Together they license
Covariates must be pretreatment. Adjusting for what the treatment influenced can remove part of the effect or open a non-causal path.
The asymmetry. Overlap and post-adjustment balance are visible in the data. Unconfoundedness is not, at any sample size, by any test.
When overlap fails. Estimates rest on extrapolation and a few heavily weighted units. Trimming helps and changes the estimand, which must then be restated.
When unconfoundedness fails. Nothing looks wrong. The table balances, the interval is tight, the answer is biased.
The recurring error. Presenting a balance table as evidence that adjustment sufficed.