Module 4 of 4 · Lesson 2 of 4
Propensity Scores
The propensity score as a summary of the covariate vector, and the balance checks that assess it.
What you will be able to do
Given an observational study, the learner can say what the propensity score is for, judge a fitted score by balance and overlap rather than by predictive accuracy, and state what it does not repair.
Orientation
A model can predict treatment assignment extremely well and be a poor tool for causal adjustment. That is not a paradox; it is the central fact about propensity scores, and it catches people who arrive with a machine-learning instinct that better prediction is better.
The score exists to make units comparable, not to classify them. So the questions that matter are whether comparable units exist at all, and whether conditioning on the score balanced the covariates, neither of which a discrimination metric answers.
Intuition
One score, two studies, and what balance actually showed
The argument for the score is short. What it looks like when it works, and when it quietly does not, takes a worked case.
A study where it worked. A job-training programme, 1,200 applicants, 400 enrolled. Enrollees were younger, less educated and had lower prior earnings, all three differing by more than half a standard deviation. Fitting the score on those covariates and comparing within strata of it, the standardised differences fell to under 0.05 on every covariate, and the two arms' score distributions overlapped across the whole range except the top 2%. That last detail is the one to notice: the 2% with no comparable untreated units were dropped, and the estimate that followed was for the remaining 98%, not for the original population. The estimand changed, and saying so is part of the result.
A study where the diagnostics looked better and meant less. A hospital compared a new protocol against standard care using a score fitted on 40 covariates. The model separated the arms almost perfectly, AUC 0.94, and balance on all 40 covariates was excellent. Both facts are consistent with a study that cannot answer the question: near-perfect separation means most treated patients had scores close to 1, where few untreated patients live, so the comparison rested on a thin slice of the data. The balance statistics were computed on that slice and were excellent about it. Good balance on a sample that no longer represents anyone is not a repair.
What the contrast shows. Balance is checked after adjusting, and it is checked on whoever is left. Two questions have to be asked together, never one without the other: is the balance good, and who is it good for? A high AUC does not condemn a specification, but it does indicate where to look: at the score distributions, and at how many units survived the overlap restriction.
The part the theorem does not cover. Balance on measured covariates is achievable and checkable. It implies nothing about covariates nobody recorded. A study can produce flawless balance tables and remain confounded by a variable that never entered the model, and no diagnostic computed from the data will say so.
Definition
What the balancing property does and does not deliver
Three parts of the definition above carry more weight than their length suggests.
"For the true score" is the load-bearing qualifier.
"In expectation" is the second qualifier. The property says the covariate distributions agree in expectation within a stratum of the score, not that they agree in any particular sample. In a small stratum they can differ considerably by chance, which is why balance is assessed with standardized differences and variance ratios rather than by eye.
The dimension reduction moves the problem; it does not solve it. Unconfoundedness given
On the estimator. Logistic regression is conventional rather than required. Any method producing a probability will do, and the choice between them is settled by the balance achieved across a region where both arms are represented, not by fit statistics.
Example
Two models of the same data
A study of a training programme has 2,000 participants and eight covariates. Two assignment models are fitted.
Model A. Eight main effects, logistic. Scores run from 0.14 to 0.83, the two arms' distributions overlap across that whole range, and after adjustment all standardized differences fall below 0.06. Classification is unimpressive: AUC 0.68.
Model B. Eight main effects, all pairwise interactions, and several polynomial terms. Scores run from 0.001 to 0.998, with 15% of controls below 0.01. Classification is excellent: AUC 0.94.
Which is the better adjustment? Model A, decisively. It achieves what the score exists to achieve, balance on the measured covariates, across a region where both arms are represented.
Model B has found combinations of covariates that all but determine assignment. That is a genuine discovery about the data, and it is bad news rather than good: those 15% of controls have almost no comparable treated units, so for them the comparison has no empirical basis. Model B has not improved the adjustment; it has revealed, and then worsened, an overlap problem.
What generalises, and what does not. Not a law that better prediction means worse adjustment. What generalises is that discrimination and adjustment quality are different measurements: strong predictive performance often reflects substantial separation between the arms, and substantial separation is where overlap becomes doubtful, so a high AUC is a reason to inspect the score distributions, never a diagnosis on its own. Model B was rejected on its 15% of controls below 0.01 and on nothing else. Had its scores stayed inside a region both arms occupied, its AUC would have been no argument against it.
Worked example
Fitting and checking a score
Problem. An analyst estimates the effect of a workplace wellness programme on sick days. Participation was voluntary. Available variables: age, sex, department, job grade, prior-year sick days, baseline health-risk questionnaire score, and this year's gym attendance.
Produce a propensity-score workflow and say what would make you reject the fitted model.
Goal. A fitted score whose adjustment is defensible, with the checks that would show it was not.
Relevant principle. The score is built to balance measured pretreatment covariates. Its quality is judged by balance and overlap, never by how well it predicts participation.
Step 1: choose covariates. Age, sex, department, job grade, prior-year sick days and baseline questionnaire score all plausibly affect both participation and sick days, and all precede the programme. Include them.
Reason: the set should contain common causes of treatment and outcome, chosen from subject-matter knowledge rather than from what improves fit.
Step 2: exclude this year's gym attendance. It is measured after the programme began and the programme plausibly increased it.
Reason: conditioning on a variable the treatment influenced removes part of the effect being estimated. Its predictive value is exactly why it is tempting and irrelevant to whether it belongs.
Step 3: fit without consulting outcomes. Estimate
Reason: choosing a specification by the answer it produces invalidates the reported uncertainty, since the selection is then a function of the outcome data.
Step 4: inspect overlap before anything else. Plot
Reason: if overlap fails there is no adjustment worth performing, and no later diagnostic repairs it.
Step 5: adjust, then check balance on the covariates. After weighting or stratification, compute standardized differences and variance ratios on the six covariates, plus important transformations and interactions.
Reason: the balancing property is a theorem about the true score, and
Step 6: revise on balance, not on fit. If a covariate remains unbalanced, add flexibility for that covariate and refit.
Reason: the criterion for revision is the balance the model failed to achieve, not its likelihood or classification accuracy.
Result. A workflow whose decisions at every step are driven by balance and overlap.
Check. What would make me reject the fitted model? Persistent imbalance on prior-year sick days after adjustment; a substantial share of either arm with scores beyond 0.05 or 0.95; or the discovery that a strong specification was retained because it raised classification accuracy.
Interpretation. Even a model passing every one of these checks establishes balance on six measured covariates. Whether those six suffice is unconfoundedness, which no diagnostic here addresses. Motivation to participate is unmeasured and plausibly affects sick days directly, so the result deserves a sensitivity analysis alongside it.
Non-example
Things a propensity score does not do
It does not address unmeasured confounding. The score is a function of the covariates it was given. If motivation, health literacy or family circumstances drive both treatment and outcome and none was recorded, the score cannot represent them and adjustment cannot remove them.
It does not make an observational study randomized. Randomization is a mechanism with known probabilities. An estimated score reconstructs an assignment probability under an assumption that cannot be checked; the arithmetic can be made to resemble an experiment while the epistemic position remains entirely different.
A well-predicting model is not a well-adjusting one. Strong discrimination often accompanies sharp covariate separation, and sharp separation is what pushes scores toward zero and one, so a high AUC is a reason to inspect overlap, not a measurement of it. AUC alone diagnoses neither overlap nor adjustment quality; the score distributions by arm and the post-adjustment balance do.
Balance on the score alone is not balance. Two arms can match on
It does not license including everything measured. Post-treatment variables and instruments both damage the adjustment. The first by blocking part of the effect, the second by amplifying bias from whatever confounding remains.
Trimming is not a repair. Removing units with extreme scores stabilises the estimate and changes the population it describes. That is a redefinition of the estimand and must be reported as one.
Contrast
Prediction and balance are different goals
| Prediction goal | Balance goal | |
|---|---|---|
| Question asked | Who was treated? | Which units are comparable? |
| Success measure | AUC, accuracy, cross-validated loss | Standardized differences, variance ratios after adjustment |
| What a good result looks like | Sharp separation of the arms | Substantial overlap of the arms |
| Effect of strong covariates | Improves the metric | Can push scores toward 0 and 1, where overlap needs checking |
| Variables to include | Anything predictive, including post-treatment ones | Pretreatment common causes only |
| Uses outcomes | Freely | Never |
Why the confusion is so natural. The score is a predicted probability, produced by tools whose default diagnostics are all classification metrics. Every habit a modelling course instils points the wrong way here.
What the right answer looks like. A propensity model with a mediocre AUC and excellent post-adjustment balance is doing its job. One with a superb AUC and 15% of controls below 0.01 is reporting that the comparison cannot be made for those units.
Against the randomized case. Under complete or Bernoulli randomization the true score is known by design and does not depend on the pretreatment covariates, say 0.5 for everyone, which gives perfect overlap and balance in expectation, and an attainable AUC of 0.5. (Under blocked or stratified assignment it is known but constant only within a design stratum, and may differ across strata.) The design under which causal inference is easiest is therefore one whose assignment is least predictable from covariates. That is not a rule that lower AUC is better: it is a demonstration that discrimination is measuring something other than what the adjustment needs.
Exercise
1: fully structured. A study fits a propensity model and reports scores from 0.22 to 0.79, AUC 0.64, and post-adjustment standardized differences all below 0.04.
(a) Is the AUC a problem? (b) What do the other two figures tell you? (c) What remains unestablished?
Check: (a) no, but not because a low AUC is itself good, discrimination is not what a propensity model is for, so the figure is close to irrelevant on its own. What settles the overlap question here is the score range: nothing near zero or one. A low AUC accompanied by scores piled at the extremes would still be an overlap failure, and a low AUC from a model that simply omitted the covariates that drive assignment would be a worse problem than a high one; (b) the score range shows no units pressed against zero or one, and the standardized differences show the adjustment balanced the measured covariates; (c) whether those covariates suffice, which is unconfoundedness and is not addressed by any of these numbers.
2: partly structured. An analyst is choosing between two specifications. Model 1: AUC 0.71, worst standardized difference after weighting 0.03, no scores outside [0.08, 0.88]. Model 2: AUC 0.89, worst standardized difference 0.02, 11% of units with scores below 0.02.
(a) Which would you use? (b) What is Model 2's better balance figure worth here? (c) What would you report if forced to use Model 2?
Check: (a) Model 1. It achieves balance across a region where both arms are represented; (b) little, because balance computed over a sample containing units with no comparable counterparts describes an adjustment resting on extrapolation for those units; (c) the overlap failure, the effective sample size, the largest weights, and the restricted population the estimate actually describes after any trimming.
3: unstructured. A team building a churn-prevention evaluation writes: "We used gradient boosting for the propensity model because it substantially outperformed logistic regression on held-out AUC (0.91 versus 0.73). We also included last-month engagement, which was the strongest predictor of receiving the retention offer."
Assess the two decisions. Say what you would check, and what you would change.
Check: both decisions optimise prediction at the expense of the adjustment. The AUC improvement signals that the model has found near-deterministic assignment patterns, so overlap must be inspected before anything else, expect scores near zero and one, extreme weights, and a much reduced effective sample size. Last-month engagement is more serious: if the retention offer was triggered by it, it is a pretreatment cause and admissible, but if it was measured after the offer went out, it is post-treatment and conditioning on it removes part of the effect. Establish its timing first. The change: select the specification by post-adjustment balance on the original covariates, use the flexible model only if it balances better across a region of genuine overlap, and report balance, score distributions and effective sample size rather than AUC.
Warning
Overlap failure, imbalance, and what each requires reporting
Some findings are not problems to work around. They are the analysis telling you what it cannot do.
Overlap fails. A share of one arm sits at scores near zero or one, with no comparable units in the other arm. For those units there is no comparison to make, not a noisy one, none. No weighting scheme, no richer specification and no larger sample repairs it, because the data contain no counterparts to borrow from.
What you must do: report the share affected and the score regions involved, before reporting any estimate.
Trimming was used to restore overlap. Dropping units with extreme scores does stabilise the estimate. It also changes who the estimate is about: the answer now describes the trimmed population, not the one you set out to study.
What you must do: state the trimming rule, how many units it removed from each arm, and name the restricted population as the estimand. A trimmed analysis reported as though it answered the original question is the more serious error, because nothing in the output reveals it.
A covariate stays unbalanced after adjustment. The score was built to balance the measured covariates and did not manage it for this one.
What you must do: add flexibility for that covariate and refit, on balance, never on classification accuracy or on the effect estimate. If it remains unbalanced, say so; an adjustment that failed on a covariate you named as a confounder has not done its job.
And the one that never appears in a diagnostic. Every check here concerns the covariates you measured. Unconfoundedness is a claim about the ones you did not, and no balance table, overlap plot or effective sample size speaks to it. A flawless diagnostic report establishes that the measured covariates were handled well, and nothing whatever about whether they were the right ones.