Propensity Scores
The probability of treatment given the covariates is a single number that can stand in for all of them. Conditioning on it balances the covariates it was built from, which turns a high-dimensional adjustment problem into a scalar one. It does nothing whatever about variables nobody measured. Judge a specification by covariate balance and overlap rather than by how well it predicts treatment, which measures something the adjustment does not need.
Definition
The propensity score is the conditional probability of treatment given the observed pretreatment covariates,
Formal statement
Assumptions and scope
The balancing property holds for the true propensity score. Practice uses
, so balance on the original covariates must be checked after adjustment rather than inferred from the theorem. The score is a device for acting on unconfoundedness, not a substitute for it. It balances the covariates it was built from and is silent about every variable nobody measured.
The modelling goal is balance, not classification. High predictive accuracy often accompanies substantial separation between the arms, which can push scores toward zero and one and leave comparable units scarce; discrimination alone diagnoses neither overlap nor adjustment quality, in either direction.
Only pretreatment covariates belong in the model. Including anything the treatment could have influenced can block part of the effect or open a non-causal path.
Outcomes must not be used when specifying or revising the assignment model. Selecting a specification by the effect estimate it produces invalidates the reported uncertainty.
Balance on
alone is insufficient evidence. Balance must be assessed on the original covariates and on important transformations and interactions of them.
Worked material
Example
Two models of the same data
A study of a training programme has 2,000 participants and eight covariates. Two assignment models are fitted.
Model A. Eight main effects, logistic. Scores run from 0.14 to 0.83, the two arms' distributions overlap across that whole range, and after adjustment all standardized differences fall below 0.06. Classification is unimpressive: AUC 0.68.
Model B. Eight main effects, all pairwise interactions, and several polynomial terms. Scores run from 0.001 to 0.998, with 15% of controls below 0.01. Classification is excellent: AUC 0.94.
Which is the better adjustment? Model A, decisively. It achieves what the score exists to achieve, balance on the measured covariates, across a region where both arms are represented.
Model B has found combinations of covariates that all but determine assignment. That is a genuine discovery about the data, and it is bad news rather than good: those 15% of controls have almost no comparable treated units, so for them the comparison has no empirical basis. Model B has not improved the adjustment; it has revealed, and then worsened, an overlap problem.
What generalises, and what does not. Not a law that better prediction means worse adjustment. What generalises is that discrimination and adjustment quality are different measurements: strong predictive performance often reflects substantial separation between the arms, and substantial separation is where overlap becomes doubtful, so a high AUC is a reason to inspect the score distributions, never a diagnosis on its own. Model B was rejected on its 15% of controls below 0.01 and on nothing else. Had its scores stayed inside a region both arms occupied, its AUC would have been no argument against it.
Non-example
Things a propensity score does not do
It does not address unmeasured confounding. The score is a function of the covariates it was given. If motivation, health literacy or family circumstances drive both treatment and outcome and none was recorded, the score cannot represent them and adjustment cannot remove them.
It does not make an observational study randomized. Randomization is a mechanism with known probabilities. An estimated score reconstructs an assignment probability under an assumption that cannot be checked; the arithmetic can be made to resemble an experiment while the epistemic position remains entirely different.
A well-predicting model is not a well-adjusting one. Strong discrimination often accompanies sharp covariate separation, and sharp separation is what pushes scores toward zero and one, so a high AUC is a reason to inspect overlap, not a measurement of it. AUC alone diagnoses neither overlap nor adjustment quality; the score distributions by arm and the post-adjustment balance do.
Balance on the score alone is not balance. Two arms can match on
It does not license including everything measured. Post-treatment variables and instruments both damage the adjustment. The first by blocking part of the effect, the second by amplifying bias from whatever confounding remains.
Trimming is not a repair. Removing units with extreme scores stabilises the estimate and changes the population it describes. That is a redefinition of the estimand and must be reported as one.
Contrast
Prediction and balance are different goals
| Prediction goal | Balance goal | |
|---|---|---|
| Question asked | Who was treated? | Which units are comparable? |
| Success measure | AUC, accuracy, cross-validated loss | Standardized differences, variance ratios after adjustment |
| What a good result looks like | Sharp separation of the arms | Substantial overlap of the arms |
| Effect of strong covariates | Improves the metric | Can push scores toward 0 and 1, where overlap needs checking |
| Variables to include | Anything predictive, including post-treatment ones | Pretreatment common causes only |
| Uses outcomes | Freely | Never |
Why the confusion is so natural. The score is a predicted probability, produced by tools whose default diagnostics are all classification metrics. Every habit a modelling course instils points the wrong way here.
What the right answer looks like. A propensity model with a mediocre AUC and excellent post-adjustment balance is doing its job. One with a superb AUC and 15% of controls below 0.01 is reporting that the comparison cannot be made for those units.
Against the randomized case. Under complete or Bernoulli randomization the true score is known by design and does not depend on the pretreatment covariates, say 0.5 for everyone, which gives perfect overlap and balance in expectation, and an attainable AUC of 0.5. (Under blocked or stratified assignment it is known but constant only within a design stratum, and may differ across strata.) The design under which causal inference is easiest is therefore one whose assignment is least predictable from covariates. That is not a rule that lower AUC is better: it is a demonstration that discrimination is measuring something other than what the adjustment needs.
Common errors
Common misconception
A propensity-score model should be judged by how well it predicts treatment, so a higher AUC, accuracy or cross-validated fit means a better adjustment.
Related units
Requires
Connected
- Randomized Assignment (contrasts with)
- Inverse Probability Weighting (suggested next)