Propensity Scores

What you will be able to do

Given an observational study, the learner can say what the propensity score is for, judge a fitted score by balance and overlap rather than by predictive accuracy, and state what it does not repair.

Orientation

A model can predict treatment assignment extremely well and be a poor tool for causal adjustment. That is not a paradox; it is the central fact about propensity scores, and it catches people who arrive with a machine-learning instinct that better prediction is better.

The score exists to make units comparable, not to classify them. So the questions that matter are whether comparable units exist at all, and whether conditioning on the score balanced the covariates, neither of which a discrimination metric answers.

Intuition

One score, two studies, and what balance actually showed

The argument for the score is short. What it looks like when it works, and when it quietly does not, takes a worked case.

A study where it worked. A job-training programme, 1,200 applicants, 400 enrolled. Enrollees were younger, less educated and had lower prior earnings, all three differing by more than half a standard deviation. Fitting the score on those covariates and comparing within strata of it, the standardised differences fell to under 0.05 on every covariate, and the two arms' score distributions overlapped across the whole range except the top 2%. That last detail is the one to notice: the 2% with no comparable untreated units were dropped, and the estimate that followed was for the remaining 98%, not for the original population. The estimand changed, and saying so is part of the result.

A study where the diagnostics looked better and meant less. A hospital compared a new protocol against standard care using a score fitted on 40 covariates. The model separated the arms almost perfectly, AUC 0.94, and balance on all 40 covariates was excellent. Both facts are consistent with a study that cannot answer the question: near-perfect separation means most treated patients had scores close to 1, where few untreated patients live, so the comparison rested on a thin slice of the data. The balance statistics were computed on that slice and were excellent about it. Good balance on a sample that no longer represents anyone is not a repair.

What the contrast shows. Balance is checked after adjusting, and it is checked on whoever is left. Two questions have to be asked together, never one without the other: is the balance good, and who is it good for? A high AUC does not condemn a specification, but it does indicate where to look: at the score distributions, and at how many units survived the overlap restriction.

The part the theorem does not cover. Balance on measured covariates is achievable and checkable. It implies nothing about covariates nobody recorded. A study can produce flawless balance tables and remain confounded by a variable that never entered the model, and no diagnostic computed from the data will say so.

Definition

What the balancing property does and does not deliver

Three parts of the definition above carry more weight than their length suggests.

"For the true score" is the load-bearing qualifier. W ⟂ X ∣ e ( X ) is a theorem, and it is a theorem about e ( X ) , not about e ^ ( X ) . Every practical analysis substitutes an estimate, and the estimate inherits none of the guarantee. This is why balance is something checked after adjusting rather than something the theorem establishes, and why a balance table is not a formality but the only evidence that the substitution worked.

"In expectation" is the second qualifier. The property says the covariate distributions agree in expectation within a stratum of the score, not that they agree in any particular sample. In a small stratum they can differ considerably by chance, which is why balance is assessed with standardized differences and variance ratios rather than by eye.

The dimension reduction moves the problem; it does not solve it. Unconfoundedness given X implies unconfoundedness given e ( X ) . The implication runs in that direction only. The score turns an unmanageable adjustment on many covariates into a manageable one on a single number, and it is silent about whether those covariates were the right ones. Nothing about a propensity model, however carefully built, produces evidence for the assumption it operates under.

On the estimator. Logistic regression is conventional rather than required. Any method producing a probability will do, and the choice between them is settled by the balance achieved across a region where both arms are represented, not by fit statistics.

Example

Two models of the same data

A study of a training programme has 2,000 participants and eight covariates. Two assignment models are fitted.

Model A. Eight main effects, logistic. Scores run from 0.14 to 0.83, the two arms' distributions overlap across that whole range, and after adjustment all standardized differences fall below 0.06. Classification is unimpressive: AUC 0.68.

Model B. Eight main effects, all pairwise interactions, and several polynomial terms. Scores run from 0.001 to 0.998, with 15% of controls below 0.01. Classification is excellent: AUC 0.94.

Which is the better adjustment? Model A, decisively. It achieves what the score exists to achieve, balance on the measured covariates, across a region where both arms are represented.

Model B has found combinations of covariates that all but determine assignment. That is a genuine discovery about the data, and it is bad news rather than good: those 15% of controls have almost no comparable treated units, so for them the comparison has no empirical basis. Model B has not improved the adjustment; it has revealed, and then worsened, an overlap problem.

What generalises, and what does not. Not a law that better prediction means worse adjustment. What generalises is that discrimination and adjustment quality are different measurements: strong predictive performance often reflects substantial separation between the arms, and substantial separation is where overlap becomes doubtful, so a high AUC is a reason to inspect the score distributions, never a diagnosis on its own. Model B was rejected on its 15% of controls below 0.01 and on nothing else. Had its scores stayed inside a region both arms occupied, its AUC would have been no argument against it.

Worked example

Fitting and checking a score

Problem. An analyst estimates the effect of a workplace wellness programme on sick days. Participation was voluntary. Available variables: age, sex, department, job grade, prior-year sick days, baseline health-risk questionnaire score, and this year's gym attendance.

Produce a propensity-score workflow and say what would make you reject the fitted model.

Goal. A fitted score whose adjustment is defensible, with the checks that would show it was not.

Relevant principle. The score is built to balance measured pretreatment covariates. Its quality is judged by balance and overlap, never by how well it predicts participation.

Step 1: choose covariates. Age, sex, department, job grade, prior-year sick days and baseline questionnaire score all plausibly affect both participation and sick days, and all precede the programme. Include them.

Reason: the set should contain common causes of treatment and outcome, chosen from subject-matter knowledge rather than from what improves fit.

Step 2: exclude this year's gym attendance. It is measured after the programme began and the programme plausibly increased it.

Reason: conditioning on a variable the treatment influenced removes part of the effect being estimated. Its predictive value is exactly why it is tempting and irrelevant to whether it belongs.

Step 3: fit without consulting outcomes. Estimate e ^ ( X ) by logistic regression on the six retained covariates, never inspecting the effect estimate while deciding the specification.

Reason: choosing a specification by the answer it produces invalidates the reported uncertainty, since the selection is then a function of the outcome data.

Step 4: inspect overlap before anything else. Plot e ^ ( X ) by arm. Look for participants below 0.05 or above 0.95, for regions occupied by one arm alone, and for the proportion of each arm in those regions.

Reason: if overlap fails there is no adjustment worth performing, and no later diagnostic repairs it.

Step 5: adjust, then check balance on the covariates. After weighting or stratification, compute standardized differences and variance ratios on the six covariates, plus important transformations and interactions.

Reason: the balancing property is a theorem about the true score, and e ^ ( X ) is an estimate. Balance is therefore something to verify, not something the theorem delivers.

Step 6: revise on balance, not on fit. If a covariate remains unbalanced, add flexibility for that covariate and refit.

Reason: the criterion for revision is the balance the model failed to achieve, not its likelihood or classification accuracy.

Result. A workflow whose decisions at every step are driven by balance and overlap.

Check. What would make me reject the fitted model? Persistent imbalance on prior-year sick days after adjustment; a substantial share of either arm with scores beyond 0.05 or 0.95; or the discovery that a strong specification was retained because it raised classification accuracy.

Interpretation. Even a model passing every one of these checks establishes balance on six measured covariates. Whether those six suffice is unconfoundedness, which no diagnostic here addresses. Motivation to participate is unmeasured and plausibly affects sick days directly, so the result deserves a sensitivity analysis alongside it.

Non-example

Things a propensity score does not do

It does not address unmeasured confounding. The score is a function of the covariates it was given. If motivation, health literacy or family circumstances drive both treatment and outcome and none was recorded, the score cannot represent them and adjustment cannot remove them.

It does not make an observational study randomized. Randomization is a mechanism with known probabilities. An estimated score reconstructs an assignment probability under an assumption that cannot be checked; the arithmetic can be made to resemble an experiment while the epistemic position remains entirely different.

A well-predicting model is not a well-adjusting one. Strong discrimination often accompanies sharp covariate separation, and sharp separation is what pushes scores toward zero and one, so a high AUC is a reason to inspect overlap, not a measurement of it. AUC alone diagnoses neither overlap nor adjustment quality; the score distributions by arm and the post-adjustment balance do.

Balance on the score alone is not balance. Two arms can match on e ^ ( X ) while differing on a covariate whose effect the model mis-specified. Balance must be examined on the original covariates and on their important transformations.

It does not license including everything measured. Post-treatment variables and instruments both damage the adjustment. The first by blocking part of the effect, the second by amplifying bias from whatever confounding remains.

Trimming is not a repair. Removing units with extreme scores stabilises the estimate and changes the population it describes. That is a redefinition of the estimand and must be reported as one.

Contrast

Prediction and balance are different goals

Prediction goalBalance goal
Question askedWho was treated?Which units are comparable?
Success measureAUC, accuracy, cross-validated lossStandardized differences, variance ratios after adjustment
What a good result looks likeSharp separation of the armsSubstantial overlap of the arms
Effect of strong covariatesImproves the metricCan push scores toward 0 and 1, where overlap needs checking
Variables to includeAnything predictive, including post-treatment onesPretreatment common causes only
Uses outcomesFreelyNever

Why the confusion is so natural. The score is a predicted probability, produced by tools whose default diagnostics are all classification metrics. Every habit a modelling course instils points the wrong way here.

What the right answer looks like. A propensity model with a mediocre AUC and excellent post-adjustment balance is doing its job. One with a superb AUC and 15% of controls below 0.01 is reporting that the comparison cannot be made for those units.

Against the randomized case. Under complete or Bernoulli randomization the true score is known by design and does not depend on the pretreatment covariates, say 0.5 for everyone, which gives perfect overlap and balance in expectation, and an attainable AUC of 0.5. (Under blocked or stratified assignment it is known but constant only within a design stratum, and may differ across strata.) The design under which causal inference is easiest is therefore one whose assignment is least predictable from covariates. That is not a rule that lower AUC is better: it is a demonstration that discrimination is measuring something other than what the adjustment needs.

Exercise

1: fully structured. A study fits a propensity model and reports scores from 0.22 to 0.79, AUC 0.64, and post-adjustment standardized differences all below 0.04.

(a) Is the AUC a problem? (b) What do the other two figures tell you? (c) What remains unestablished?

Check: (a) no, but not because a low AUC is itself good, discrimination is not what a propensity model is for, so the figure is close to irrelevant on its own. What settles the overlap question here is the score range: nothing near zero or one. A low AUC accompanied by scores piled at the extremes would still be an overlap failure, and a low AUC from a model that simply omitted the covariates that drive assignment would be a worse problem than a high one; (b) the score range shows no units pressed against zero or one, and the standardized differences show the adjustment balanced the measured covariates; (c) whether those covariates suffice, which is unconfoundedness and is not addressed by any of these numbers.

2: partly structured. An analyst is choosing between two specifications. Model 1: AUC 0.71, worst standardized difference after weighting 0.03, no scores outside [0.08, 0.88]. Model 2: AUC 0.89, worst standardized difference 0.02, 11% of units with scores below 0.02.

(a) Which would you use? (b) What is Model 2's better balance figure worth here? (c) What would you report if forced to use Model 2?

Check: (a) Model 1. It achieves balance across a region where both arms are represented; (b) little, because balance computed over a sample containing units with no comparable counterparts describes an adjustment resting on extrapolation for those units; (c) the overlap failure, the effective sample size, the largest weights, and the restricted population the estimate actually describes after any trimming.

3: unstructured. A team building a churn-prevention evaluation writes: "We used gradient boosting for the propensity model because it substantially outperformed logistic regression on held-out AUC (0.91 versus 0.73). We also included last-month engagement, which was the strongest predictor of receiving the retention offer."

Assess the two decisions. Say what you would check, and what you would change.

Check: both decisions optimise prediction at the expense of the adjustment. The AUC improvement signals that the model has found near-deterministic assignment patterns, so overlap must be inspected before anything else, expect scores near zero and one, extreme weights, and a much reduced effective sample size. Last-month engagement is more serious: if the retention offer was triggered by it, it is a pretreatment cause and admissible, but if it was measured after the offer went out, it is post-treatment and conditioning on it removes part of the effect. Establish its timing first. The change: select the specification by post-adjustment balance on the original covariates, use the flexible model only if it balances better across a region of genuine overlap, and report balance, score distributions and effective sample size rather than AUC.

Warning

Overlap failure, imbalance, and what each requires reporting

Some findings are not problems to work around. They are the analysis telling you what it cannot do.

Overlap fails. A share of one arm sits at scores near zero or one, with no comparable units in the other arm. For those units there is no comparison to make, not a noisy one, none. No weighting scheme, no richer specification and no larger sample repairs it, because the data contain no counterparts to borrow from.

What you must do: report the share affected and the score regions involved, before reporting any estimate.

Trimming was used to restore overlap. Dropping units with extreme scores does stabilise the estimate. It also changes who the estimate is about: the answer now describes the trimmed population, not the one you set out to study.

What you must do: state the trimming rule, how many units it removed from each arm, and name the restricted population as the estimand. A trimmed analysis reported as though it answered the original question is the more serious error, because nothing in the output reveals it.

A covariate stays unbalanced after adjustment. The score was built to balance the measured covariates and did not manage it for this one.

What you must do: add flexibility for that covariate and refit, on balance, never on classification accuracy or on the effect estimate. If it remains unbalanced, say so; an adjustment that failed on a covariate you named as a confounder has not done its job.

And the one that never appears in a diagnostic. Every check here concerns the covariates you measured. Unconfoundedness is a claim about the ones you did not, and no balance table, overlap plot or effective sample size speaks to it. A flawless diagnostic report establishes that the measured covariates were handled well, and nothing whatever about whether they were the right ones.

Next step

Practice Propensity Scores

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.