Practice: Propensity Scores
Question
Recognition · Interpretation
The balancing property of the propensity score states that, within a level of
2 hints available, least help first.
Hint 1: Retrieval cue
Write out
Hint 2: Concept cue
Ask which quantities actually entered the model that produced the score.
Method selection · Direct application · Evaluation
Two propensity models are fitted to the same study.
Model A: scores from 0.14 to 0.83, arms overlapping throughout, all post-adjustment standardized differences below 0.06, AUC 0.68.
Model B: scores from 0.001 to 0.998 with 15% of controls below 0.01, all post-adjustment standardized differences below 0.02, AUC 0.94.
Choose one and justify the choice.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what the score is built to achieve, then ask which figures measure that.
Hint 2: Strategy cue
Consider what an AUC near 1 implies about the availability of comparable units.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
Choose Model A.
Why the AUC comparison does not favour Model B. The propensity score exists to balance covariates, not to predict treatment. An AUC of 0.94 says the covariates nearly determine assignment, which is a statement about how little comparable variation remains, bad news for adjustment, not good.
Why Model B's better balance figure is worth little. Its standardized differences are computed over a sample in which 15% of controls have almost no comparable treated units. Balance achieved across a region where one arm is essentially absent describes an adjustment resting on extrapolation for those units. A balance statistic cannot distinguish genuine comparison from model-based extension.
What Model A achieves. Balance below 0.06 on every covariate, across a score range where both arms are represented throughout. That is the adjustment doing its job on a comparison the data actually support.
What neither model establishes. That the measured covariates suffice. Both balance what they were given; unconfoundedness concerns what was not measured, and no diagnostic here speaks to it.
A prior question for both models. Whether each specification contains only pretreatment covariates. The balancing property is a claim about measured pretreatment variables; a score built with anything the treatment could have influenced does not support it, whatever its balance table shows.
A complete answer does each of these:
- states balancing property
- judges by balance not prediction
- bounds what it repairs
- respects covariate timing
Comparison · Evaluation
Which finding would be the strongest reason to revise a fitted propensity model?
2 hints available, least help first.
Hint 1: Retrieval cue
Ask which of these measures balance.
Hint 2: Concept cue
One option lets the outcome influence the assignment model. Find it.
Error diagnosis · Explanation · Evaluation
A report states:
We compared several specifications for the propensity model and selected the gradient-boosted version, which achieved the highest cross-validated AUC (0.93). This gives us confidence that treatment assignment has been accurately modelled and that the resulting adjustment is reliable.
Identify the error and say what should have driven the choice.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
State what the score is built to achieve, then ask whether AUC measures it.
Hint 2: Concept cue
Ask what an AUC near 1 implies about where comparable units can be found.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The error. The specification was selected by a prediction criterion. The propensity score is not a forecasting device: it exists so that units with similar scores are comparable on the measured covariates, and classification accuracy does not measure that.
Why a high AUC is a prompt to check, not a verdict. An AUC of 0.93 says the measured covariates separate the arms sharply. That often means scores pushed toward zero and one, where the opposite arm is thinly populated, but the AUC does not establish it, and it is the score distributions by arm and the post-adjustment balance that settle whether the comparison is supported. Selecting a specification for accuracy optimises a criterion unrelated to the adjustment; it does not by itself prove the overlap is bad, which is why the report should have shown the diagnostics that would.
The clearest illustration. Under complete or Bernoulli randomization the true propensity score is known by design and does not depend on the pretreatment covariates, so the best attainable AUC is 0.5. The design under which causal inference is easiest is one whose assignment is least predictable from covariates, which shows that discrimination measures something other than what the adjustment requires, not that a lower AUC is itself better.
What should have driven the choice. Covariate balance after adjustment, standardized differences and variance ratios on the original covariates and their important transformations, assessed alongside score overlap between the arms, the effective sample size, and the magnitude of the largest weights. A specification earns its place by balancing better across a region where both arms are represented.
What no specification search can deliver. Any confidence that the measured covariates suffice. That is unconfoundedness, and it is untouched by model selection of any kind.
What else to check in that specification. Which variables it used. A boosted model chasing AUC will take any predictive column, including ones recorded after treatment began. The score is defined on pretreatment covariates; fitting it with later information balances the arms on consequences of treatment rather than on their prior comparability.
A complete answer does each of these:
- states balancing property
- judges by balance not prediction
- bounds what it repairs
- respects covariate timing
Transfer · Evaluation · Explanation
A lending team evaluates whether a financial-literacy course reduces default. Enrolment was voluntary. They fit a model for enrolment and write:
The enrolment model reaches 91% accuracy on held-out data. We also added the customer's number of logins in the month after the course launched, which improved accuracy further. Because the model is so accurate, we are confident the matched comparison is sound.
Assess both decisions and say what you would check first.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Check each variable's timing relative to the treatment before assessing any model statistic.
Hint 2: Strategy cue
Ask separately: was this variable admissible, and was the model judged by the right criterion?
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The accuracy claim. Accuracy is the wrong criterion, and 91% is a reason for concern rather than confidence. It indicates the covariates nearly determine enrolment, so scores concentrate near zero and one and comparable customers become scarce exactly where they are most needed. The model should be judged by balance on the original covariates after matching, together with score overlap, the number of unmatched units, and the effective sample size.
The post-launch logins variable. This is the more serious error. Logins in the month after launch could have been influenced by the course itself, making it a post-treatment variable. Conditioning on it removes part of the effect under study and can open a non-causal path, biasing the comparison in a direction that is hard to predict. That it improved accuracy is precisely what makes it tempting and is irrelevant to whether it belongs; only pretreatment variables are admissible.
What I would check first. The timing of every variable in the model, since an inadmissible covariate invalidates the adjustment regardless of any other diagnostic. Then the score distributions by arm, then balance on the original covariates among matched units, then how many units went unmatched and which population the matched sample now represents.
What remains after all of it. Whether the measured covariates suffice. Financial sophistication and motivation plausibly drive both enrolment and default, and neither is in the data. A well-balanced, well-overlapping matched sample would still rest on that assumption, so a sensitivity analysis belongs in the report.
Before reading the model at all. Confirm it was fitted on information available before enrolment. Variables recorded after the course began, attendance, later account activity, would make the score describe consequences of enrolment, and adjusting on it would not deliver the comparability the method is used to obtain.
A complete answer does each of these:
- states balancing property
- judges by balance not prediction
- bounds what it repairs
- respects covariate timing
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.