Practice: Inverse Probability Weighting
Question
Recognition · Direct application
A control unit has estimated propensity score
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what the probability of this unit's observed treatment was.
Hint 2: Concept cue
The weight is one over the probability of the treatment the unit actually got.
Direct application · Interpretation · Evaluation
Four units, with treatment
| Unit | |||
|---|---|---|---|
| 1 | 1 | 30 | 0.50 |
| 2 | 1 | 20 | 0.25 |
| 3 | 0 | 12 | 0.50 |
| 4 | 0 | 18 | 0.75 |
Compute the Horvitz–Thompson estimate, and assess the weights.
Then suppose a fifth unit is added: a control with
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Weight each unit by one over the probability of the treatment it received, then difference the arms.
Hint 2: Strategy cue
Compute each arm's weighted total first, then divide the difference by
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
Treated contributions. Unit 1:
The weights. Treated:
A complete answer does each of these:
- forms weights correctly
- reads extreme weights
- restates estimand after trimming
- reports weighting diagnostics
Comparison · Evaluation
Two weighted analyses, each on 1,000 units. A: max weight 5.6, effective sample size 890, standardized differences below 0.04. B: max weight 250, effective sample size 96, standardized differences below 0.02. Which is more trustworthy?
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what an effective sample size of 96 out of 1,000 implies.
Hint 2: Concept cue
Where does a weight of 250 come from, and what does that say about comparable units?
Error diagnosis · Explanation · Evaluation
A report states:
Weights ranged from 1.1 to 640. We winsorised at the 99th percentile to control the variance. The weighted analysis shows a 14% mortality reduction, and weighted covariate balance was excellent (all standardized differences below 0.03).
Identify the error and say what should be reported instead. Say also whether winsorising and trimming are the same move, and what each one does to the quantity being estimated.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Work backwards from a weight of 640 to the propensity score that produced it.
Hint 2: Concept cue
Ask whether winsorising creates the comparable units that were missing.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The error. A maximum weight of 640 is being read as a variance problem. It is an overlap failure: a weight that size implies an estimated propensity near 0.0016, meaning some units had essentially no chance of the treatment status they had, and the data contain almost nothing about what comparable units would have done otherwise.
What winsorising does. It replaces the largest weights with smaller ones, which stabilises the variance and silently changes the estimand. The affected units are no longer represented as themselves. The comparison for those units still has no empirical basis; only the symptom has been treated.
Why the balance figure does not rescue it. Standardized differences below 0.03 are computed over the altered weighted sample. Balance cannot distinguish genuine comparison from model-based extrapolation, and it speaks only to the measured covariates in any case.
What should be reported. The propensity-score distributions by arm and the proportion of each arm beyond 0.05 and 0.95. The effective sample size, which with weights of this magnitude is likely a small fraction of the cohort. If units are to be removed, explicit trimming to a region of overlap, with the estimand restated as the effect within that subpopulation rather than the original ATE. Uncertainty estimates that account for the weights having been estimated. And a statement that unconfoundedness remains assumed throughout, with a sensitivity analysis, health-seeking behaviour plausibly drives both screening and mortality and is rarely measured.
What the report should carry. Weighted covariate balance, the maximum and tail weights, and the effective sample size before and after winsorising. A weight of 640 with an unreported effective sample size hides how few observations carry the estimate, and stabilising the variance does not restore the comparable units that were never there.
A complete answer does each of these:
- forms weights correctly
- reads extreme weights
- reports weighting diagnostics
Transfer · Evaluation · Explanation
A market-research team reports:
Our panel under-represents rural respondents, so we weighted responses by the inverse of each region's sampling fraction. Some rural respondents received weights above 400. The weighted estimate of product preference is 38%, and the weighted regional composition now matches the census exactly.
Assess the analysis. Say what the large weights indicate and what you would report.
The team proposes two remedies: cap every weight at 100, or drop the respondents whose weights exceed 100. Say what each does to the quantity being estimated, and which you would report and how.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask how many distinct people a weight of 400 represents, and how many such people are in the sample.
Hint 2: Strategy cue
Map each survey term onto its counterpart: sampling fraction, composition matching, non-response.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The structure is the same. Weighting by the inverse of a sampling probability to build a pseudo-population that matches a target composition is the mechanism from this unit, wearing survey vocabulary. What weights above 400 indicate. A handful of rural respondents are standing in for hundreds of people each. The estimate for that segment rests on a few individuals' answers, so it will move substantially if any one of them differs. This is the same limitation as extreme propensity weights, and it is about the information available, not the arithmetic. Why matching the census exactly is not reassurance. Composition matching is the analogue of covariate balance: it shows the weighting achieved what it was asked to achieve on region. It says nothing about whether the few rural respondents resemble rural non-respondents in their product preferences. The analogue of unconfoundedness, and equally unverifiable from these data. What I would report. The effective sample size overall and within region, which is the honest measure of how much information the weighted estimate carries; the number of distinct rural respondents behind the rural estimate; the maximum weight; and an interval that accounts for the weighting rather than treating the weighted count as the sample size. If rural coverage is too thin to support a segment estimate, that should be stated rather than concealed by a composition that matches the census. The remedy is design. Recruiting more rural respondents addresses the problem; no reweighting of the existing panel does. Diagnostics for a survey weight too. Report the maximum weight, the effective sample size and balance on the weighting variables. A design weight and a propensity weight raise the same question, how much of the estimate rests on how few respondents, even though only one of them is estimated. Capping at 100. The high-weight respondents stay in the sample and their influence is reduced by fiat. That changes the estimator, buying stability at the price of bias, and what it now estimates has no clean description as the preference of any stated population. The census composition it was built to match is no longer reproduced either. It is the harder of the two to report honestly, because nothing in the output announces it. Dropping them. This changes the population rather than the estimator. The estimate becomes product preference among respondents whose weights fall below the threshold, which, given how the weights arose, means substantially excluding the rural segment the weighting existed to represent. That is defensible only when stated: the rule, the number dropped, and the retained population named as what the 38% now describes. What I would report. Neither as a repair. Both remedies address the arithmetic and neither creates the rural respondents the panel lacks. Report the unmodified weighted estimate with its effective sample size, the number of distinct rural respondents behind the rural figure, and the maximum weight, and if a trimmed figure is reported alongside it, report it as an estimate for the retained population rather than as a corrected national one.
A complete answer does each of these:
- forms weights correctly
- reads extreme weights
- restates estimand after trimming
- reports weighting diagnostics
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.