Practice: Inverse Probability Weighting

Recognition · Direct application

A control unit has estimated propensity score e ^ ( X ) = 0.80 . What is its weight in an ATE analysis?

2 hints available, least help first.

Hint 1: Retrieval cue

Ask what the probability of this unit's observed treatment was.

Hint 2: Concept cue

The weight is one over the probability of the treatment the unit actually got.

Direct application · Interpretation · Evaluation

Four units, with treatment W , outcome Y , and estimated propensity score e ^ :

Unit W Y e ^
11300.50
21200.25
30120.50
40180.75

Compute the Horvitz–Thompson estimate, and assess the weights.

Then suppose a fifth unit is added: a control with e ^ = 0.98 and Y = 25 . Compute its weight, and say what would change about the quantity you are estimating if you dropped that unit rather than reporting it.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Weight each unit by one over the probability of the treatment it received, then difference the arms.

Hint 2: Strategy cue

Compute each arm's weighted total first, then divide the difference by N .

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

Treated contributions. Unit 1: 30 / 0.50 = 60 . Unit 2: 20 / 0.25 = 80 . Sum = 140 . Control contributions. Unit 3: 12 / ( 1 − 0.50 ) = 12 / 0.50 = 24 . Unit 4: 18 / ( 1 − 0.75 ) = 18 / 0.25 = 72 . Sum = 96 . The estimate.

τ ^ I P W = 1 4 ( 140 − 96 ) = 44 4 = 11.

The weights. Treated: 2 and 4 . Controls: 2 and 4 . The largest is 4, and no score approaches 0 or 1, so nothing here indicates an overlap problem, though with four units the diagnostics are illustrative rather than informative. ** Unit 4 carries a weight of 4 because it looked 75% likely to be treated and was not; it is doing a quarter of the work in this small example. In a real sample the figures to report alongside the estimate are the maximum weight, the effective sample size, and weighted balance on the original covariates. What to report beside it. A weighted point estimate is not readable alone. Report weighted covariate balance, the maximum weight and the effective sample size, together with sensitivity to the propensity specification. With four units these are illustrative; in a real sample they are what decides whether the estimate means anything. The fifth unit. A control with e ^ = 0.98 carries weight 1 / ( 1 − 0.98 ) = 50 . One respondent standing in for fifty. It is a control who looked almost certain to be treated, and there are few such controls precisely because the covariates nearly determined assignment there. What dropping it would change. Trimming that unit does not clean up the estimate; it changes what is being estimated. The remaining sample is everyone with e ^ below the trimming threshold, so the answer becomes an average effect within that subpopulation, and the original ATE over all five units is no longer what the number describes. That is a legitimate move only if it is reported as one: state the rule, how many units it removed from each arm, and name the retained population as the estimand. Why this is not the same as capping it.** Setting the weight to, say, 10 keeps the unit in the sample and alters the estimator instead. The capped quantity has no tidy description as an effect for any stated population, which makes it harder to report honestly than trimming, not easier, and either way, the comparable units that were missing are still missing.

A complete answer does each of these:

  • forms weights correctly
  • reads extreme weights
  • restates estimand after trimming
  • reports weighting diagnostics

Comparison · Evaluation

Two weighted analyses, each on 1,000 units. A: max weight 5.6, effective sample size 890, standardized differences below 0.04. B: max weight 250, effective sample size 96, standardized differences below 0.02. Which is more trustworthy?

2 hints available, least help first.

Hint 1: Retrieval cue

Ask what an effective sample size of 96 out of 1,000 implies.

Hint 2: Concept cue

Where does a weight of 250 come from, and what does that say about comparable units?

Error diagnosis · Explanation · Evaluation

A report states:

Weights ranged from 1.1 to 640. We winsorised at the 99th percentile to control the variance. The weighted analysis shows a 14% mortality reduction, and weighted covariate balance was excellent (all standardized differences below 0.03).

Identify the error and say what should be reported instead. Say also whether winsorising and trimming are the same move, and what each one does to the quantity being estimated.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Work backwards from a weight of 640 to the propensity score that produced it.

Hint 2: Concept cue

Ask whether winsorising creates the comparable units that were missing.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The error. A maximum weight of 640 is being read as a variance problem. It is an overlap failure: a weight that size implies an estimated propensity near 0.0016, meaning some units had essentially no chance of the treatment status they had, and the data contain almost nothing about what comparable units would have done otherwise.

What winsorising does. It replaces the largest weights with smaller ones, which stabilises the variance and silently changes the estimand. The affected units are no longer represented as themselves. The comparison for those units still has no empirical basis; only the symptom has been treated.

Why the balance figure does not rescue it. Standardized differences below 0.03 are computed over the altered weighted sample. Balance cannot distinguish genuine comparison from model-based extrapolation, and it speaks only to the measured covariates in any case.

What should be reported. The propensity-score distributions by arm and the proportion of each arm beyond 0.05 and 0.95. The effective sample size, which with weights of this magnitude is likely a small fraction of the cohort. If units are to be removed, explicit trimming to a region of overlap, with the estimand restated as the effect within that subpopulation rather than the original ATE. Uncertainty estimates that account for the weights having been estimated. And a statement that unconfoundedness remains assumed throughout, with a sensitivity analysis, health-seeking behaviour plausibly drives both screening and mortality and is rarely measured.

What the report should carry. Weighted covariate balance, the maximum and tail weights, and the effective sample size before and after winsorising. A weight of 640 with an unreported effective sample size hides how few observations carry the estimate, and stabilising the variance does not restore the comparable units that were never there.

A complete answer does each of these:

  • forms weights correctly
  • reads extreme weights
  • reports weighting diagnostics

Transfer · Evaluation · Explanation

A market-research team reports:

Our panel under-represents rural respondents, so we weighted responses by the inverse of each region's sampling fraction. Some rural respondents received weights above 400. The weighted estimate of product preference is 38%, and the weighted regional composition now matches the census exactly.

Assess the analysis. Say what the large weights indicate and what you would report.

The team proposes two remedies: cap every weight at 100, or drop the respondents whose weights exceed 100. Say what each does to the quantity being estimated, and which you would report and how.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask how many distinct people a weight of 400 represents, and how many such people are in the sample.

Hint 2: Strategy cue

Map each survey term onto its counterpart: sampling fraction, composition matching, non-response.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The structure is the same. Weighting by the inverse of a sampling probability to build a pseudo-population that matches a target composition is the mechanism from this unit, wearing survey vocabulary. What weights above 400 indicate. A handful of rural respondents are standing in for hundreds of people each. The estimate for that segment rests on a few individuals' answers, so it will move substantially if any one of them differs. This is the same limitation as extreme propensity weights, and it is about the information available, not the arithmetic. Why matching the census exactly is not reassurance. Composition matching is the analogue of covariate balance: it shows the weighting achieved what it was asked to achieve on region. It says nothing about whether the few rural respondents resemble rural non-respondents in their product preferences. The analogue of unconfoundedness, and equally unverifiable from these data. What I would report. The effective sample size overall and within region, which is the honest measure of how much information the weighted estimate carries; the number of distinct rural respondents behind the rural estimate; the maximum weight; and an interval that accounts for the weighting rather than treating the weighted count as the sample size. If rural coverage is too thin to support a segment estimate, that should be stated rather than concealed by a composition that matches the census. The remedy is design. Recruiting more rural respondents addresses the problem; no reweighting of the existing panel does. Diagnostics for a survey weight too. Report the maximum weight, the effective sample size and balance on the weighting variables. A design weight and a propensity weight raise the same question, how much of the estimate rests on how few respondents, even though only one of them is estimated. Capping at 100. The high-weight respondents stay in the sample and their influence is reduced by fiat. That changes the estimator, buying stability at the price of bias, and what it now estimates has no clean description as the preference of any stated population. The census composition it was built to match is no longer reproduced either. It is the harder of the two to report honestly, because nothing in the output announces it. Dropping them. This changes the population rather than the estimator. The estimate becomes product preference among respondents whose weights fall below the threshold, which, given how the weights arose, means substantially excluding the rural segment the weighting existed to represent. That is defensible only when stated: the rule, the number dropped, and the retained population named as what the 38% now describes. What I would report. Neither as a repair. Both remedies address the arithmetic and neither creates the rural respondents the panel lacks. Report the unmodified weighted estimate with its effective sample size, the number of distinct rural respondents behind the rural figure, and the maximum weight, and if a trimmed figure is reported alongside it, report it as an estimate for the retained population rather than as a corrected national one.

A complete answer does each of these:

  • forms weights correctly
  • reads extreme weights
  • restates estimand after trimming
  • reports weighting diagnostics
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.