Module 4 of 4 · Lesson 3 of 4
Inverse Probability Weighting
Inverse-probability weighting into a pseudo-population, and what extreme weights indicate about overlap.
What you will be able to do
Given estimated propensity scores, the learner can form inverse-probability weights, compute the weighted estimate, and diagnose what extreme weights indicate about the comparison.
Orientation
If a unit had almost no chance of the treatment it received, it stands in for many units like it. That is what a weight is, and why a weight of 640 is a warning rather than a number.
Computing the weighted estimate takes a few lines. The harder part is reading a handful of enormous weights as a finding rather than as a numerical problem to be tidied away: they report how little comparable data the study contains.
Intuition
Extreme weights, computed
A ten-unit sample, the weights it produces, and what the arithmetic shows about where the estimate comes from.
| Unit | Weight | Share of its arm | ||
|---|---|---|---|---|
| 1 | 1 | |||
| 2 | 1 | |||
| 3 | 1 | |||
| 4 | 1 | |||
| 5 | 1 |
The treated arm has five units and a total weight of
What that does to the estimate. The weighted treated mean is
so unit 4's outcome receives two-thirds of the weight. Suppose the other four treated units have outcomes near
The effective sample size quantifies this. Kish's formula gives
Five treated units carrying the information of roughly 1.35. That number, not the nominal count, is what the standard error should be read against.
Why trimming does not repair it. Dropping unit 4 removes the variance problem and changes the question: the estimand becomes an effect over units with
The diagnostic to report alongside any IPW estimate. Maximum weight, effective sample size, and weighted balance on the original covariates. A large maximum weight with a small ESS is an overlap failure written in the vocabulary of weights.
Definition
Reading the weights, and how the two forms differ
The canonical definition gives the Horvitz–Thompson estimator, the raw ATE weights, the stabilized variant and the Hájek form. This block covers how those weights behave.
What a weight represents. A treated unit with
Why that concentrates variance. A single unit with weight
Stabilized weights reduce this variability by replacing the numerators with marginal treatment probabilities. They change the variance, not the overlap: a covariate region containing no controls still contains none after stabilization.
Horvitz–Thompson against Hájek. The Horvitz–Thompson weights are not forced to sum to the arm's sample size, so its estimate can fall outside the range of the observed outcomes. The Hájek form normalizes within each arm and is usually better behaved in finite samples. The two are not algebraically identical, though both target the ATE under the same conditions.
What weighting does not supply. Consistency, unconfoundedness and overlap are assumptions the estimator acts under; none of them is established by weighting.
Example
Two weight distributions
Two studies both estimate an effect by inverse-probability weighting on 1,000 units.
Study A. Estimated propensity scores run from 0.18 to 0.79. The largest weight is 5.6. The effective sample size is 890. After weighting, every standardized difference is below 0.04.
Study B. Scores run from 0.004 to 0.991. The largest weight is 250. The effective sample size is 96. After weighting, every standardized difference is below 0.02.
Which estimate would you trust? Study A, despite its slightly worse balance figures.
Study B's balance looks better, but its effective sample size says the estimate rests on the equivalent of 96 equally weighted observations, not 1,000. A single unit carrying a weight of 250 contributes a quarter of that. Move that one outcome and the answer moves with it.
What Study B's numbers are reporting. Not instability to be smoothed, but the absence of comparable units. Scores of 0.004 mean units who had essentially no chance of the treatment they received, and nothing in the data describes what similar units would have done under the other condition.
What trimming would do. Dropping units below 0.02 and above 0.98 would shrink the largest weight and raise the effective sample size, and the estimate would then describe the subpopulation that remains, not the original 1,000. That is a legitimate move, and it is a change of question, not a repair.
Worked example
Weights from a small table
Problem. Six units, with treatment
| Unit | |||
|---|---|---|---|
| 1 | 1 | 20 | 0.50 |
| 2 | 1 | 24 | 0.25 |
| 3 | 1 | 18 | 0.80 |
| 4 | 0 | 14 | 0.40 |
| 5 | 0 | 10 | 0.20 |
| 6 | 0 | 16 | 0.75 |
Compute the Horvitz–Thompson estimate and assess the weights.
Goal. A weighted estimate, and a judgement about whether it means anything.
Relevant principle. A treated unit is weighted by
Step 1: treated weights and contributions.
Reason: each treated unit stands in for
Step 2: control weights and contributions.
Reason: a control is weighted by
Step 3: assemble.
Step 4: inspect the weights. They are
Reason: the weight distribution, not the point estimate, says whether the comparison rests on data.
Result.
Check. The Hájek form normalizes within arm: treated
Interpretation. Report the estimate with the maximum weight, the effective sample size and weighted covariate balance. With six units these figures are illustrative only; with a real sample they are what decides whether the estimate is worth reporting at all.
Non-example
Things weighting does not accomplish
It does not repair unmeasured confounding. Every weight is a function of the estimated propensity score, which is a function of the measured covariates. A variable nobody recorded enters nowhere.
Capping weights does not restore overlap. Setting every weight above 10 to 10 makes the arithmetic well-behaved and reduces variance at the price of bias. It is a change to the estimator, not a clean move to a new target: the capped quantity has no simple description as the effect for some stated population, which is what makes it harder to report honestly than trimming, not easier. The units that had no comparable counterparts still have none.
Stabilized weights do not fix an overlap failure. They reduce variance under the corresponding normalization. A unit with
A well-balanced weighted sample is not sufficient. Balance on the measured covariates after weighting shows the procedure did its job on those covariates. Sufficiency of the covariate set is unconfoundedness, and no diagnostic addresses it.
Treating the weights as known understates uncertainty.
A large
Contrast
What a large weight indicates
| Read as a numerical problem | Read as an overlap failure | |
|---|---|---|
| What the weight means | An unstable computation | Few or no comparable units for this observation |
| Natural response | Cap, winsorise, or trim to stabilise | Ask whether the comparison is supported at all |
| Effect on the target | Estimator altered; no clean named target | Estimand restated explicitly, and reported |
| What gets reported | A point estimate with a smaller variance | Overlap diagnostics, effective sample size, restricted population |
| What the reader learns | That the analysis was tidy | Which units the estimate actually describes |
Why the first reading is so tempting. The symptom really is numerical, a variance that will not settle, an estimate that jumps when one observation changes, and every numerical tool offers a fix. The fix works on the symptom.
What the second reading supplies. It connects the weights back to the assumption they depend on. A propensity score of 0.004 means the data contain almost nothing about what units like this one would have done under the other treatment. That is not a defect of the estimator; it is a limit of the evidence.
The legitimate move. Trimming is a reasonable response, provided the estimand is restated. An analysis of the units with
Exercise
1: fully structured. A treated unit has
(a) What weight does each receive for an ATE analysis? (b) Which is more influential? (c) What does the second weight indicate about that unit?
Check: (a) the treated unit gets
2: partly structured. An analysis of 4,000 units reports a maximum weight of 380 and an effective sample size of 210.
(a) What does the effective sample size tell you? (b) Is capping the weights at 20 an adequate response? (c) What would you report?
Check: (a) the estimate carries roughly the information of 210 equally weighted observations, not 4,000, so its precision is far lower than the raw count suggests; (b) no, capping improves the variance and leaves the units without comparable counterparts exactly where they were. It is an estimator change with a bias cost, and the capped quantity answers no cleanly stated question; trimming to a declared overlap region would at least name the population the estimate describes; (c) the propensity-score distributions by arm, the proportion of units beyond 0.05 and 0.95, the effective sample size, and, if trimming is chosen, an explicit statement of the restricted population the estimate now describes.
3: unstructured. A public-health team weights a cohort to estimate the effect of a screening programme on five-year mortality. They write: "Weights ranged from 1.1 to 640. We winsorised at the 99th percentile to control variance. The weighted analysis shows a 14% mortality reduction, and weighted covariate balance was excellent (all standardized differences below 0.03)."
Assess the analysis and say what you would change.
Check: a maximum weight of 640 means some units had a propensity near 0.0016, essentially no chance of the screening status they had, so the data contain almost nothing about their counterfactual. Winsorising treats this as variance control. It trades bias for stability, and the resulting quantity is not the ATE and not the effect for any stated subpopulation either, unlike trimming, which at least names the units it keeps. The excellent balance figure is computed over that altered sample and cannot distinguish genuine comparison from extrapolation. What to change: report the propensity distributions by arm and the proportion in the tails; report the effective sample size, which with weights of this size is probably a small fraction of the cohort; replace winsorising with explicit trimming to a region of genuine overlap and restate the estimand as the effect in that subpopulation; use uncertainty estimates accounting for the estimated weights; and state that unconfoundedness remains assumed regardless, with a sensitivity analysis, since health-seeking behaviour plausibly drives both screening and mortality and is rarely measured.
What to carry forward
The estimator.
The weights.
The pseudo-population. A reweighted sample in which measured covariates no longer predict assignment.
Hájek form. Normalizing within arm; usually better behaved in finite samples than Horvitz–Thompson, which can fall outside the range of observed outcomes.
Stabilized weights. A variance device, not an overlap remedy.
Extreme weights are an overlap failure. They arise because comparable units are scarce, and capping them creates no missing comparison. Capping and winsorising alter the estimator, trading variance for bias; what they estimate is no longer the ATE and has no tidy name.
Trimming is legitimate, and it changes the question. State the retained subpopulation rather than continuing to claim the original ATE.
Diagnostics to report. Weighted covariate balance, maximum and tail weights, effective sample size, sensitivity to specification, and uncertainty that accounts for the weights being estimated.
The recurring error. Treating enormous weights as a numerical nuisance rather than as the study reporting the limits of its evidence.