Inverse Probability Weighting
Weighting each unit by the inverse of its probability of receiving the treatment it got targets a pseudo-population in which assignment is unrelated to the measured covariates. Under the true propensity score and the identification assumptions that is a population-level result; in practice the weights are estimated, so the balance actually achieved in the weighted sample is something to check rather than assume. The arithmetic is simple and its failure mode is specific: when some units had almost no chance of their observed treatment, their weights explode, and an estimate that rests on a handful of heavily weighted observations is reporting an overlap problem rather than an effect.
Definition
With estimated propensity scores
The raw ATE weights are
The two are not algebraically identical in finite samples, though both target the ATE under suitable conditions.
Formal statement
Assumptions and scope
Identification requires consistency, unconfoundedness and overlap. Weighting acts on those assumptions and supplies none of them.
Extreme weights are an overlap failure reported in another vocabulary, not a numerical nuisance. Treating them as something to smooth away hides the condition that makes the estimate unreliable.
Trimming units with extreme scores changes the target population. The estimand must be restated as an effect over the retained subpopulation rather than reported as the original ATE.
The weights are estimated, so uncertainty estimates must account for the estimation of
; treating the weights as known understates the standard error. The Horvitz–Thompson form need not produce weights summing to the sample size in either arm, so it can fall outside the range of the observed outcomes. The Hájek form normalizes within arm and is usually better behaved in finite samples.
Stabilized weights reduce variance while preserving the target only under the corresponding normalization; they do not repair an overlap failure.
Balance must be assessed in the weighted sample, on the original covariates, alongside the maximum weight and the effective sample size.
Worked material
Example
Two weight distributions
Two studies both estimate an effect by inverse-probability weighting on 1,000 units.
Study A. Estimated propensity scores run from 0.18 to 0.79. The largest weight is 5.6. The effective sample size is 890. After weighting, every standardized difference is below 0.04.
Study B. Scores run from 0.004 to 0.991. The largest weight is 250. The effective sample size is 96. After weighting, every standardized difference is below 0.02.
Which estimate would you trust? Study A, despite its slightly worse balance figures.
Study B's balance looks better, but its effective sample size says the estimate rests on the equivalent of 96 equally weighted observations, not 1,000. A single unit carrying a weight of 250 contributes a quarter of that. Move that one outcome and the answer moves with it.
What Study B's numbers are reporting. Not instability to be smoothed, but the absence of comparable units. Scores of 0.004 mean units who had essentially no chance of the treatment they received, and nothing in the data describes what similar units would have done under the other condition.
What trimming would do. Dropping units below 0.02 and above 0.98 would shrink the largest weight and raise the effective sample size, and the estimate would then describe the subpopulation that remains, not the original 1,000. That is a legitimate move, and it is a change of question, not a repair.
Non-example
Things weighting does not accomplish
It does not repair unmeasured confounding. Every weight is a function of the estimated propensity score, which is a function of the measured covariates. A variable nobody recorded enters nowhere.
Capping weights does not restore overlap. Setting every weight above 10 to 10 makes the arithmetic well-behaved and reduces variance at the price of bias. It is a change to the estimator, not a clean move to a new target: the capped quantity has no simple description as the effect for some stated population, which is what makes it harder to report honestly than trimming, not easier. The units that had no comparable counterparts still have none.
Stabilized weights do not fix an overlap failure. They reduce variance under the corresponding normalization. A unit with
A well-balanced weighted sample is not sufficient. Balance on the measured covariates after weighting shows the procedure did its job on those covariates. Sufficiency of the covariate set is unconfoundedness, and no diagnostic addresses it.
Treating the weights as known understates uncertainty.
A large
Contrast
What a large weight indicates
| Read as a numerical problem | Read as an overlap failure | |
|---|---|---|
| What the weight means | An unstable computation | Few or no comparable units for this observation |
| Natural response | Cap, winsorise, or trim to stabilise | Ask whether the comparison is supported at all |
| Effect on the target | Estimator altered; no clean named target | Estimand restated explicitly, and reported |
| What gets reported | A point estimate with a smaller variance | Overlap diagnostics, effective sample size, restricted population |
| What the reader learns | That the analysis was tidy | Which units the estimate actually describes |
Why the first reading is so tempting. The symptom really is numerical, a variance that will not settle, an estimate that jumps when one observation changes, and every numerical tool offers a fix. The fix works on the symptom.
What the second reading supplies. It connects the weights back to the assumption they depend on. A propensity score of 0.004 means the data contain almost nothing about what units like this one would have done under the other treatment. That is not a defect of the estimator; it is a limit of the evidence.
The legitimate move. Trimming is a reasonable response, provided the estimand is restated. An analysis of the units with
Common errors
Common misconception
Extreme inverse-probability weights are a numerical instability to be fixed by trimming or capping, after which the estimate can be reported as before, and capping and trimming are interchangeable ways of doing it.
Related units
Requires
Connected
- Matching for Causal Inference (contrasts with)