Inverse Probability Weighting

What you will be able to do

Given estimated propensity scores, the learner can form inverse-probability weights, compute the weighted estimate, and diagnose what extreme weights indicate about the comparison.

Orientation

If a unit had almost no chance of the treatment it received, it stands in for many units like it. That is what a weight is, and why a weight of 640 is a warning rather than a number.

Computing the weighted estimate takes a few lines. The harder part is reading a handful of enormous weights as a finding rather than as a numerical problem to be tidied away: they report how little comparable data the study contains.

Intuition

Extreme weights, computed

A ten-unit sample, the weights it produces, and what the arithmetic shows about where the estimate comes from.

Unit W e ^ ( X ) WeightShare of its arm
11 0.50 2.0 8 %
21 0.40 2.5 10 %
31 0.50 2.0 8 %
41 0.02 50.0 67 %
51 0.60 1.67 7 %

The treated arm has five units and a total weight of 58.2 . Unit 4 alone carries 50 of it, two thirds of the arm, because its covariates made treatment very unlikely and it was treated anyway.

What that does to the estimate. The weighted treated mean is

∑ i w i Y i ∑ i w i ,

so unit 4's outcome receives two-thirds of the weight. Suppose the other four treated units have outcomes near 10 and unit 4 has Y = 30 : the weighted mean is about 23 , against an unweighted mean of 14 . Change unit 4's outcome to Y = 4 and the weighted mean falls to about 7 . One observation moves the arm mean by 16 units.

The effective sample size quantifies this. Kish's formula gives

ESS = ( ∑ i w i ) 2 ∑ i w i 2 = 58.2 2 2 2 + 2.5 2 + 2 2 + 50 2 + 1.67 2 ≈ 3387 2517 ≈ 1.35 .

Five treated units carrying the information of roughly 1.35. That number, not the nominal count, is what the standard error should be read against.

Why trimming does not repair it. Dropping unit 4 removes the variance problem and changes the question: the estimand becomes an effect over units with e ^ ( X ) above the trimming threshold, a population that excludes the covariate profiles unit 4 represented. The estimate becomes stable and is no longer answering what was originally asked, which must be stated rather than absorbed silently.

The diagnostic to report alongside any IPW estimate. Maximum weight, effective sample size, and weighted balance on the original covariates. A large maximum weight with a small ESS is an overlap failure written in the vocabulary of weights.

Definition

Reading the weights, and how the two forms differ

The canonical definition gives the Horvitz–Thompson estimator, the raw ATE weights, the stabilized variant and the Hájek form. This block covers how those weights behave.

What a weight represents. A treated unit with e ^ ( X i ) = 0.25 carries weight 1 / 0.25 = 4 : it stands for itself and for roughly three similar units that were not treated. A control unit with e ^ ( X i ) = 0.9 carries weight 1 / ( 1 − 0.9 ) = 10 . Large weights arise where treatment was unlikely and occurred, or likely and did not.

Why that concentrates variance. A single unit with weight 10 contributes as much to the estimate as ten units of weight 1 , so its outcome, and any measurement error in it, is amplified accordingly. Inspecting the largest weights is therefore a diagnostic: a few extreme weights indicate regions of the covariate space where one arm is nearly absent.

Stabilized weights reduce this variability by replacing the numerators with marginal treatment probabilities. They change the variance, not the overlap: a covariate region containing no controls still contains none after stabilization.

Horvitz–Thompson against Hájek. The Horvitz–Thompson weights are not forced to sum to the arm's sample size, so its estimate can fall outside the range of the observed outcomes. The Hájek form normalizes within each arm and is usually better behaved in finite samples. The two are not algebraically identical, though both target the ATE under the same conditions.

What weighting does not supply. Consistency, unconfoundedness and overlap are assumptions the estimator acts under; none of them is established by weighting.

Example

Two weight distributions

Two studies both estimate an effect by inverse-probability weighting on 1,000 units.

Study A. Estimated propensity scores run from 0.18 to 0.79. The largest weight is 5.6. The effective sample size is 890. After weighting, every standardized difference is below 0.04.

Study B. Scores run from 0.004 to 0.991. The largest weight is 250. The effective sample size is 96. After weighting, every standardized difference is below 0.02.

Which estimate would you trust? Study A, despite its slightly worse balance figures.

Study B's balance looks better, but its effective sample size says the estimate rests on the equivalent of 96 equally weighted observations, not 1,000. A single unit carrying a weight of 250 contributes a quarter of that. Move that one outcome and the answer moves with it.

What Study B's numbers are reporting. Not instability to be smoothed, but the absence of comparable units. Scores of 0.004 mean units who had essentially no chance of the treatment they received, and nothing in the data describes what similar units would have done under the other condition.

What trimming would do. Dropping units below 0.02 and above 0.98 would shrink the largest weight and raise the effective sample size, and the estimate would then describe the subpopulation that remains, not the original 1,000. That is a legitimate move, and it is a change of question, not a repair.

Worked example

Weights from a small table

Problem. Six units, with treatment W , observed outcome Y , and estimated propensity score e ^ :

Unit W Y e ^
11200.50
21240.25
31180.80
40140.40
50100.20
60160.75

Compute the Horvitz–Thompson estimate and assess the weights.

Goal. A weighted estimate, and a judgement about whether it means anything.

Relevant principle. A treated unit is weighted by 1 / e ^ ; a control by 1 / ( 1 − e ^ ) .

Step 1: treated weights and contributions.

Unit 1:  20 0.50 = 40 , Unit 2:  24 0.25 = 96 , Unit 3:  18 0.80 = 22.5 .

Reason: each treated unit stands in for 1 / e ^ units like it, so unit 2, treated despite a 25% chance, carries four times its own outcome.

Step 2: control weights and contributions.

Unit 4:  14 0.60 = 23.33 , Unit 5:  10 0.80 = 12.5 , Unit 6:  16 0.25 = 64.

Reason: a control is weighted by 1 / ( 1 − e ^ ) ; unit 6, with a 75% chance of treatment and untreated, represents many such units.

Step 3: assemble.

τ ^ I P W = 1 6 [ ( 40 + 96 + 22.5 ) − ( 23.33 + 12.5 + 64 ) ] = 158.5 − 99.83 6 = 58.67 6 ≈ 9.78 .

Step 4: inspect the weights. They are 2.0 ,   4.0 ,   1.25 for the treated and 1.67 ,   1.25 ,   4.0 for the controls. The largest is 4.0, and no score is near 0 or 1.

Reason: the weight distribution, not the point estimate, says whether the comparison rests on data.

Result. τ ^ I P W ≈ 9.78 , with well-behaved weights.

Check. The Hájek form normalizes within arm: treated 158.5 2.0 + 4.0 + 1.25 = 158.5 7.25 ≈ 21.86 ; control 99.83 1.67 + 1.25 + 4.0 = 99.83 6.92 ≈ 14.43 ; difference ≈ 7.43 . The two differ substantially at N = 6 , which is exactly the finite-sample divergence to expect, and a reason to prefer the Hájek form in small samples.

Interpretation. Report the estimate with the maximum weight, the effective sample size and weighted covariate balance. With six units these figures are illustrative only; with a real sample they are what decides whether the estimate is worth reporting at all.

Non-example

Things weighting does not accomplish

It does not repair unmeasured confounding. Every weight is a function of the estimated propensity score, which is a function of the measured covariates. A variable nobody recorded enters nowhere.

Capping weights does not restore overlap. Setting every weight above 10 to 10 makes the arithmetic well-behaved and reduces variance at the price of bias. It is a change to the estimator, not a clean move to a new target: the capped quantity has no simple description as the effect for some stated population, which is what makes it harder to report honestly than trimming, not easier. The units that had no comparable counterparts still have none.

Stabilized weights do not fix an overlap failure. They reduce variance under the corresponding normalization. A unit with e ^ = 0.002 remains a unit with almost no comparable counterparts.

A well-balanced weighted sample is not sufficient. Balance on the measured covariates after weighting shows the procedure did its job on those covariates. Sufficiency of the covariate set is unconfoundedness, and no diagnostic addresses it.

Treating the weights as known understates uncertainty. e ^ ( X ) was estimated, and inference that ignores this reports a standard error smaller than the procedure warrants.

A large N is not a large effective sample size. Ten thousand units with a few dominant weights can carry the information of a few hundred. The raw count is not the relevant number.

Contrast

What a large weight indicates

Read as a numerical problemRead as an overlap failure
What the weight meansAn unstable computationFew or no comparable units for this observation
Natural responseCap, winsorise, or trim to stabiliseAsk whether the comparison is supported at all
Effect on the targetEstimator altered; no clean named targetEstimand restated explicitly, and reported
What gets reportedA point estimate with a smaller varianceOverlap diagnostics, effective sample size, restricted population
What the reader learnsThat the analysis was tidyWhich units the estimate actually describes

Why the first reading is so tempting. The symptom really is numerical, a variance that will not settle, an estimate that jumps when one observation changes, and every numerical tool offers a fix. The fix works on the symptom.

What the second reading supplies. It connects the weights back to the assumption they depend on. A propensity score of 0.004 means the data contain almost nothing about what units like this one would have done under the other treatment. That is not a defect of the estimator; it is a limit of the evidence.

The legitimate move. Trimming is a reasonable response, provided the estimand is restated. An analysis of the units with e ^ ∈ [ 0.05 , 0.95 ] estimates an effect for that subpopulation, which may be exactly the practical question. What is not legitimate is trimming and continuing to describe the result as the average treatment effect over everyone.

Exercise

1: fully structured. A treated unit has e ^ = 0.20 ; a control unit has e ^ = 0.90 .

(a) What weight does each receive for an ATE analysis? (b) Which is more influential? (c) What does the second weight indicate about that unit?

Check: (a) the treated unit gets 1 / 0.20 = 5 ; the control gets 1 / ( 1 − 0.90 ) = 10 ; (b) the control, at double the weight; (c) it looked 90% likely to be treated and was not, so it stands in for many similar treated units, and if many controls look like this, the control arm is thinly populated where the treated arm is dense.

2: partly structured. An analysis of 4,000 units reports a maximum weight of 380 and an effective sample size of 210.

(a) What does the effective sample size tell you? (b) Is capping the weights at 20 an adequate response? (c) What would you report?

Check: (a) the estimate carries roughly the information of 210 equally weighted observations, not 4,000, so its precision is far lower than the raw count suggests; (b) no, capping improves the variance and leaves the units without comparable counterparts exactly where they were. It is an estimator change with a bias cost, and the capped quantity answers no cleanly stated question; trimming to a declared overlap region would at least name the population the estimate describes; (c) the propensity-score distributions by arm, the proportion of units beyond 0.05 and 0.95, the effective sample size, and, if trimming is chosen, an explicit statement of the restricted population the estimate now describes.

3: unstructured. A public-health team weights a cohort to estimate the effect of a screening programme on five-year mortality. They write: "Weights ranged from 1.1 to 640. We winsorised at the 99th percentile to control variance. The weighted analysis shows a 14% mortality reduction, and weighted covariate balance was excellent (all standardized differences below 0.03)."

Assess the analysis and say what you would change.

Check: a maximum weight of 640 means some units had a propensity near 0.0016, essentially no chance of the screening status they had, so the data contain almost nothing about their counterfactual. Winsorising treats this as variance control. It trades bias for stability, and the resulting quantity is not the ATE and not the effect for any stated subpopulation either, unlike trimming, which at least names the units it keeps. The excellent balance figure is computed over that altered sample and cannot distinguish genuine comparison from extrapolation. What to change: report the propensity distributions by arm and the proportion in the tails; report the effective sample size, which with weights of this size is probably a small fraction of the cohort; replace winsorising with explicit trimming to a region of genuine overlap and restate the estimand as the effect in that subpopulation; use uncertainty estimates accounting for the estimated weights; and state that unconfoundedness remains assumed regardless, with a sensitivity analysis, since health-seeking behaviour plausibly drives both screening and mortality and is rarely measured.

What to carry forward

The estimator. τ ^ I P W = 1 N ∑ i [ W i Y i e ^ i − ( 1 − W i ) Y i 1 − e ^ i ] .

The weights. 1 / e ^ for a treated unit, 1 / ( 1 − e ^ ) for a control. Each observed unit stands in for the similar units whose opposite outcome is missing.

The pseudo-population. A reweighted sample in which measured covariates no longer predict assignment.

Hájek form. Normalizing within arm; usually better behaved in finite samples than Horvitz–Thompson, which can fall outside the range of observed outcomes.

Stabilized weights. A variance device, not an overlap remedy.

Extreme weights are an overlap failure. They arise because comparable units are scarce, and capping them creates no missing comparison. Capping and winsorising alter the estimator, trading variance for bias; what they estimate is no longer the ATE and has no tidy name.

Trimming is legitimate, and it changes the question. State the retained subpopulation rather than continuing to claim the original ATE.

Diagnostics to report. Weighted covariate balance, maximum and tail weights, effective sample size, sensitivity to specification, and uncertainty that accounts for the weights being estimated.

The recurring error. Treating enormous weights as a numerical nuisance rather than as the study reporting the limits of its evidence.

Next step

Practice Inverse Probability Weighting

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.