Potential Outcomes and the Fundamental Problem

Each unit has an outcome under treatment and an outcome under control; the effect for that unit is their difference. Exactly one of the two is ever observed, so no individual effect is ever computed. Every method in causal inference is a way of replacing the missing half with a credible comparison, and every such method is a claim about an average rather than about a person.

Definition

For each unit i define Y i ( 1 ) , the outcome that would be observed if i received treatment, and Y i ( 0 ) , the outcome that would be observed if i received control. The individual treatment effect is τ i = Y i ( 1 ) − Y i ( 0 ) . With treatment indicator W i ∈ { 0 , 1 } , the observed outcome is Y i obs = W i Y i ( 1 ) + ( 1 − W i ) Y i ( 0 ) , which selects one potential outcome and discards the other.

Formal statement

τ i = Y i ( 1 ) − Y i ( 0 ) ; Y i obs = W i Y i ( 1 ) + ( 1 − W i ) Y i ( 0 ) ; τ S = 1 N ∑ i = 1 N [ Y i ( 1 ) − Y i ( 0 ) ] for a finite sample of N units, and ATE = E [ Y ( 1 ) − Y ( 0 ) ] for a population.

Assumptions and scope

  • Potential outcomes are defined for every unit whether or not that treatment is assigned. Y i ( 1 ) is a property of unit i and the treatment, not of the experiment that happened to be run; writing it does not assume the unit was treated.

  • The notation assumes a unit's potential outcomes depend on its own assignment alone, and that the treatment has one well-defined version. Interference between units, or several hidden versions of the same treatment, makes Y i ( 1 ) ambiguous and requires richer notation before anything below applies.

  • The finite-sample effect τ S and the population ATE are different estimands. Which one an experiment targets depends on whether the units are treated as the population of interest or as a sample drawn from one, and the two carry different variances even when they share an estimator.

  • The fundamental problem is about identification, not about sample size. Collecting more units never reveals a missing counterfactual for any unit already in hand; it only improves an average over units.

Forms this is expressed in

The same content in several forms. Each makes something visible that the others leave implicit, so moving between them is part of understanding the topic rather than a presentation choice.

tabular

The science table: every unit's two potential outcomes side by side, with the assignment and the resulting observed outcome. In any real study exactly one of the two outcome columns is available per row, and this representation exists to make that fact visible rather than to suggest both are obtainable.

Unit Y i ( 0 ) Y i ( 1 ) τ i W i Y i obs Unobserved
110144114 Y 1 ( 0 )
212153012 Y 2 ( 1 )
39112111 Y 3 ( 0 )
415183015 Y 4 ( 1 )
511132011 Y 5 ( 1 )
613174117 Y 6 ( 0 )

Reading across a row gives one unit's causal effect and is impossible in practice. Reading down the observed column gives what a study collects. The whole of experimental design concerns how much the second can say about the first.

Worked material

Example

A potential-outcomes table with unobservable entries

Six units, with the potential outcomes stipulated so the arithmetic is visible. In a real study the shaded half of this table does not exist.

Unit Y i ( 0 ) Y i ( 1 ) τ i W i Y i obs
110144114
212153012
39112111
415183015
511132011
613174117

The finite-sample average effect is τ S = ( 4 + 3 + 2 + 3 + 2 + 4 ) / 6 = 3 .

Now discard what a study never sees. The treated units are 1, 3, 6 with observed outcomes 14, 11, 17, giving Y ¯ 1 = 14 . The control units are 2, 4, 5 with 12, 15, 11, giving Y ¯ 0 ≈ 12.67 . The difference in means is τ ^ ≈ 1.33 .

The estimate is not 3, and nothing went wrong. This assignment happened to place units with high Y ( 0 ) in control. Across every allowed assignment the difference in means averages to τ S ; on any single one it need not. The gap is randomization error, not a flaw in the design and not evidence of bias.

Non-example

Four things that are not a potential-outcomes contrast

A before-and-after comparison. A clinic measures pain before the drug and two hours after, and reports the change as the effect. The two measurements are the same patient at two times, not the same patient under two treatments. Y i ( 0 ) is what the pain would have been at two hours without the drug, which is not the pain before taking it, pain drifts, resolves, and responds to rest. A before/after difference is a potential-outcomes contrast only under the additional assumption that nothing else changed, which is usually the assumption in question.

A variable measured after treatment. A trial reports the effect of the drug on pain among patients who reported no side effects. Side effects are caused by the treatment, so conditioning on them splits the sample by something the treatment determined. There is no Y i ( 0 ) for "this patient, untreated, in the no-side-effect group", because the group itself does not exist without the treatment.

An outcome that depends on another unit's assignment. In a vaccine trial in one household, an untreated person's infection risk falls when a housemate is vaccinated. Then Y i ( 0 ) is not well defined: the notation assumes a unit's outcome depends on its own assignment alone. This is a failure of the stable-treatment assumption, and it requires richer notation rather than more data.

A prediction from a model. A regression predicts what an untreated patient "would have scored" and the fitted value is called the counterfactual. The model output is an estimate of a conditional mean over units with similar covariates. It may be a reasonable stand-in, but it is not Y i ( 0 ) , and calling it so hides every assumption that makes the substitution credible.

Contrast

What the missing half is not

The missing potential outcome is not any of the things learners reach for to fill it.

Not the other group's outcome. Unit 1 was treated and scored 14. Unit 2 was a control and scored 12. The difference, 2, is not unit 1's effect and not unit 2's: it compares two different units, and τ 1 happens to be 4.

Not the other arm's mean. Substituting Y ¯ 0 for each treated unit's Y i ( 0 ) produces a number, and that number is the difference in means again. A sample average, carrying no information about any individual.

Not recoverable with more data. A thousand more units supply a thousand more half-rows. They sharpen the estimate of an average; they do not complete a single row already in hand.

Not a measurement problem. Better instruments measure what happened more precisely. No instrument measures what would have happened under a treatment that was not given.

Common errors

Common misconception

The difference between a treated unit's outcome and a control unit's outcome is that unit's treatment effect, so with enough data the effect on each individual can be read off directly.

Related units

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.