Module 1 of 4 · Lesson 1 of 1

Potential Outcomes and the Fundamental Problem

Two potential outcomes per unit, and why assignment leaves only one of them observable.

What you will be able to do

Given a description of units, an assignment, and observed data, the learner can write each unit's potential outcomes and observed outcome, state which quantities are known and which are missing, and explain why no individual treatment effect can be computed from the data.

Orientation

Every unit has an outcome under treatment and an outcome without it. You observe exactly one of them, ever. Everything in causal inference follows from that.

That sounds like bookkeeping. It is the distinction the whole subject rests on: every design, estimator and test that follows is a way of standing in for a number that was never observed, and a learner who cannot see which number is missing cannot judge whether the substitute is credible.

Intuition

Two potential outcomes, one of them unobservable

Give every unit two sealed envelopes. One says what happens to it if treated; the other says what happens if it is not. Both are written before anyone decides anything. They are facts about the unit, not predictions.

Assigning a treatment opens one envelope and destroys the other.

The causal effect for that unit is the difference between the two numbers. It is well defined, and it is not observable for any unit in any study, at any sample size.

This is why causal statements are about averages. Not because averages are more interesting than individuals, but because the individual quantity is structurally unavailable and the average is not.

Definition

Potential outcomes, observed outcomes, and effects

For each unit i and a binary treatment, define two potential outcomes:

Y i ( 1 ) = the outcome if  i  is treated , Y i ( 0 ) = the outcome if  i  is not .

Both are defined for every unit, whichever treatment is actually assigned. The individual treatment effect is their difference:

τ i = Y i ( 1 ) − Y i ( 0 ) .

Let W i ∈ { 0 , 1 } record the assignment. The observed outcome is selected by it:

Y i obs = W i Y i ( 1 ) + ( 1 − W i ) Y i ( 0 ) .

When W i = 1 this returns Y i ( 1 ) ; when W i = 0 it returns Y i ( 0 ) . The other potential outcome is not merely unknown. It is not generated by the study at all.

The fundamental problem of causal inference is that τ i requires both halves and every study supplies one. Averages remain reachable: for a finite sample of N units,

τ S = 1 N ∑ i = 1 N [ Y i ( 1 ) − Y i ( 0 ) ] ,

and for a population, ATE = E [ Y ( 1 ) − Y ( 0 ) ] .

Example

A potential-outcomes table with unobservable entries

Six units, with the potential outcomes stipulated so the arithmetic is visible. In a real study the shaded half of this table does not exist.

Unit Y i ( 0 ) Y i ( 1 ) τ i W i Y i obs
110144114
212153012
39112111
415183015
511132011
613174117

The finite-sample average effect is τ S = ( 4 + 3 + 2 + 3 + 2 + 4 ) / 6 = 3 .

Now discard what a study never sees. The treated units are 1, 3, 6 with observed outcomes 14, 11, 17, giving Y ¯ 1 = 14 . The control units are 2, 4, 5 with 12, 15, 11, giving Y ¯ 0 ≈ 12.67 . The difference in means is τ ^ ≈ 1.33 .

The estimate is not 3, and nothing went wrong. This assignment happened to place units with high Y ( 0 ) in control. Across every allowed assignment the difference in means averages to τ S ; on any single one it need not. The gap is randomization error, not a flaw in the design and not evidence of bias.

Worked example

A four-patient trial with a sign-reversed estimate

Problem. A clinic enrols four patients in a trial of a new analgesic. The outcome is pain score after two hours, lower being better. For the purpose of this example the full schedule of potential outcomes is stipulated; a real trial would supply only one column.

Patient Y i ( 0 ) Y i ( 1 ) W i
A741
B650
C881
D520

Goal. Report what the trial observes, what it can estimate, and what it cannot.

Relevant principle. The assignment selects one potential outcome per unit; the other is not measured, not zero, and not recoverable from other units.

Step 1: apply the selection identity.
Y i obs = W i Y i ( 1 ) + ( 1 − W i ) Y i ( 0 ) .
Patient A: W = 1 , so Y obs = 4 , and Y A ( 0 ) = 7 goes unobserved.
Patient B: W = 0 , so Y obs = 6 , and Y B ( 1 ) goes unobserved.
Patient C: W = 1 , so Y obs = 8 , and Y C ( 0 ) goes unobserved.
Patient D: W = 0 , so Y obs = 5 , and Y D ( 1 ) goes unobserved.
Reason: the indicator picks a column; nothing about the other column is measured.

Step 2: compute the estimand from the stipulated table.
τ A = − 3 , τ B = − 1 , τ C = 0 , τ D = − 3 , so

τ S = − 3 − 1 + 0 − 3 4 = − 1.75 .

Reason: the finite-sample average effect needs both columns, which is exactly why it is unavailable in practice.

Step 3: compute what the trial can actually compute.
Treated: A and C, with observed 4 and 8 , so Y ¯ 1 = 6 .
Control: B and D, with observed 6 and 5 , so Y ¯ 0 = 5.5 .

τ ^ = 6 − 5.5 = 0.5 .

Reason: only the observed column is available, and the estimator compares groups rather than a unit with itself.

Result. The trial reports τ ^ = 0.5 while τ S = − 1.75 . The estimate has the wrong sign.

Check. Nothing was miscalculated. Patient C, for whom the drug does nothing, was assigned to treatment and has the worst pain score in the table; patient D, who benefits most, was assigned to control and never received the drug. The assignment happened to pair a high baseline with treatment.

Interpretation. With four units this is unremarkable: the difference in means is unbiased over repeated assignments, not correct on any one of them. The lesson is not that the estimator is broken but that a single small experiment carries real randomization error, and that τ S was never observable to check it against. Precision comes from units, not from better arithmetic.

Non-example

Four things that are not a potential-outcomes contrast

A before-and-after comparison. A clinic measures pain before the drug and two hours after, and reports the change as the effect. The two measurements are the same patient at two times, not the same patient under two treatments. Y i ( 0 ) is what the pain would have been at two hours without the drug, which is not the pain before taking it, pain drifts, resolves, and responds to rest. A before/after difference is a potential-outcomes contrast only under the additional assumption that nothing else changed, which is usually the assumption in question.

A variable measured after treatment. A trial reports the effect of the drug on pain among patients who reported no side effects. Side effects are caused by the treatment, so conditioning on them splits the sample by something the treatment determined. There is no Y i ( 0 ) for "this patient, untreated, in the no-side-effect group", because the group itself does not exist without the treatment.

An outcome that depends on another unit's assignment. In a vaccine trial in one household, an untreated person's infection risk falls when a housemate is vaccinated. Then Y i ( 0 ) is not well defined: the notation assumes a unit's outcome depends on its own assignment alone. This is a failure of the stable-treatment assumption, and it requires richer notation rather than more data.

A prediction from a model. A regression predicts what an untreated patient "would have scored" and the fitted value is called the counterfactual. The model output is an estimate of a conditional mean over units with similar covariates. It may be a reasonable stand-in, but it is not Y i ( 0 ) , and calling it so hides every assumption that makes the substitution credible.

Contrast

What the missing half is not

The missing potential outcome is not any of the things learners reach for to fill it.

Not the other group's outcome. Unit 1 was treated and scored 14. Unit 2 was a control and scored 12. The difference, 2, is not unit 1's effect and not unit 2's: it compares two different units, and τ 1 happens to be 4.

Not the other arm's mean. Substituting Y ¯ 0 for each treated unit's Y i ( 0 ) produces a number, and that number is the difference in means again. A sample average, carrying no information about any individual.

Not recoverable with more data. A thousand more units supply a thousand more half-rows. They sharpen the estimate of an average; they do not complete a single row already in hand.

Not a measurement problem. Better instruments measure what happened more precisely. No instrument measures what would have happened under a treatment that was not given.

Exercise

Work these in order. The first supplies most of the structure; the last supplies none.

1: fully structured. Three units, with the table given in full.

Unit Y i ( 0 ) Y i ( 1 ) W i
112151
29140
311111

(a) Write Y i obs for each unit. (b) Which potential outcome is missing for unit 2? (c) Compute τ S . (d) Compute τ ^ from the observed data.

Check: observed are 15 , 9 , 11 ; unit 2 is missing Y 2 ( 1 ) = 14 ; τ S = ( 3 + 5 + 0 ) / 3 ≈ 2.67 ; τ ^ = 13 − 9 = 4 .

2: partly structured. A study of six units reports Y ¯ 1 = 22 from four treated units and Y ¯ 0 = 19 from two control units. A reader asks what the effect was for the third treated unit.

Answer the reader, and say which quantity the study does support.

Check: the study supports τ ^ = 3 as an estimate of an average effect. It supports nothing about the third treated unit, whose Y ( 0 ) was never generated.

3: unstructured. A colleague proposes: "Run the experiment, then for every treated unit subtract the mean of the controls. That gives each treated unit's individual effect, and averaging them recovers the average effect."

The second half of the claim is arithmetically fine. Explain precisely what is wrong with the first half, and why the two halves can differ in validity even though the same subtraction appears in both.

Check: subtracting a group mean produces a number per unit, but that number estimates nothing about the unit, its Y ( 0 ) was never observed, and the control mean carries no unit-specific information. Averaging those numbers gives Y ¯ 1 − Y ¯ 0 , which is a valid estimator of the average effect. A valid average can be assembled from quantities that are individually meaningless.

What to carry forward

The two quantities. Every unit has Y i ( 1 ) and Y i ( 0 ) . Both are properties of the unit; both are defined whether or not the unit is assigned that treatment.

The effect. τ i = Y i ( 1 ) − Y i ( 0 ) , for that unit.

What a study sees. Y i obs = W i Y i ( 1 ) + ( 1 − W i ) Y i ( 0 ) . The assignment selects one; the other is never generated.

The fundamental problem. No individual effect is identified, in any study, at any sample size. This is structural, not statistical.

What remains reachable. Averages: τ S = 1 N ∑ i [ Y i ( 1 ) − Y i ( 0 ) ] for a fixed set of units, or ATE = E [ Y ( 1 ) − Y ( 0 ) ] for a population.

The three words to keep apart. The estimand is the quantity wanted; the estimator is the rule applied to the sample; the estimate is the number that rule returns. "The effect" names all three in careless writing and none of them precisely.

The recurring error. Filling a missing Y i ( 0 ) with the other arm's outcome, the other arm's mean, a pre-treatment measurement, or a model's fitted value, and then speaking about an individual.

Next step

Practice Potential Outcomes and the Fundamental Problem

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Randomized Assignment

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.