Potential Outcomes and the Fundamental Problem
What you will be able to do
Given a description of units, an assignment, and observed data, the learner can write each unit's potential outcomes and observed outcome, state which quantities are known and which are missing, and explain why no individual treatment effect can be computed from the data.
Orientation
Every unit has an outcome under treatment and an outcome without it. You observe exactly one of them, ever. Everything in causal inference follows from that.
That sounds like bookkeeping. It is the distinction the whole subject rests on: every design, estimator and test that follows is a way of standing in for a number that was never observed, and a learner who cannot see which number is missing cannot judge whether the substitute is credible.
Intuition
Two potential outcomes, one of them unobservable
Give every unit two sealed envelopes. One says what happens to it if treated; the other says what happens if it is not. Both are written before anyone decides anything. They are facts about the unit, not predictions.
Assigning a treatment opens one envelope and destroys the other.
The causal effect for that unit is the difference between the two numbers. It is well defined, and it is not observable for any unit in any study, at any sample size.
This is why causal statements are about averages. Not because averages are more interesting than individuals, but because the individual quantity is structurally unavailable and the average is not.
Definition
Potential outcomes, observed outcomes, and effects
For each unit
Both are defined for every unit, whichever treatment is actually assigned. The individual treatment effect is their difference:
Let
When
The fundamental problem of causal inference is that
and for a population,
Example
A potential-outcomes table with unobservable entries
Six units, with the potential outcomes stipulated so the arithmetic is visible. In a real study the shaded half of this table does not exist.
| Unit | |||||
|---|---|---|---|---|---|
| 1 | 10 | 14 | 4 | 1 | 14 |
| 2 | 12 | 15 | 3 | 0 | 12 |
| 3 | 9 | 11 | 2 | 1 | 11 |
| 4 | 15 | 18 | 3 | 0 | 15 |
| 5 | 11 | 13 | 2 | 0 | 11 |
| 6 | 13 | 17 | 4 | 1 | 17 |
The finite-sample average effect is
Now discard what a study never sees. The treated units are 1, 3, 6 with observed outcomes 14, 11, 17, giving
The estimate is not 3, and nothing went wrong. This assignment happened to place units with high
Worked example
A four-patient trial with a sign-reversed estimate
Problem. A clinic enrols four patients in a trial of a new analgesic. The outcome is pain score after two hours, lower being better. For the purpose of this example the full schedule of potential outcomes is stipulated; a real trial would supply only one column.
| Patient | |||
|---|---|---|---|
| A | 7 | 4 | 1 |
| B | 6 | 5 | 0 |
| C | 8 | 8 | 1 |
| D | 5 | 2 | 0 |
Goal. Report what the trial observes, what it can estimate, and what it cannot.
Relevant principle. The assignment selects one potential outcome per unit; the other is not measured, not zero, and not recoverable from other units.
Step 1: apply the selection identity.
Patient A:
Patient B:
Patient C:
Patient D:
Reason: the indicator picks a column; nothing about the other column is measured.
Step 2: compute the estimand from the stipulated table.
Reason: the finite-sample average effect needs both columns, which is exactly why it is unavailable in practice.
Step 3: compute what the trial can actually compute.
Treated: A and C, with observed
Control: B and D, with observed
Reason: only the observed column is available, and the estimator compares groups rather than a unit with itself.
Result. The trial reports
Check. Nothing was miscalculated. Patient C, for whom the drug does nothing, was assigned to treatment and has the worst pain score in the table; patient D, who benefits most, was assigned to control and never received the drug. The assignment happened to pair a high baseline with treatment.
Interpretation. With four units this is unremarkable: the difference in means is unbiased over repeated assignments, not correct on any one of them. The lesson is not that the estimator is broken but that a single small experiment carries real randomization error, and that
Non-example
Four things that are not a potential-outcomes contrast
A before-and-after comparison. A clinic measures pain before the drug and two hours after, and reports the change as the effect. The two measurements are the same patient at two times, not the same patient under two treatments.
A variable measured after treatment. A trial reports the effect of the drug on pain among patients who reported no side effects. Side effects are caused by the treatment, so conditioning on them splits the sample by something the treatment determined. There is no
An outcome that depends on another unit's assignment. In a vaccine trial in one household, an untreated person's infection risk falls when a housemate is vaccinated. Then
A prediction from a model. A regression predicts what an untreated patient "would have scored" and the fitted value is called the counterfactual. The model output is an estimate of a conditional mean over units with similar covariates. It may be a reasonable stand-in, but it is not
Contrast
What the missing half is not
The missing potential outcome is not any of the things learners reach for to fill it.
Not the other group's outcome. Unit 1 was treated and scored 14. Unit 2 was a control and scored 12. The difference, 2, is not unit 1's effect and not unit 2's: it compares two different units, and
Not the other arm's mean. Substituting
Not recoverable with more data. A thousand more units supply a thousand more half-rows. They sharpen the estimate of an average; they do not complete a single row already in hand.
Not a measurement problem. Better instruments measure what happened more precisely. No instrument measures what would have happened under a treatment that was not given.
Exercise
Work these in order. The first supplies most of the structure; the last supplies none.
1: fully structured. Three units, with the table given in full.
| Unit | |||
|---|---|---|---|
| 1 | 12 | 15 | 1 |
| 2 | 9 | 14 | 0 |
| 3 | 11 | 11 | 1 |
(a) Write
Check: observed are
2: partly structured. A study of six units reports
Answer the reader, and say which quantity the study does support.
Check: the study supports
3: unstructured. A colleague proposes: "Run the experiment, then for every treated unit subtract the mean of the controls. That gives each treated unit's individual effect, and averaging them recovers the average effect."
The second half of the claim is arithmetically fine. Explain precisely what is wrong with the first half, and why the two halves can differ in validity even though the same subtraction appears in both.
Check: subtracting a group mean produces a number per unit, but that number estimates nothing about the unit, its
What to carry forward
The two quantities. Every unit has
The effect.
What a study sees.
The fundamental problem. No individual effect is identified, in any study, at any sample size. This is structural, not statistical.
What remains reachable. Averages:
The three words to keep apart. The estimand is the quantity wanted; the estimator is the rule applied to the sample; the estimate is the number that rule returns. "The effect" names all three in careless writing and none of them precisely.
The recurring error. Filling a missing