Practice: Potential Outcomes and the Fundamental Problem

Recognition · Interpretation

A unit has Y i ( 0 ) = 7 and Y i ( 1 ) = 11 , and is assigned W i = 0 . What does the study observe for this unit, and what is its individual treatment effect?

2 hints available, least help first.

Hint 1: Retrieval cue

Apply Y i obs = W i Y i ( 1 ) + ( 1 − W i ) Y i ( 0 ) with W i = 0 .

Hint 2: Concept cue

Separate what the effect IS from what the study can compute.

Interpretation · Direct application · Explanation

Four units have the following stipulated potential outcomes and assignments.

Unit Y i ( 0 ) Y i ( 1 ) W i
120261
218210
325251
422300

(a) Write Y i obs for each unit and mark which potential outcome is unobserved.

(b) Compute the finite-sample average treatment effect τ S from the full table.

(c) Compute the difference in means τ ^ from the observed data alone, and explain why it differs from τ S .

(d) Unit 3 has τ 3 = 0 . State what a study of these four units could report about unit 3 specifically, and why.

(e) Name the quantity your answer to (c) targets and the rule you used to approximate it, and say which of the two the stipulated table lets you compute exactly.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

For each row, let the assignment select one column and strike out the other.

Hint 2: Strategy cue

Compute τ S from the full stipulated table, then recompute using only the surviving column, and compare what the two calculations had access to.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) Unit 1 is treated: Y obs = 26 , with Y 1 ( 0 ) = 20 unobserved. Unit 2 is control: 18 , with Y 2 ( 1 ) unobserved. Unit 3 is treated: 25 , with Y 3 ( 0 ) unobserved. Unit 4 is control: 22 , with Y 4 ( 1 ) unobserved.

(b) The individual effects are 6 , 3 , 0 , 8 , so τ S = 17 / 4 = 4.25 .

(c) Treated observed outcomes are 26 and 25 , giving Y ¯ 1 = 25.5 . Control observed outcomes are 18 and 22 , giving Y ¯ 0 = 20 . So τ ^ = 5.5 , against τ S = 4.25 . The two differ because the estimator compares different units rather than the same unit under both treatments: this assignment placed unit 4, whose Y ( 1 ) is high, in control, and unit 3, whose effect is zero, in treatment. Over all allowed assignments the difference in means averages to τ S ; on this one realisation it does not.

(d) Nothing about unit 3's effect. The study observes Y 3 ( 1 ) = 25 and nothing else for that unit; τ 3 = 0 is visible here only because the table stipulates Y 3 ( 0 ) , which a real study would not supply. No amount of data on the other three units identifies it.

(d) The target is the average treatment effect τ = 1 4 ∑ i ( Y i ( 1 ) − Y i ( 0 ) ) , a property of these four units. The rule is the difference in means τ ^ = Y ¯ 1 − Y ¯ 0 , applied to whichever units happened to be treated. Because the table stipulates both potential outcomes, τ can be computed exactly here, which is precisely what no real experiment permits, and why the estimator exists.

A complete answer does each of these:

  • selects observed outcome
  • marks counterfactual missing
  • individual effect not identified
  • names estimand

Comparison · Classification · Interpretation

A physiotherapy clinic reports four comparisons. For each, state whether it is a contrast between potential outcomes for the same units, and if it is not, say precisely which quantity is being compared instead.

(a) Mean mobility score of patients who received the new programme, minus the mean score of patients who received the standard programme, where the programme was assigned by a coin flip.

(b) Each patient's mobility score after the programme, minus that patient's score at intake.

(c) Mean score of patients who completed the full programme, minus the mean of those who dropped out partway.

(d) Mean score of patients at a clinic that adopted the programme, minus the mean at a clinic that did not, where the clinics chose for themselves.

For each, name the quantity the comparison targets, if any, and the rule being used to approximate it.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

For each comparison, ask what the second term would have to be for it to be Y i ( 0 ) for those same units.

Hint 2: Concept cue

Two of these compare different units; one compares the same units at different times; one is split by something the treatment itself influenced.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) A contrast between potential outcomes. Random assignment makes the treated group's observed mean an estimate of E [ Y ( 1 ) ] and the control group's an estimate of E [ Y ( 0 ) ] for the same population, so the difference estimates an average causal effect.

(b) Not a causal contrast. This compares the same patient at two times, so the second term is the intake score rather than Y i ( 0 ) , what the score would have been at follow-up without the programme. Those differ whenever anything else changes over the period, including natural recovery.

(c) Not a causal contrast, and the worst of the four. Completion is determined after treatment begins and is plausibly caused by how the patient was responding, so the groups are defined by a post-treatment variable. There is no well-defined Y ( 0 ) for 'this patient, untreated, in the completer group'.

(d) A contrast between groups that may differ for reasons other than the programme, because the clinics selected themselves. It compares E [ Y ∣ adopting clinic ] with E [ Y ∣ non-adopting clinic ] . It becomes a causal contrast only under an identification assumption about how clinics chose, which the design does not supply.

Targets and rules. A contrast between potential outcomes for the same units targets an average treatment effect, with a difference in means as its estimator. The others target something else entirely, a difference between populations, or a change over time, and no rule applied to them approximates a causal effect. Saying which quantity is wanted before which rule is applied is what keeps the two from merging into "the effect".

A complete answer does each of these:

  • selects observed outcome
  • marks counterfactual missing
  • individual effect not identified
  • names estimand

Error diagnosis · Explanation · Transfer

A colleague reports on a randomized trial of a tutoring programme:

Student A was tutored and scored 78. Student B was not tutored and scored 71. So tutoring was worth 7 points for Student A. Averaging these unit-level effects across all pairs gives the average treatment effect, and with a large enough sample we could report each student's personal gain.

There are three distinct errors in this passage. Identify each, say what the quantity being described actually is, and state what the trial can legitimately report about Student A.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Write down which potential outcomes belong to Student A, and which of them the trial observed.

Hint 2: Concept cue

Consider separately: whose outcomes are being subtracted, how the average effect is actually obtained, and whether more data could change either answer.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

First error: a between-unit difference is called a within-unit effect. Student A's effect is τ A = Y A ( 1 ) − Y A ( 0 ) , both terms belonging to Student A. The figure 7 is Y A obs − Y B obs , a comparison of two different students who differ in more than tutoring. It is not τ A , and nothing licenses treating it as such.

Second error: the averaging story is backwards. The difference in means is not an average of computed individual effects, because no individual effect was ever computed. It is a difference of two group averages, and randomization makes it unbiased for the average effect directly, not by averaging quantities that were each separately identified.

Third error: sample size is offered as the remedy. More students do not complete any existing student's row. Each additional unit contributes one observed potential outcome and one missing one. Sample size improves precision about an average; it never identifies an individual effect, because the obstacle is structural rather than statistical.

What the trial can report about Student A: her observed score under tutoring, Y A ( 1 ) = 78 , and nothing else about her personally. Any statement about what she would have scored untutored is an inference from other students, justified, if at all, by the design, and it is a statement about an average rather than about her.

Estimand against estimator. The report uses "the effect" for three different things: the quantity it wants (the average of Y i ( 1 ) − Y i ( 0 ) over students), the rule it applied (a difference between two students' scores), and the number that came out. Naming them separately is what exposes the error. The rule it used does not approximate the quantity it wants, because the two students are not each other's counterfactuals.

A complete answer does each of these:

  • selects observed outcome
  • marks counterfactual missing
  • individual effect not identified
  • names estimand

Transfer · Evaluation · Construction

A city introduces a cycling-infrastructure programme in six districts. Two years later a report states:

Districts with the new lanes saw cycling rise by 31%. Districts without them saw a rise of 9%. The programme therefore caused a 22-percentage-point increase in cycling, and each district that built lanes gained roughly 22 points from having done so.

Without using the words 'potential outcome', set out (a) what quantity the report has actually computed, (b) what quantity would answer the question it asks, (c) why the final clause about each district is a different and stronger claim than the one before it, and (d) what would have to be true about how districts were selected for the comparison to support even the weaker claim.

State the quantity the report would need in order to make its claim, and the rule it actually applied.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask what the lane-building districts would have done had they not built lanes, and whether anything in the report measures it.

Hint 2: Strategy cue

Treat districts as the units. Separate the claim about their average from the claim about each one, and ask what each would require.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) A difference between two group averages: the average change among districts that built lanes and the average change among those that did not. Both are averages over observed districts.

(b) The question asks what would have happened in the lane-building districts had they not built lanes. That quantity was never observed for those districts, and the non-building districts are a stand-in for it rather than a measurement of it.

(c) The 22-point figure is an average over districts. The claim that each district gained roughly 22 points asserts something about every individual district, which requires knowing what each would have done without lanes. A quantity no study of this kind produces. Districts almost certainly differ: a dense district with existing cycling culture and a car-dependent suburb need not respond alike, and the average can be composed of very different district-level changes, including some near zero.

(d) The comparison supports the weaker average claim only if the districts that built lanes would, absent the programme, have changed like those that did not. Since the city chose where to build, plausibly where cycling was already growing, or where campaigning was strongest, that is an assumption about the selection process, not something the data establish. Without it the 22 points mixes the effect of the lanes with whatever made those districts get lanes first.

What was wanted and what was done. The claim needs the average over districts of Y d ( 1 ) − Y d ( 0 ) : cycling with the lanes against cycling in the same districts without them. What was computed is a difference between districts that got lanes and districts that did not. A rule that approximates the wanted quantity only if the two groups would have moved alike, which is exactly what choosing where to build lanes makes doubtful.

A complete answer does each of these:

  • selects observed outcome
  • marks counterfactual missing
  • individual effect not identified
  • names estimand
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.