Practice: Potential Outcomes and the Fundamental Problem
Question
Recognition · Interpretation
A unit has
2 hints available, least help first.
Hint 1: Retrieval cue
Apply
Hint 2: Concept cue
Separate what the effect IS from what the study can compute.
Interpretation · Direct application · Explanation
Four units have the following stipulated potential outcomes and assignments.
| Unit | |||
|---|---|---|---|
| 1 | 20 | 26 | 1 |
| 2 | 18 | 21 | 0 |
| 3 | 25 | 25 | 1 |
| 4 | 22 | 30 | 0 |
(a) Write
(b) Compute the finite-sample average treatment effect
(c) Compute the difference in means
(d) Unit 3 has
(e) Name the quantity your answer to (c) targets and the rule you used to approximate it, and say which of the two the stipulated table lets you compute exactly.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
For each row, let the assignment select one column and strike out the other.
Hint 2: Strategy cue
Compute
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) Unit 1 is treated:
(b) The individual effects are
(c) Treated observed outcomes are
(d) Nothing about unit 3's effect. The study observes
(d) The target is the average treatment effect
A complete answer does each of these:
- selects observed outcome
- marks counterfactual missing
- individual effect not identified
- names estimand
Comparison · Classification · Interpretation
A physiotherapy clinic reports four comparisons. For each, state whether it is a contrast between potential outcomes for the same units, and if it is not, say precisely which quantity is being compared instead.
(a) Mean mobility score of patients who received the new programme, minus the mean score of patients who received the standard programme, where the programme was assigned by a coin flip.
(b) Each patient's mobility score after the programme, minus that patient's score at intake.
(c) Mean score of patients who completed the full programme, minus the mean of those who dropped out partway.
(d) Mean score of patients at a clinic that adopted the programme, minus the mean at a clinic that did not, where the clinics chose for themselves.
For each, name the quantity the comparison targets, if any, and the rule being used to approximate it.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
For each comparison, ask what the second term would have to be for it to be
Hint 2: Concept cue
Two of these compare different units; one compares the same units at different times; one is split by something the treatment itself influenced.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) A contrast between potential outcomes. Random assignment makes the treated group's observed mean an estimate of
(b) Not a causal contrast. This compares the same patient at two times, so the second term is the intake score rather than
(c) Not a causal contrast, and the worst of the four. Completion is determined after treatment begins and is plausibly caused by how the patient was responding, so the groups are defined by a post-treatment variable. There is no well-defined
(d) A contrast between groups that may differ for reasons other than the programme, because the clinics selected themselves. It compares
Targets and rules. A contrast between potential outcomes for the same units targets an average treatment effect, with a difference in means as its estimator. The others target something else entirely, a difference between populations, or a change over time, and no rule applied to them approximates a causal effect. Saying which quantity is wanted before which rule is applied is what keeps the two from merging into "the effect".
A complete answer does each of these:
- selects observed outcome
- marks counterfactual missing
- individual effect not identified
- names estimand
Error diagnosis · Explanation · Transfer
A colleague reports on a randomized trial of a tutoring programme:
Student A was tutored and scored 78. Student B was not tutored and scored 71. So tutoring was worth 7 points for Student A. Averaging these unit-level effects across all pairs gives the average treatment effect, and with a large enough sample we could report each student's personal gain.
There are three distinct errors in this passage. Identify each, say what the quantity being described actually is, and state what the trial can legitimately report about Student A.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Write down which potential outcomes belong to Student A, and which of them the trial observed.
Hint 2: Concept cue
Consider separately: whose outcomes are being subtracted, how the average effect is actually obtained, and whether more data could change either answer.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
First error: a between-unit difference is called a within-unit effect. Student A's effect is
Second error: the averaging story is backwards. The difference in means is not an average of computed individual effects, because no individual effect was ever computed. It is a difference of two group averages, and randomization makes it unbiased for the average effect directly, not by averaging quantities that were each separately identified.
Third error: sample size is offered as the remedy. More students do not complete any existing student's row. Each additional unit contributes one observed potential outcome and one missing one. Sample size improves precision about an average; it never identifies an individual effect, because the obstacle is structural rather than statistical.
What the trial can report about Student A: her observed score under tutoring,
Estimand against estimator. The report uses "the effect" for three different things: the quantity it wants (the average of
A complete answer does each of these:
- selects observed outcome
- marks counterfactual missing
- individual effect not identified
- names estimand
Transfer · Evaluation · Construction
A city introduces a cycling-infrastructure programme in six districts. Two years later a report states:
Districts with the new lanes saw cycling rise by 31%. Districts without them saw a rise of 9%. The programme therefore caused a 22-percentage-point increase in cycling, and each district that built lanes gained roughly 22 points from having done so.
Without using the words 'potential outcome', set out (a) what quantity the report has actually computed, (b) what quantity would answer the question it asks, (c) why the final clause about each district is a different and stronger claim than the one before it, and (d) what would have to be true about how districts were selected for the comparison to support even the weaker claim.
State the quantity the report would need in order to make its claim, and the rule it actually applied.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what the lane-building districts would have done had they not built lanes, and whether anything in the report measures it.
Hint 2: Strategy cue
Treat districts as the units. Separate the claim about their average from the claim about each one, and ask what each would require.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a) A difference between two group averages: the average change among districts that built lanes and the average change among those that did not. Both are averages over observed districts.
(b) The question asks what would have happened in the lane-building districts had they not built lanes. That quantity was never observed for those districts, and the non-building districts are a stand-in for it rather than a measurement of it.
(c) The 22-point figure is an average over districts. The claim that each district gained roughly 22 points asserts something about every individual district, which requires knowing what each would have done without lanes. A quantity no study of this kind produces. Districts almost certainly differ: a dense district with existing cycling culture and a car-dependent suburb need not respond alike, and the average can be composed of very different district-level changes, including some near zero.
(d) The comparison supports the weaker average claim only if the districts that built lanes would, absent the programme, have changed like those that did not. Since the city chose where to build, plausibly where cycling was already growing, or where campaigning was strongest, that is an assumption about the selection process, not something the data establish. Without it the 22 points mixes the effect of the lanes with whatever made those districts get lanes first.
What was wanted and what was done. The claim needs the average over districts of
A complete answer does each of these:
- selects observed outcome
- marks counterfactual missing
- individual effect not identified
- names estimand
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.