Practice: Regression Adjustment in Experiments
Question
Recognition · Interpretation
In a randomized experiment, OLS is used to fit
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what OLS fits when the only regressor takes two values.
Hint 2: Concept cue
A line through two group means passes through both. What is its slope?
Method selection · Direct application · Explanation
A randomized trial of a reading intervention proposes:
Decide which covariates belong, state what the adjustment contributes, and say where the causal claim comes from.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Take each covariate and ask whether it was fixed before assignment.
Hint 2: Strategy cue
For any covariate the treatment could have changed, ask what conditioning on it removes.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
Admissible covariates. Baseline score and school are fixed before assignment and belong. Baseline score in particular should explain a large share of outcome variation, so it is the one most likely to sharpen the estimate.
Inadmissible covariate. Books read during term is measured after the intervention began, and a reading intervention plausibly changes it. It is post-treatment and must be dropped.
What including it would do. Books read is very likely a mediator, part of how the intervention works is by increasing reading. Conditioning on it removes that pathway, so
What adjustment contributes. Precision only. The arms are already comparable in expectation because of randomization, so the estimate should not move much; what changes is the residual variance and hence the standard error.
Where the causal claim comes from. The assignment mechanism. It would be equally causal with no covariates at all, using the raw difference in means, merely less precise. No covariate in this specification is carrying identification.
Matching the analysis to the design. With individual randomization, heteroskedasticity-robust standard errors are appropriate; conventional ones assume a homoskedasticity there is no reason to expect. Had assignment been blocked, the block indicators belong in the regression; had it been clustered, the errors must be clustered at the unit of assignment; unequal assignment probabilities would call for weights.
A complete answer does each of these:
- locates identification
- screens covariate timing
- bounds what adjustment buys
- matches analysis to design
Comparison · Evaluation
Which statement correctly distinguishes covariate adjustment in a randomized experiment from adjustment in an observational study?
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what would go wrong in each setting if the covariate set were wrong.
Hint 2: Concept cue
In one setting the cost is precision; in the other it is correctness.
Error diagnosis · Explanation · Evaluation
A report states:
Randomization was compromised for the first three weeks, when site staff allocated participants manually. We addressed this by controlling for site, week of enrolment, age, sex and baseline severity. The adjusted treatment effect remains significant, so we are confident in the causal interpretation.
Identify the error and say what the report should state instead.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what was supplying the causal interpretation before the breakdown, and whether it still is.
Hint 2: Concept cue
If the covariates are now doing the identifying, what assumption has the study quietly taken on?
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The error. Covariates are being credited with repairing a failed design. For the affected period, treatment was not assigned by a known mechanism, so nothing about the design licenses a causal reading, and adding controls does not restore it.
What the analysis has silently become. An observational study. The five covariates are now carrying identification, which means the estimate rests on unconfoundedness given site, week, age, sex and baseline severity. That assumption was never planned for, and it is implausible here: staff allocating manually may have steered patients on clinical judgement, urgency or perceived likely benefit, none of which is recorded. If any such consideration affects the outcome, the adjusted estimate remains confounded.
Why significance is not reassurance. A confounded estimate can be precisely estimated and significant. The p-value speaks to sampling variability, not to whether the comparison identifies an effect.
What the report should state. The breakdown, plainly, with the number of participants affected. The primary analysis restricted to the properly randomized period, where the design still supports the claim. The full-sample adjusted analysis presented as a secondary, observational estimate with its assumption stated explicitly. A sensitivity analysis for the manually allocated portion, asking how strong an unmeasured confounder would need to be to overturn the conclusion. And, if possible, an account from the staff of the criteria they actually used, which is evidence about the assignment mechanism that no statistical adjustment can substitute for.
A further point about the specification. Even where randomization held, the regression has to match the mechanism: site-level allocation makes site a design feature rather than a nuisance covariate, and errors would need clustering at the level assignment occurred. Controlling for site as though it were an ordinary covariate treats a design fact as a statistical adjustment.
A complete answer does each of these:
- locates identification
- screens covariate timing
- bounds what adjustment buys
- matches analysis to design
Transfer · Evaluation · Explanation
An engineering team runs a properly randomized A/B test of a new checkout flow, with purchase completion as the outcome. Their analyst writes:
The raw difference was +1.2pp but noisy. We added controls for account age, device type, and number of pages viewed in the session, which tightened the estimate to +1.4pp with a much smaller standard error. We report the adjusted figure.
Assess the analysis and say what you would change.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Check when each control was measured relative to the user seeing the variant.
Hint 2: Strategy cue
Ask where the large precision gain actually came from, and whether that is good news.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
Two of the three controls are fine. Account age and device type are fixed before the user enters the experiment and are admissible precision covariates.
Pages viewed in the session is not. It is measured during the session, after exposure to the new checkout flow, and the flow plausibly changes how many pages a user views. Conditioning on it removes part of the effect, if the new flow works partly by reducing the pages needed to complete a purchase, that pathway is being held fixed, and can induce association between the variant and unobserved user characteristics. Much of the apparent precision gain is likely coming from this variable, which makes the tightened standard error misleading rather than reassuring.
On reporting only the adjusted figure. Both should be reported. Showing +1.2pp raw beside the adjusted estimate lets a reader see how much of the conclusion depends on the modelling. Suppressing the raw difference hides exactly what most needs checking, particularly when a post-treatment control is in the specification.
Where the causal claim comes from. The randomization, which was sound. The raw +1.2pp was already a valid causal estimate; the covariates were only ever going to sharpen it.
What I would change. Drop pages viewed. Refit with account age and device type. Report the raw difference alongside the adjusted estimate. If session engagement is itself of interest, analyse it as a secondary outcome. A legitimate question, and a different one from the effect on purchase completion.
What the design requires here. Individual randomization with a binary outcome: robust standard errors, not conventional ones, since a binary outcome is heteroskedastic by construction. No block indicators or clustering are needed because assignment was neither blocked nor grouped, which should be stated rather than assumed.
A complete answer does each of these:
- locates identification
- screens covariate timing
- bounds what adjustment buys
- matches analysis to design
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.