Regression Adjustment in Experiments

A regression of outcome on a treatment indicator reproduces the difference in means exactly. Adding pretreatment covariates can sharpen the estimate by absorbing outcome variation the treatment had nothing to do with. What adjustment cannot do is supply identification: randomization already did that, and where randomization is absent no covariate on the right-hand side restores it.

Definition

A regression of the outcome on an intercept and a binary treatment indicator, Y i = α + τ W i + ε i , yields by OLS the algebraic identity τ ^ = Y ¯ 1 − Y ¯ 0 . Additive adjustment fits Y i = α + τ W i + X i ⊤ β + ε i with X i pretreatment, absorbing outcome variation explained by X and typically reducing the standard error. Fully interacted adjustment fits Y i = α + τ W i + X i ⊤ β + W i X i ⊤ γ + ε i , allowing the covariate-outcome relationship to differ by arm; with covariates centred at the full-sample mean, τ is interpretable at that reference point and arm-specific fitted means can be averaged to estimate an average effect.

Formal statement

Y i = α + τ W i + ε i ⇒ τ ^ = Y ¯ 1 − Y ¯ 0 ; additive Y i = α + τ W i + X i ⊤ β + ε i ; interacted Y i = α + τ W i + X i ⊤ β + W i X i ⊤ γ + ε i .

Assumptions and scope

  • Only pretreatment covariates are admissible. Conditioning on a variable the treatment could have influenced removes part of the effect or opens a non-causal path, and no gain in precision offsets that.

  • Adjustment improves precision; it does not supply identification. The causal reading of τ ^ rests on the assignment mechanism in every specification.

  • The specification should be chosen before treatment effects are examined. Selecting covariates by the estimate they produce invalidates the reported uncertainty.

  • The analysis must respect the design: blocking, clustering and unequal assignment probabilities all change what a correct regression and standard error look like.

  • Standard errors must suit the design; heteroskedasticity-robust errors are usual for individually randomized experiments.

  • Adjustment on a covariate unrelated to the outcome need not improve precision, and each additional covariate costs a degree of freedom.

  • Report the unadjusted difference alongside any adjusted estimate, so the contribution of the modelling is visible rather than buried.

Worked material

Example

What adjustment moves, and what it does not

A randomized trial of a tutoring programme has 400 students. The outcome is an end-of-year test score; a baseline score from the previous year is available.

Unadjusted. Y ¯ 1 − Y ¯ 0 = 4.2 points, S E = 1.9 .

Adjusted for baseline score. τ ^ = 4.3 points, S E = 0.8 .

Read the two numbers. The estimate barely moved, which is what randomization leads you to expect: the arms were already comparable on baseline score, so removing its influence does not shift the comparison. The standard error more than halved, because baseline score explains much of the variation in end-of-year score, and that variation was previously sitting in the residual.

That is adjustment working exactly as advertised, same quantity, sharper estimate.

A contrasting case. Suppose the adjusted estimate had come back at 7.1 instead. Nothing has been corrected; something has been revealed. A shift that large says the arms differed substantially on baseline score, which in a properly randomized trial of 400 students is an unlucky draw worth reporting rather than quietly adjusting away. The pre-specified analysis still stands, and the discrepancy belongs in the write-up.

Why the unadjusted number is still reported. Presenting only the adjusted figure makes the modelling invisible. Showing both lets a reader see how much of the result depends on the specification, which is precisely the thing a reader cannot otherwise check.

Non-example

Adjustments that do not do what is claimed

Adjusting after a compromised randomization. If assignment was subverted, staff steering certain applicants, a broken allocation sequence, covariates do not restore it. The result is an adjusted estimate of a confounded comparison, and the assumption it now rests on is unconfoundedness, which the experiment was designed to avoid needing.

Adjusting for differential attrition. When dropout differs by arm, the units remaining are a treatment-affected selection. Controlling for baseline characteristics of the survivors does not undo that selection, because the selection operated on things the baseline does not capture.

Controlling for a mediator to 'isolate the direct effect'. Conditioning on a post-treatment variable does not generally identify a direct effect. It removes one pathway and can simultaneously induce association through unobserved common causes.

Choosing the covariates that most raise significance. A specification selected by the estimate it produces has a reported uncertainty that no longer describes the procedure actually followed.

Reading a large shift in the estimate as a correction. In a properly randomized trial, adjustment should barely move the estimate. A large movement is information about the draw, to be reported rather than absorbed.

Reporting only the adjusted figure. Suppressing the unadjusted difference hides how much of the finding is the modelling, which is exactly what a reader most needs to judge.

Contrast

Same regression, two different jobs

In a randomized experimentIn an observational study
What identifies the effectThe assignment mechanismThe covariates, under unconfoundedness
What covariates contributePrecisionIdentification
Are they optional?Yes — the unadjusted difference is already validNo — omitting a confounder biases the estimate
Expected effect on the estimateLittle movementMovement is expected and is the point
If the covariate set is wrongPrecision is suboptimalThe estimate is biased
Post-treatment covariatesInadmissibleInadmissible

Why the confusion is natural. The command is the same in every statistical package. Nothing in the syntax records which of these two situations you are in, and the output looks identical.

The consequential asymmetry. In an experiment, getting the covariate set wrong costs precision. In an observational study, getting it wrong costs correctness. A single habit applied to both makes the difference invisible, and it is the difference between a suboptimal analysis and a wrong answer.

Where the misconception bites hardest. A trial whose randomization broke down is not converted into an observational study that happens to be well adjusted. It becomes a study relying on an assumption nobody planned for, about covariates nobody chose with identification in mind. The correct report says the design failed and states what is now being assumed.

Common errors

Common misconception

Adding covariates to the regression corrects for problems in how treatment was assigned, so an adjusted estimate can be read causally even when randomization failed or was never performed.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.