Linear Regression for Experimental Research

Least squares fits a line, or a hyperplane, by minimising squared residuals, and supplies coefficients with standard errors from which tests and intervals follow. Two things it does not supply: a causal reading, which comes from the design rather than the fit, and evidence of a good model, which R-squared does not measure. A randomized treatment effect can be entirely credible with a low R-squared.

Definition

The simple linear model is Y i = β 0 + β 1 X i + ε i , and ordinary least squares minimises ∑ i ( Y i − β 0 − β 1 X i ) 2 , giving

β ^ 1 = ∑ i ( X i − X ¯ ) ( Y i − Y ¯ ) ∑ i ( X i − X ¯ ) 2 , β ^ 0 = Y ¯ − β ^ 1 X ¯ .

Fitted values and residuals are Y ^ i = β ^ 0 + β ^ 1 X i and e i = Y i − Y ^ i , with residual variance estimate σ ^ 2 = ∑ i e i 2 / ( n − 2 ) . Coefficient inference uses t = β ^ j − β j , 0 S E ( β ^ j ) on the residual degrees of freedom. With S S E = ∑ i ( Y i − Y ^ i ) 2 and S S T = ∑ i ( Y i − Y ¯ ) 2 , the coefficient of determination is R 2 = 1 − S S E / S S T . In matrix form Y = X β + ε and, when X ⊤ X is invertible, β ^ = ( X ⊤ X ) − 1 X ⊤ Y , with Var ⁡ ( β ^ ∣ X ) = σ 2 ( X ⊤ X ) − 1 under homoskedastic uncorrelated errors.

Formal statement

β ^ 1 = ∑ i ( X i − X ¯ ) ( Y i − Y ¯ ) ∑ i ( X i − X ¯ ) 2 ; β ^ = ( X ⊤ X ) − 1 X ⊤ Y ; t j = β ^ j / S E ( β ^ j ) ; R 2 = 1 − S S E / S S T .

Assumptions and scope

  • A coefficient is an association. A causal reading requires a design or an identification argument, and the fit supplies neither.

  • R 2 measures in-sample fit. It is not evidence of correctness, and a low value is compatible with a credible and precisely estimated treatment effect.

  • Each coefficient is conditional on the other predictors in the model. Adding or removing a variable changes what the remaining coefficients mean.

  • The classical variance formula Var ⁡ ( β ^ ∣ X ) = σ 2 ( X ⊤ X ) − 1 assumes homoskedastic uncorrelated errors. Robust or design-based standard errors are appropriate when those assumptions fail.

  • β ^ = ( X ⊤ X ) − 1 X ⊤ Y requires X ⊤ X to be invertible; perfectly collinear predictors make the coefficients individually undefined.

  • Residual degrees of freedom are n minus the number of estimated coefficients, which is n − 2 for an intercept and one slope.

  • Extrapolating the fitted line beyond the range of the observed predictors asserts a relationship the data do not cover.

Worked material

Example

The same effect at two very different R-squareds

Two regressions from the same randomized trial of a tutoring programme, n = 400 .

Regression 1: outcome on treatment alone. τ ^ = 4.2 points, S E = 1.9 , R 2 = 0.012 .

Regression 2: outcome on treatment, baseline score, and school. τ ^ = 4.3 points, S E = 0.8 , R 2 = 0.68 .

What changed and what did not. The estimated effect barely moved, which randomization leads you to expect. The standard error more than halved, because baseline score explains much of the outcome variation that was previously in the residual. R 2 rose from near zero to 0.68.

What the R 2 jump does not mean. That the second regression is more causally trustworthy. Both estimate the same causal quantity, and both get their causal warrant from the same randomization. The second is merely more precise.

Reading Regression 1 on its own. An R 2 of 0.012 says the treatment accounts for about 1% of the variation in test scores. That is unsurprising, scores depend overwhelmingly on prior attainment, school and a hundred other things, and it is no argument against the estimate. The effect is real, causally identified, and small relative to everything else that moves the outcome. Those are compatible facts.

R 2 answers "how much of this outcome's variation does the model track?" Causal credibility answers "where did this comparison come from?" Neither bears on the other.

Non-example

Claims a fitted regression does not support

" R 2 = 0.89 , so the coefficients can be trusted." R 2 compares fitted to observed values. A confounded model can track an outcome closely while attributing its movement to the wrong variable.

" R 2 = 0.04 , so the treatment effect is meaningless." Most outcomes have many determinants. A treatment can have a real, precisely estimated causal effect while accounting for a small share of total variation.

Selecting the specification with the highest R 2 . Adding predictors never lowers R 2 , so maximising it selects the largest model rather than the best one, and if the choice is made after seeing the estimates, the reported uncertainty no longer describes the procedure.

Reading a coefficient without naming what is held fixed. "The effect of advertising is 1.85" omits that the comparison is between stores alike on the other predictors. Change the predictor set and the number changes meaning.

Using classical standard errors when the variance is not constant. Var ⁡ ( β ^ ∣ X ) = σ 2 ( X ⊤ X ) − 1 assumes homoskedastic uncorrelated errors; clustered or heteroskedastic data need robust or design-based alternatives.

Predicting outside the observed range of the predictors. The fitted line is estimated where the data are. Extending it asserts a relationship over a region the data do not cover.

Contrast

Explained variation against causal warrant

What R 2 reportsWhat a causal reading requires
Question answeredHow much outcome variation does the model track here?What would happen under an intervention?
Computed fromFitted versus observed valuesNothing in the output — it comes from the design
Improved byAdding predictors, alwaysRandomization, or an identification argument
High value impliesGood in-sample trackingNothing
Low value impliesMuch variation unexplainedNothing

Why the two get conflated. They appear in the same output, and "explains 71% of the variance" sounds like a claim about explanation in the ordinary sense. It is not; it is a claim about squared distances.

The clean separation. Ask where the comparison came from. If treatment was randomized, the coefficient is causal at any R 2 . If it was chosen by the units, the coefficient is an association at any R 2 . The fit statistic never moves the answer.

A useful case to hold onto. A well-run trial with R 2 = 0.01 gives a credible causal estimate that explains almost none of the outcome's variation. An observational regression with R 2 = 0.95 gives a precisely tracked surface with no causal content. Both are ordinary.

Where R 2 is genuinely useful. Comparing how much variation competing models track on the same data, and judging whether predictions are likely to be useful. A prediction question, not an estimation one.

Common errors

Common misconception

A high R-squared shows the model is correct and its coefficients can be trusted, while a low R-squared means the analysis has failed and its estimates are not worth reporting.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.