Module 3 of 3 · Lesson 1 of 3

Linear Regression for Experimental Research

Regression coefficients and standard errors, and the causal and predictive claims they do not support.

What you will be able to do

Given regression output, the learner can state what a coefficient means as an average change per unit of the predictor, conditional on the others in the model and with its units, and treat R 2 as in-sample explained variation rather than evidence of correctness.

What you will be able to do

Given regression output and a described study, the learner can identify the design or identification argument, not the fit, as the source of any causal reading, and say when the classical standard errors are inappropriate and what should replace them.

Orientation

Regression output reports several different things, and the temptation is to read all of them as evidence the model is right. R 2 is not that. Neither is a small p-value.

Fitting is mechanical and has a closed form. The difficulty is refusing two upgrades the output invites: from association to cause, and from explained variation to a well-specified model.

Intuition

Reading a coefficient when another predictor is added

The canonical intuition states that a slope is an association holding the model's other predictors fixed, and that R 2 does not establish causal standing. This block works through what happens to a coefficient when the model changes.

A regression of monthly spending on advertising alone returns a slope of 4.20 per pound spent. Adding store footfall to the same model returns 1.60 for advertising and 0.85 per visitor.

Neither number is wrong, and they answer different questions. The first is the average difference in spending between months differing by one pound of advertising, with nothing held fixed. The second is that difference among months with equal footfall.

Which one is wanted depends on the question. If advertising works partly by drawing people into the store, then footfall is on the causal path, and conditioning on it removes part of the effect being estimated. If footfall is instead driven by seasonal trade that also affects advertising budgets, holding it fixed removes a confounding path and the second number is closer to what was wanted.

The regression cannot distinguish these two cases. Both produce the same drop from 4.20 to 1.60 , and the decision between them comes from knowing how the variables arose.

R 2 does not help here. Adding footfall raises it from 0.31 to 0.68 , and that rise occurs in both scenarios. A model can account for far more variation while estimating a quantity further from the one intended.

Definition

Least squares in matrix form, and the variance assumptions

The canonical definition gives the estimator, the residual variance, the t statistic, R 2 and the matrix form. This block covers the conditions attached to them.

The divisor in σ ^ 2 . The denominator n − 2 is n minus the number of estimated coefficients, so a model with p coefficients uses n − p . Using n would understate the residual variance, because the fitted line was chosen to make the residuals small.

What the t statistic assumes. The reference distribution requires the errors to be uncorrelated with constant variance, and either normal errors or a sample large enough for the central limit theorem. Clustered or serially correlated observations violate the first condition, and the reported standard error is then too small regardless of sample size.

What R 2 is a proportion of. S S T measures variation about the sample mean, so R 2 compares the model against predicting that mean. It does not compare the model against any alternative model, and it cannot fall when a predictor is added.

Where the variance formula comes from. Var ⁡ ( β ^ ∣ X ) = σ 2 ( X ⊤ X ) − 1 is derived from the model's error assumptions, so it is model-based. A randomized experiment supports an alternative derivation from the assignment mechanism, which does not require homoskedasticity. Where the two disagree, they are answering different questions about where the uncertainty comes from.

When X ⊤ X is not invertible. Exact collinearity among predictors makes the inverse undefined and no unique solution exists. Near-collinearity makes it defined and unstable, which is the condition the shrinkage unit addresses.

Example

The same effect at two very different R-squareds

Two regressions from the same randomized trial of a tutoring programme, n = 400 .

Regression 1: outcome on treatment alone. τ ^ = 4.2 points, S E = 1.9 , R 2 = 0.012 .

Regression 2: outcome on treatment, baseline score, and school. τ ^ = 4.3 points, S E = 0.8 , R 2 = 0.68 .

What changed and what did not. The estimated effect barely moved, which randomization leads you to expect. The standard error more than halved, because baseline score explains much of the outcome variation that was previously in the residual. R 2 rose from near zero to 0.68.

What the R 2 jump does not mean. That the second regression is more causally trustworthy. Both estimate the same causal quantity, and both get their causal warrant from the same randomization. The second is merely more precise.

Reading Regression 1 on its own. An R 2 of 0.012 says the treatment accounts for about 1% of the variation in test scores. That is unsurprising, scores depend overwhelmingly on prior attainment, school and a hundred other things, and it is no argument against the estimate. The effect is real, causally identified, and small relative to everything else that moves the outcome. Those are compatible facts.

R 2 answers "how much of this outcome's variation does the model track?" Causal credibility answers "where did this comparison come from?" Neither bears on the other.

Worked example

Reading a regression table

Problem. A regression of monthly revenue (in thousands) on advertising spend (in thousands) and a store-size indicator gives:

TermEstimateSE
Intercept12.43.1
Advertising1.850.42
Large store8.702.90

with n = 60 , R 2 = 0.71 . A manager concludes: "Advertising causes £1,850 of revenue per £1,000 spent, and with 71% explanatory power we can rely on it."

Goal. State what each number supports, and correct the conclusion.

Relevant principle. A coefficient is a conditional association; the causal reading comes from the design; R 2 measures in-sample fit.

Step 1: read the advertising coefficient. For two stores of the same size category, one spending £1,000 more per month is associated with £1,850 more revenue on average.

Reason: the coefficient is conditional on the other predictors, so the comparison is between stores alike on store size, and the units are those of the variables as entered.

Step 2: test it. t = 1.85 / 0.42 ≈ 4.40 on 60 − 3 = 57 residual degrees of freedom, comfortably beyond any conventional critical value.

Reason: the same template as any other test, applied to β ^ j against a null of zero.

Step 3: reject the causal claim. Advertising spend was chosen by the stores, not assigned. Stores expecting a strong month plausibly advertise more, so the coefficient mixes advertising's effect with whatever drove the spending decision.

Reason: nothing about least squares distinguishes a cause from a correlate; that distinction is a property of how the data arose.

Step 4: reject the R 2 claim. An R 2 of 0.71 says the two predictors track 71% of revenue variation in these 60 observations. It is not evidence the model is correct, that no confounder is missing, or that the coefficient is unbiased.

Reason: R 2 compares fitted to observed values; a confounded model can track the outcome closely and still misattribute the effect.

Step 5: check the standard errors. Revenue variance plausibly grows with store size, so homoskedasticity is doubtful, and robust standard errors are the safer choice here.

Result. A well-estimated conditional association of £1,850 per £1,000, with no causal warrant and an R 2 that speaks only to in-sample fit.

Check. Would a higher R 2 help the causal claim? No. A model with R 2 = 0.95 built the same way would have the same problem. Only a design change, randomizing advertising budgets across stores, or finding a natural experiment, addresses it.

Interpretation. Report the association with robust standard errors, state plainly that spending was not assigned, and propose a randomized budget trial if the causal question is the one that matters.

Non-example

Claims a fitted regression does not support

" R 2 = 0.89 , so the coefficients can be trusted." R 2 compares fitted to observed values. A confounded model can track an outcome closely while attributing its movement to the wrong variable.

" R 2 = 0.04 , so the treatment effect is meaningless." Most outcomes have many determinants. A treatment can have a real, precisely estimated causal effect while accounting for a small share of total variation.

Selecting the specification with the highest R 2 . Adding predictors never lowers R 2 , so maximising it selects the largest model rather than the best one, and if the choice is made after seeing the estimates, the reported uncertainty no longer describes the procedure.

Reading a coefficient without naming what is held fixed. "The effect of advertising is 1.85" omits that the comparison is between stores alike on the other predictors. Change the predictor set and the number changes meaning.

Using classical standard errors when the variance is not constant. Var ⁡ ( β ^ ∣ X ) = σ 2 ( X ⊤ X ) − 1 assumes homoskedastic uncorrelated errors; clustered or heteroskedastic data need robust or design-based alternatives.

Predicting outside the observed range of the predictors. The fitted line is estimated where the data are. Extending it asserts a relationship over a region the data do not cover.

Contrast

Explained variation against causal warrant

What R 2 reportsWhat a causal reading requires
Question answeredHow much outcome variation does the model track here?What would happen under an intervention?
Computed fromFitted versus observed valuesNothing in the output — it comes from the design
Improved byAdding predictors, alwaysRandomization, or an identification argument
High value impliesGood in-sample trackingNothing
Low value impliesMuch variation unexplainedNothing

Why the two get conflated. They appear in the same output, and "explains 71% of the variance" sounds like a claim about explanation in the ordinary sense. It is not; it is a claim about squared distances.

The clean separation. Ask where the comparison came from. If treatment was randomized, the coefficient is causal at any R 2 . If it was chosen by the units, the coefficient is an association at any R 2 . The fit statistic never moves the answer.

A useful case to hold onto. A well-run trial with R 2 = 0.01 gives a credible causal estimate that explains almost none of the outcome's variation. An observational regression with R 2 = 0.95 gives a precisely tracked surface with no causal content. Both are ordinary.

Where R 2 is genuinely useful. Comparing how much variation competing models track on the same data, and judging whether predictions are likely to be useful. A prediction question, not an estimation one.

Exercise

1: fully structured. A regression of yield on fertiliser (kg) gives β ^ 1 = 2.4 with S E = 0.6 , n = 40 , R 2 = 0.31 .

(a) State what β ^ 1 means, with units. (b) Test it against zero. (c) Does R 2 = 0.31 weaken the finding?

Check: (a) each additional kilogram of fertiliser is associated with 2.4 more units of yield on average; (b) t = 2.4 / 0.6 = 4.0 on 40 − 2 = 38 degrees of freedom, well beyond the critical value of about 2.02, so reject; (c) no. It says fertiliser tracks 31% of yield variation, and yield depends on soil, weather and much else. The estimate's precision is reported by its standard error, not by R 2 .

2: partly structured. An analyst compares two models on the same data: one with 3 predictors and R 2 = 0.62 , one with 12 and R 2 = 0.71 .

(a) Does the higher R 2 show the larger model is better? (b) What would you look at instead? (c) What if the 12 predictors were chosen after examining which improved R 2 ?

Check: (a) no, R 2 never falls when predictors are added, so a larger model wins mechanically; (b) whether each predictor belongs on subject-matter grounds, whether any is post-treatment, out-of-sample performance if prediction is the aim, and the standard errors of the coefficients of interest; (c) then the reported uncertainty is not trustworthy, since the specification is a function of the data it is being tested on.

3: unstructured. A city analyst reports: "Regressing neighbourhood crime rates on police patrol hours gives a coefficient of + 0.31 ( S E = 0.09 , R 2 = 0.44 ). More patrols are associated with more crime. Since the model explains 44% of the variance, we are confident in this relationship and recommend reducing patrols."

Assess the analysis and say what you would report.

Check: the coefficient is almost certainly reverse causation, patrols are allocated to high-crime neighbourhoods, so patrol hours are a consequence of crime as much as a potential cause. The positive sign is what that allocation produces, and reducing patrols on this basis could be actively harmful. R 2 = 0.44 is irrelevant to the problem: it reports in-sample tracking, and a model can track closely while reversing cause and effect. The standard error shows the association is precisely estimated, which is not the same as correctly interpreted. What to report: the association with an explicit statement that patrol allocation is endogenous; whether any variation in patrol hours arose independently of local crime, a staffing shock, a policy change, a boundary redrawing, which could support a credible design; and, absent that, no recommendation about patrol levels from this analysis at all.

What to carry forward

The estimator. OLS minimises ∑ i ( Y i − β 0 − β 1 X i ) 2 , giving β ^ 1 = ∑ i ( X i − X ¯ ) ( Y i − Y ¯ ) ∑ i ( X i − X ¯ ) 2 and, in matrix form, β ^ = ( X ⊤ X ) − 1 X ⊤ Y .

What a coefficient says. The average change in the outcome per unit change in that predictor, holding the model's other predictors fixed. An association, with units.

Coefficient inference. t = β ^ j / S E ( β ^ j ) on the residual degrees of freedom. The same template as every other test.

R 2 = 1 − S S E / S S T . In-sample explained variation. It never falls when predictors are added, so maximising it selects the largest model.

What R 2 does not establish. Correctness, causality, or unbiasedness. A low value is compatible with a credible, precisely estimated treatment effect.

Where causality comes from. The design or an identification argument. Never from the fit.

Variance assumptions. σ 2 ( X ⊤ X ) − 1 assumes homoskedastic uncorrelated errors; use robust or design-based standard errors when they fail.

Conditionality. Each coefficient depends on what else is in the model. Changing the predictor set changes what the others mean.

The recurring error. Reading a high R 2 as evidence the model is right.

Next step

Practice Linear Regression for Experimental Research

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Indicator Variables and Interactions

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.