Module 3 of 3 · Lesson 1 of 3
Linear Regression for Experimental Research
Regression coefficients and standard errors, and the causal and predictive claims they do not support.
What you will be able to do
Given regression output, the learner can state what a coefficient means as an average change per unit of the predictor, conditional on the others in the model and with its units, and treat
What you will be able to do
Given regression output and a described study, the learner can identify the design or identification argument, not the fit, as the source of any causal reading, and say when the classical standard errors are inappropriate and what should replace them.
Orientation
Regression output reports several different things, and the temptation is to read all of them as evidence the model is right.
Fitting is mechanical and has a closed form. The difficulty is refusing two upgrades the output invites: from association to cause, and from explained variation to a well-specified model.
Intuition
Reading a coefficient when another predictor is added
The canonical intuition states that a slope is an association holding the model's other predictors fixed, and that
A regression of monthly spending on advertising alone returns a slope of
Neither number is wrong, and they answer different questions. The first is the average difference in spending between months differing by one pound of advertising, with nothing held fixed. The second is that difference among months with equal footfall.
Which one is wanted depends on the question. If advertising works partly by drawing people into the store, then footfall is on the causal path, and conditioning on it removes part of the effect being estimated. If footfall is instead driven by seasonal trade that also affects advertising budgets, holding it fixed removes a confounding path and the second number is closer to what was wanted.
The regression cannot distinguish these two cases. Both produce the same drop from
Definition
Least squares in matrix form, and the variance assumptions
The canonical definition gives the estimator, the residual variance, the
The divisor in
What the
What
Where the variance formula comes from.
When
Example
The same effect at two very different R-squareds
Two regressions from the same randomized trial of a tutoring programme,
Regression 1: outcome on treatment alone.
Regression 2: outcome on treatment, baseline score, and school.
What changed and what did not. The estimated effect barely moved, which randomization leads you to expect. The standard error more than halved, because baseline score explains much of the outcome variation that was previously in the residual.
What the
Reading Regression 1 on its own. An
Worked example
Reading a regression table
Problem. A regression of monthly revenue (in thousands) on advertising spend (in thousands) and a store-size indicator gives:
| Term | Estimate | SE |
|---|---|---|
| Intercept | 12.4 | 3.1 |
| Advertising | 1.85 | 0.42 |
| Large store | 8.70 | 2.90 |
with
Goal. State what each number supports, and correct the conclusion.
Relevant principle. A coefficient is a conditional association; the causal reading comes from the design;
Step 1: read the advertising coefficient. For two stores of the same size category, one spending £1,000 more per month is associated with £1,850 more revenue on average.
Reason: the coefficient is conditional on the other predictors, so the comparison is between stores alike on store size, and the units are those of the variables as entered.
Step 2: test it.
Reason: the same template as any other test, applied to
Step 3: reject the causal claim. Advertising spend was chosen by the stores, not assigned. Stores expecting a strong month plausibly advertise more, so the coefficient mixes advertising's effect with whatever drove the spending decision.
Reason: nothing about least squares distinguishes a cause from a correlate; that distinction is a property of how the data arose.
Step 4: reject the
Reason:
Step 5: check the standard errors. Revenue variance plausibly grows with store size, so homoskedasticity is doubtful, and robust standard errors are the safer choice here.
Result. A well-estimated conditional association of £1,850 per £1,000, with no causal warrant and an
Check. Would a higher
Interpretation. Report the association with robust standard errors, state plainly that spending was not assigned, and propose a randomized budget trial if the causal question is the one that matters.
Non-example
Claims a fitted regression does not support
"
"
Selecting the specification with the highest
Reading a coefficient without naming what is held fixed. "The effect of advertising is 1.85" omits that the comparison is between stores alike on the other predictors. Change the predictor set and the number changes meaning.
Using classical standard errors when the variance is not constant.
Predicting outside the observed range of the predictors. The fitted line is estimated where the data are. Extending it asserts a relationship over a region the data do not cover.
Contrast
Explained variation against causal warrant
| What | What a causal reading requires | |
|---|---|---|
| Question answered | How much outcome variation does the model track here? | What would happen under an intervention? |
| Computed from | Fitted versus observed values | Nothing in the output — it comes from the design |
| Improved by | Adding predictors, always | Randomization, or an identification argument |
| High value implies | Good in-sample tracking | Nothing |
| Low value implies | Much variation unexplained | Nothing |
Why the two get conflated. They appear in the same output, and "explains 71% of the variance" sounds like a claim about explanation in the ordinary sense. It is not; it is a claim about squared distances.
The clean separation. Ask where the comparison came from. If treatment was randomized, the coefficient is causal at any
A useful case to hold onto. A well-run trial with
Where
Exercise
1: fully structured. A regression of yield on fertiliser (kg) gives
(a) State what
Check: (a) each additional kilogram of fertiliser is associated with 2.4 more units of yield on average; (b)
2: partly structured. An analyst compares two models on the same data: one with 3 predictors and
(a) Does the higher
Check: (a) no,
3: unstructured. A city analyst reports: "Regressing neighbourhood crime rates on police patrol hours gives a coefficient of
Assess the analysis and say what you would report.
Check: the coefficient is almost certainly reverse causation, patrols are allocated to high-crime neighbourhoods, so patrol hours are a consequence of crime as much as a potential cause. The positive sign is what that allocation produces, and reducing patrols on this basis could be actively harmful.
What to carry forward
The estimator. OLS minimises
What a coefficient says. The average change in the outcome per unit change in that predictor, holding the model's other predictors fixed. An association, with units.
Coefficient inference.
What
Where causality comes from. The design or an identification argument. Never from the fit.
Variance assumptions.
Conditionality. Each coefficient depends on what else is in the model. Changing the predictor set changes what the others mean.
The recurring error. Reading a high