Practice: Linear Regression for Experimental Research

Recognition · Interpretation

A randomized trial reports a treatment coefficient of 4.2 ( S E = 1.9 ) with R 2 = 0.012 . What does the low R 2 tell you?

2 hints available, least help first.

Hint 1: Retrieval cue

State what R 2 is computed from.

Hint 2: Concept cue

How many things besides a tutoring programme move a test score?

Direct application · Interpretation · Explanation

A regression of monthly revenue (£000s) on advertising spend (£000s) and a large-store indicator gives:

TermEstimateSE
Intercept12.43.1
Advertising1.850.42
Large store8.702.90

with n = 60 and R 2 = 0.71 . Stores chose their own advertising budgets.

State what the advertising coefficient means, test it, say whether the reported standard errors are the right ones, and say what the analysis does and does not support.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Say what is held fixed when you read the advertising coefficient.

Hint 2: Strategy cue

Ask who decided how much each store advertised, and why.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The coefficient. Among stores in the same size category, one spending £1,000 more per month is associated with £1,850 more monthly revenue on average. The comparison is conditional on store size, and the units follow those of the variables as entered.

The test. t = 1.85 / 0.42 ≈ 4.40 on 60 − 3 = 57 residual degrees of freedom, far beyond any conventional critical value. The association is precisely estimated.

What it does not support. A causal claim. Stores chose their own budgets, so a store expecting a strong month plausibly advertises more. The coefficient mixes any effect of advertising with whatever drove the spending decision, and least squares cannot separate them.

What R 2 = 0.71 does not add. It says the two predictors track 71% of revenue variation in these 60 observations. It is not evidence that the model is correct, that no confounder is missing, or that the coefficient is unbiased. A confounded model can track an outcome closely.

A further caution. Revenue variance plausibly grows with store size, so homoskedasticity is doubtful and robust standard errors are the safer choice.

What to report. The conditional association with robust standard errors, an explicit statement that spending was not assigned, and, if the causal question is the one that matters, a proposal to randomize advertising budgets across stores.

A complete answer does each of these:

  • states coefficient meaning
  • bounds r squared

Comparison · Evaluation

On the same data, model A has 3 predictors and R 2 = 0.62 ; model B has 12 and R 2 = 0.71 . What does the comparison establish?

2 hints available, least help first.

Hint 1: Retrieval cue

What happens to R 2 when you add a predictor that is pure noise?

Hint 2: Concept cue

If a statistic can only go one way, what does its direction tell you?

Recognition · Evaluation

An observational study regresses earnings on years of schooling with eight controls, reporting a coefficient of 2,400 per year. What would license reading this as the causal effect of schooling?

2 hints available, least help first.

Hint 1: Retrieval cue

Which of these is a claim about the world rather than about the fit?

Hint 2: Concept cue

Ask what would still be true if an unmeasured variable drove both schooling and earnings.

Comparison · Evaluation

A study measures 30 pupils in each of 40 classrooms, assigns a teaching method BY CLASSROOM, and regresses pupil test scores on the method indicator using ordinary standard errors. What is wrong?

2 hints available, least help first.

Hint 1: Retrieval cue

At what level did the treatment actually vary?

Hint 2: Concept cue

Count the units that could have received a different assignment.

Error diagnosis · Explanation · Evaluation

A city analyst reports:

Regressing neighbourhood crime rates on police patrol hours gives a coefficient of + 0.31 ( S E = 0.09 , R 2 = 0.44 ). More patrols are associated with more crime. Since the model explains 44% of the variance, we are confident in this relationship and recommend reducing patrols.

Identify the errors, say whether the reported standard error can be taken at face value, and say what should be reported.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask how police patrol hours get assigned to neighbourhoods.

Hint 2: Concept cue

If crime causes patrols, what sign would the coefficient take?

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The central error. The coefficient is almost certainly reverse causation. Patrols are allocated to neighbourhoods with high crime, so patrol hours are a consequence of crime at least as much as a potential cause of it. The positive sign is what that allocation process produces, and acting on it by reducing patrols could be actively harmful.

The R 2 error. An R 2 of 0.44 is offered as grounds for confidence. It reports in-sample tracking only. A model can track an outcome closely while reversing cause and effect, indeed that is exactly what happens here, because allocation makes patrols track crime well.

What the standard error shows. That the association is precisely estimated. Precision is not correctness of interpretation: a tightly estimated coefficient can be tightly estimating the wrong thing.

Whether the standard error can be trusted. Crime rates across neighbourhoods differ enormously in scale, so the error variance almost certainly grows with the neighbourhood's baseline level. Classical standard errors assume homoskedastic uncorrelated errors and would understate the uncertainty here; robust standard errors, clustered if neighbourhoods are observed repeatedly, are the appropriate choice.

What should be reported. The association, with an explicit statement that patrol allocation is determined by crime and therefore endogenous. Whether any variation in patrol hours arose independently of local crime, a staffing shock, a policy change, a boundary redrawing, a randomized pilot, since such variation could support a credible design. And absent that, no recommendation about patrol levels from this analysis at all.

Why no amount of fit would help. A specification with R 2 = 0.95 built the same way has the same problem. Only a change in where the variation comes from addresses it.

A complete answer does each of these:

  • locates causal warrant
  • flags variance assumptions

Transfer · Evaluation · Explanation

An analytics platform shows a "model quality" badge derived from R 2 . A growth team writes:

Our churn model scores 0.91 on model quality. It shows that customers contacted by support are 3x more likely to churn, so we are reducing proactive support outreach.

Assess the reasoning, say what you would want to know about how the uncertainty was computed, and say what you would report.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask what causes a customer to be contacted by support in the first place.

Hint 2: Strategy cue

Rename the badge back to what it is computed from, then reassess the claim.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The badge is R 2 renamed. Calling it "model quality" invites exactly the upgrade the statistic does not support. A score of 0.91 says the model tracks 91% of the variation in churn within this sample. It is not a measure of correctness, and it says nothing about whether any coefficient can be read causally.

The likely reverse causation. Customers contact support, or get contacted, when something is going wrong. A billing problem, a failed integration, declining usage. Those same troubles drive churn. So support contact is a marker of a customer already in difficulty, and the association runs from trouble to both contact and churn.

Why the high score makes it worse, not better. A model tracking churn well has found strong correlates of churn, and support contact is a strong correlate precisely because it flags struggling customers. The badge rewards the feature that makes the causal reading wrong.

The consequence of acting on it. Reducing outreach would remove help from the customers most at risk, and would plausibly increase churn, with the model then showing the association weakening, which could be misread as vindication.

What I would report. The association stated as such, with the reverse-causation mechanism named explicitly. The timing of contact relative to churn signals, since a contact following a usage drop is clearly a consequence. And, if the question is whether outreach helps, a randomized trial: withhold proactive outreach from a random subset of at-risk customers and compare. That is a design question, and no fit statistic on observational data substitutes for it.

How the uncertainty was computed. A churn model fitted over customers of very different sizes and tenures will have error variance that grows with account size, and customers within the same account or cohort are not independent. Classical standard errors assume neither problem exists. Before quoting any coefficient as precise, ask whether robust, and where the data are grouped, clustered, standard errors were used; a badge derived from R 2 says nothing about this.

A complete answer does each of these:

  • locates causal warrant
  • flags variance assumptions

Interpretation · Explanation

A regression of monthly household spending (£) on household income (£) and household size (people) reports:

TermEstimateSE
Intercept42085
Income0.310.04
Size14538

with R 2 = 0.46 .

State what the income coefficient and the R 2 each report, with units where they apply.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

A coefficient in a multiple regression holds the other predictors fixed.

Hint 2: Concept cue

What sample is R 2 computed on?

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The income coefficient. Households differing by £1 in monthly income differ by £0.31 in monthly spending ON AVERAGE, among households of the SAME SIZE. Three parts carry the meaning and each is load-bearing:

  • Units. Pounds of spending per pound of income, dimensionless here, but it must be stated, because the same number against income in thousands would mean something else entirely.
  • On average. It describes a mean difference across households, not what any particular household does.
  • Conditional on size. The coefficient compares households at the same household size. Without that clause it would be the unconditional association, which is a different quantity.

The R 2 . The model accounts for 46% of the variation in spending IN THIS SAMPLE. That is all it reports. It is not the probability the model is right, not a measure of how well the model would predict a new household, and not evidence about whether income causes spending.

What neither number reports. Nothing here establishes that raising a household's income by £1 would raise its spending by £0.31. That is a claim about an intervention, and it rests on the design, not on the table.

A complete answer does each of these:

  • states coefficient meaning
  • bounds r squared

Construction · Explanation

A retailer regresses weekly sales on whether a store ran a local advertising campaign, controlling for store size, region and month. The coefficient is + \pounds 4,100 per week.

(a) Write down what would have to be true for this to be the causal effect of advertising, and name one plausible violation.

(b) Say what analysis of the SAME data would be more defensible.

(c) The data hold 52 weeks for each of 300 stores. Say whether the ordinary standard errors describe this estimator, and what should replace them if not.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Who decided which stores ran a campaign, and on what basis?

Hint 2: Strategy cue

The data are repeated over weeks. What comparison does that make available?

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

What would have to be true. Conditional on store size, region and month, whether a store ran the campaign would have to be unrelated to everything else affecting its sales. Formally: within cells defined by those three controls, campaign assignment is as good as random with respect to potential sales. A plausible violation. Campaigns are not assigned at random. A regional manager chooses where to run them. If they target stores already trending upward, or stores in a catchment with a new housing development, then the campaign indicator is standing in for the growth that would have happened anyway. The coefficient then measures selection, and no amount of controlling for size and region repairs it, because the manager's reason is not in the data. Note the direction matters for the reading: targeting promising stores inflates the estimate, while targeting struggling stores would deflate it. "There is confounding" is a weaker statement than "the confounding runs this way." A more defensible analysis of the same data. Use the panel structure. Compare each store to ITSELF before and after its campaign, and difference that against stores that ran none in the same weeks. A difference-in-differences. It removes every fixed store characteristic, observed or not, and rests instead on a parallel-trends assumption, which is at least partly checkable: plot the pre-campaign trajectories of both groups and look for divergence before the campaign began. (c) The standard errors. No. There are 15,600 rows but only 300 stores, and sales within a store are correlated week to week. A store that sells well in one week sells well in the next, for reasons that have nothing to do with the campaign. Ordinary standard errors treat every row as an independent observation and will be far too small, so an effect can appear decisively estimated on the strength of repeated measurements of the same few stores. Cluster the standard errors at the STORE level, which is the level at which the campaign varied and at which the correlation lives. With 300 clusters the asymptotics are comfortable; with a dozen they would not be, and a wild bootstrap would be the alternative. The point. Neither improvement came from a better fit or more controls. One came from a design making a weaker, checkable assumption; the other from counting how many things could actually have come out differently.

A complete answer does each of these:

  • locates causal warrant
  • flags variance assumptions

Evaluation · Comparison

A study assigns a teaching method by classroom across 40 classrooms of 30 pupils each, then regresses individual pupil test scores on the method indicator, reporting a coefficient of 4.1 points with ordinary standard errors and R 2 = 0.38 .

Select every statement that is correct.

Select every option that applies

Every option that applies, and only those. The set is checked as a whole.

Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.