Practice: Indicator Variables and Interactions

Recognition · Interpretation

In Y i = β 0 + β 1 D i + ε i with D i ∈ { 0 , 1 } and no other predictors, what is β 1 ?

2 hints available, least help first.

Hint 1: Retrieval cue

Substitute D = 0 and then D = 1 into the equation.

Hint 2: Concept cue

Two values of D means two fitted values. What separates them?

Direct application · Interpretation · Explanation

A trial fits Y = 42.0 + 1.8 X + 2.4 D + 0.9 X D , where X is years of experience entered uncentred and ranging from 2 to 30, and D is the programme indicator.

State what β 2 = 2.4 reports, give the group gap at 5, 15 and 25 years, and say what centring X at its mean of 12 would change.

Finally: the trial's authors say the regression shows the programme caused the gain. Say what in the setup licenses a causal reading, and what the regression contributes.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Write the fitted line for each group and subtract.

Hint 2: Strategy cue

Ask what happens to the interaction term when X = 0 .

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

What β 2 = 2.4 reports. The group difference at X = 0 , among participants with zero years of experience. The sample runs from 2 to 30 years, so that is an extrapolation to a point no one in the study occupies.

The gap as a function of X . Setting D = 1 and subtracting the D = 0 line gives 2.4 + 0.9 X .

  • At 5 years: 2.4 + 4.5 = 6.9 points
  • At 15 years: 2.4 + 13.5 = 15.9 points
  • At 25 years: 2.4 + 22.5 = 24.9 points

The programme's benefit rises steeply with experience, and 2.4 understates it everywhere in the observed range.

What centring changes. Refitting with X c = X − 12 gives β 2 = 2.4 + 0.9 ( 12 ) = 13.2 , the gap at average experience. The fitted values, residuals, R 2 and β 3 are all unchanged, only the reference point moved, and with it the question β 2 answers.

The general lesson. β 2 was never the overall effect. It was always the effect at X = 0 , and centring chooses where that zero sits.

Algebra against warrant. With an intercept and a treatment indicator, least squares returns the difference in conditional means exactly, that is an identity, true of any data whatever produced them. What licenses reading it as an effect is that this is a trial: assignment was randomized, so the two arms are comparable in expectation. The regression contributes the arithmetic and, through the interaction, the way the gap varies with experience; it supplies none of the causal warrant.

A complete answer does each of these:

  • reads indicator as mean difference
  • conditions main effect on reference
  • recovers group quantities
  • separates algebra from causality

Comparison · Evaluation

An analyst centres a covariate at its mean and refits a model containing an interaction. Which set of quantities changes?

2 hints available, least help first.

Hint 1: Retrieval cue

Ask which coefficients are defined at the point where the covariate is zero.

Hint 2: Concept cue

If the fitted surface is identical, which numbers can possibly have changed?

Error diagnosis · Explanation · Evaluation

A clinical report states:

In our model including a treatment-by-age interaction, the treatment coefficient was 0.4 ( p = 0.61 ) and the interaction was 0.08 ( p = 0.03 ). We conclude the treatment has no overall effect but works better in older patients.

Patients range from 45 to 80 years and age was entered uncentred. Assess both conclusions.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask at what value of age the reported treatment coefficient applies.

Hint 2: Concept cue

Is that value inside the range of patients studied?

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The first conclusion is unsupported. The coefficient 0.4 is the treatment effect at age zero, which is 45 years below the youngest patient in the study. Its non-significance says nothing about the effect at any age anyone actually had, and reporting it as 'no overall effect' is the error.

The effect across the observed range. The treatment effect at age a is 0.4 + 0.08 a :

  • at 45: 0.4 + 3.6 = 4.0
  • at 65: 0.4 + 5.2 = 5.6
  • at 80: 0.4 + 6.4 = 6.8

Substantial throughout, and the opposite of what was claimed.

The second conclusion is better supported. The interaction describes how the effect varies with age and does not depend on the reference point, so it is not vulnerable to the same error.

A caveat on it. Whether the age interaction was specified in advance matters. If it emerged from exploring subgroups, it carries the multiplicity of every subgroup analysis that might have been run, and should be treated as a hypothesis for the next study.

What should be reported. The treatment effect at several substantively chosen ages with intervals, or the model refitted with age centred at its mean so the headline coefficient describes the average patient.

One thing the report gets right for the wrong reason. Fitting the model does return the group difference at the reference value exactly, but that is an algebraic property of least squares with an intercept and an indicator. It holds whether or not anyone was assigned to anything, so it cannot be what makes the coefficient an effect. The design has to do that, and a regression run on observational data returns the same identity with no warrant behind it.

A complete answer does each of these:

  • reads indicator as mean difference
  • conditions main effect on reference
  • recovers group quantities
  • separates algebra from causality

Transfer · Evaluation · Explanation

A pricing team fits demand on price, an enterprise-segment indicator, and their interaction. Price ranges from £40 to £250. They report:

Enterprise customers buy 180 more units than small-business customers (segment coefficient = 180, p < 0.001), and the price sensitivity differs between segments (interaction = −1.4). We will therefore prioritise enterprise accounts.

Assess the reasoning and say what you would report.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask at what price the reported segment coefficient applies.

Hint 2: Strategy cue

Write the segment gap as a function of price and evaluate it across the observed range.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The segment coefficient is misread. With the interaction in the model, 180 is the segment difference at a price of zero, outside the £40 to £250 range and not a price anyone is charged. It is not the overall segment gap.

The gap across the actual price range. The segment difference is 180 − 1.4 × price :

  • at £40: 180 − 56 = 124 units
  • at £145 (mid-range): 180 − 203 = − 23 units
  • at £250: 180 − 350 = − 170 units

So enterprise customers buy more only at low prices; above roughly £129 the sign reverses and small-business customers buy more. The headline of 180 describes none of this.

The consequence for the decision. Prioritising enterprise accounts on the strength of a 180-unit advantage is unsound, because that advantage does not exist at most prices in the range. The right question is what price points the business actually operates at, and the segment comparison should be evaluated there.

What I would report. The segment gap at the prices actually charged, with intervals, or the model refitted with price centred at a representative value so the segment coefficient describes a real operating point. The interaction itself is the substantive finding: price sensitivity differs between segments, and that is what should drive segment-specific pricing rather than a blanket prioritisation.

One more check. Whether prices were set by the business in response to expected demand, in which case the coefficients are conditional associations rather than a demand curve, and no pricing decision should be read off them directly.

Why the algebra is not the argument. The enterprise coefficient is exactly the difference in conditional means between segments at the reference price, by construction. Customers were not assigned to segments; they arrived in them. So the number is a description of who buys what, and reading it as what would happen if a customer were moved between segments asks the regression for something only a design could supply.

A complete answer does each of these:

  • reads indicator as mean difference
  • conditions main effect on reference
  • recovers group quantities
  • separates algebra from causality
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.