Module 3 of 3 · Lesson 2 of 3

Indicator Variables and Interactions

Indicator coding and interactions, and how the reference category determines a main effect.

What you will be able to do

Given a model containing indicators and interactions, the learner can state what each coefficient means, identify the reference point it depends on, and recover a quantity of interest from the parameterisation.

Orientation

A coefficient on an interaction term is not "the effect of the interaction". What each coefficient means depends on what the others are holding fixed, and the main effect is the one people misread.

The arithmetic is elementary. The difficulty is that a coefficient's meaning depends on how the model was parameterised, so the same number answers different questions depending on what else is in the equation.

Intuition

What the group coefficient reports before and after centring

The canonical intuition establishes that an indicator coefficient is a difference between conditional means, and that adding an interaction changes what that coefficient refers to. This block follows one fitted model through that change.

Fit Y = β 0 + β 1 X + β 2 D + β 3 X D , where X is age in years and D marks group membership, and suppose the fit gives β 2 = − 1.8 with β 3 = 0.05 .

Without the interaction term, β 2 would be the group difference at every age, because the model permits only a vertical shift.

With the interaction term present, the group difference is β 2 + β 3 X , which depends on X . At X = 0 it is − 1.8 ; at X = 40 it is − 1.8 + 0.05 ( 40 ) = 0.2 ; at X = 60 it is 1.2 . The reported β 2 describes the difference at age zero, and the groups differ in the opposite direction over most of the observed range.

After centring X at its sample mean of 42 , the same fit returns β 2 = − 1.8 + 0.05 ( 42 ) = 0.3 , the difference at the average age. The fitted values, the residuals and β 3 are unchanged; only the quantity β 2 names has changed.

Centring therefore does not improve the model. It moves the reference point of the group coefficient to a value within the data, so that the coefficient reports a comparison someone asked for.

Definition

Indicators, interactions, and what each coefficient means

The canonical statement gives the two regressions and what each coefficient equals. Three readings of it decide whether the numbers are reported correctly, and none is stated above.

The algebraic identity is not the warrant. That OLS with an intercept returns β ^ 1 = Y ¯ 1 − Y ¯ 0 is true of any data whatever produced them: it is a fact about least squares, not about the world. What licenses reading it as an effect is the assignment mechanism. A regression run on observational data returns the same identity and supports nothing causal.

An interaction moves what the main effect means. In Y = β 0 + β 1 X + β 2 D + β 3 X D + ε , β 2 is the group difference at X = 0 and nowhere else. If X is untransformed age, β 2 is the difference at birth, which is usually outside the data and always beside the point. Centring X at a value the study cares about is what makes β 2 answer a question somebody asked.

The reference category is a choice that shows up in every coefficient. With k categories and k − 1 indicators, every reported difference is against the omitted one. Change which category is omitted and all the coefficients change, though the fitted values do not. A table of indicator coefficients is uninterpretable without saying which category is missing from it.

Example

The same model, two reference points

A trial of a training programme fits, with X = years of experience entered uncentred:

Y = 42.0 + 1.8 X + 2.4 D + 0.9 X D .

The naive reading. "The programme adds 2.4 points."

What β 2 = 2.4 actually says. The group difference at X = 0 , among participants with zero years of experience. If the sample runs from 2 to 30 years, that is an extrapolation to a point no one occupies.

The difference at realistic values. The gap between the groups is 2.4 + 0.9 X :

  • at 5 years: 2.4 + 4.5 = 6.9 points
  • at 15 years: 2.4 + 13.5 = 15.9 points
  • at 25 years: 2.4 + 22.5 = 24.9 points

So the programme's benefit rises steeply with experience, and 2.4 understates it everywhere in the observed range.

What centring changes. Suppose mean experience is 12 years. Refitting with X c = X − 12 gives the same fit and a different β 2 : 2.4 + 0.9 ( 12 ) = 13.2 , the gap at average experience. The fitted lines, residuals and R 2 are identical, only the reference point moved.

β 2 was never "the effect". It was always "the effect at X = 0 ", and centring chooses where that is.

Worked example

An effect that is not what it looks like

Problem. A study of a tutoring programme reports:

score ^ = 55.0 + 0.42 baseline + 3.10 T − 0.05 ( baseline × T )

where baseline runs from 20 to 90 with mean 60, and T is the tutoring indicator. The authors write: "Tutoring raises scores by 3.1 points."

Goal. Say what the coefficients report and give the effect on a defensible scale.

Relevant principle. With an interaction present, the coefficient on T is the treatment effect at baseline = 0 , not an overall effect.

Step 1: locate the reference point. β T = 3.10 is the effect where baseline = 0 .

Reason: setting baseline to 0 zeroes the interaction term, leaving T 's coefficient alone.

Step 2: judge whether that point is meaningful. Baseline scores run from 20 to 90. Zero is outside the data entirely, so 3.10 describes a student who does not exist in this sample.

Reason: a fitted surface is estimated where the data are; reading it elsewhere extrapolates.

Step 3: write the effect as a function.

effect at baseline  b = 3.10 − 0.05 b .

Step 4: evaluate across the observed range.

  • at b = 20 : 3.10 − 1.00 = 2.10 points
  • at b = 60 (the mean): 3.10 − 3.00 = 0.10 points
  • at b = 90 : 3.10 − 4.50 = − 1.40 points

Reason: the negative interaction means the benefit shrinks as baseline rises, reversing sign in the upper range.

Step 5: restate the finding. The programme helps weaker students modestly, does essentially nothing at the average, and is associated with slightly worse outcomes for the strongest. "Raises scores by 3.1 points" is true at no observed value of baseline.

Result. The effect is 3.10 − 0.05 b , ranging from + 2.10 to − 1.40 across the sample.

Check. Centring baseline at its mean would give a T coefficient of 3.10 − 0.05 ( 60 ) = 0.10 . The effect at average baseline. Same model, same fit, and a headline number that does not mislead.

Interpretation. Report the effect at substantively chosen values of baseline, or centre and report the average-case effect alongside the interaction. Also note that this is a subgroup pattern: unless the variation by baseline was specified in advance, it is a hypothesis for the next study rather than an established finding.

Non-example

Readings the parameterisation does not support

"With an interaction in the model, the main effect is the average effect." It is the effect at the other variable's reference value. Those coincide only if the reference happens to be the mean, which is what centring arranges deliberately.

Reading an uncentred β 2 when zero is outside the data. Age zero, income zero, baseline score zero, if no unit is near that value, the coefficient describes a point the data do not cover.

Including all k indicators for a k -level category alongside an intercept. The design matrix becomes singular and the coefficients are not individually identified.

Comparing indicator coefficients across models with different reference categories. Each is a comparison against whichever category was omitted; changing the reference changes every coefficient without changing the fit.

Dropping an interaction because its p-value exceeded 0.05, then reading the main effect as an overall effect. Testing and then conditioning on the test uses the data twice, and whether effects genuinely vary is a question about the research question, not only a threshold.

Treating a treatment-covariate interaction found after the fact as an established subgroup effect. Unless specified in advance, it carries the multiplicity of all the subgroups that could have been examined.

Contrast

Main effects with and without an interaction term

No interaction in the modelInteraction present
β 2 (group indicator)The group difference, everywhereThe group difference at X = 0
SlopesForced equal across groupsFree to differ, by β 3
Group gapConstant β 2 + β 3 X , varying with X
Effect of centring X Changes β 0 onlyChanges β 2 as well
A single headline numberDefensibleRequires choosing a value of X

Why the misreading persists. Software labels β 2 a "main effect", which in ordinary language suggests the principal or overall effect. In a model with an interaction it means neither.

The diagnostic question. Ask what happens to the interaction term when the other variable is zero. It vanishes, which is exactly why β 2 describes that point and no other.

Symmetry. β 3 is equally "the slope difference between groups" and "the change in the group gap per unit of X ". Both readings are correct and the second is usually more useful when D is a treatment.

What centring does and does not change. It relocates the reference point, changing β 0 and β 2 . It leaves the fitted values, residuals, R 2 , and β 3 untouched. The model is the same; the coefficients answer questions about a different point.

Exercise

1: fully structured. A model gives Y = 20 + 3 X + 5 D + 2 X D , with X uncentred and ranging from 1 to 10.

(a) What is the fitted line for D = 0 ? For D = 1 ? (b) What does β 2 = 5 report? (c) What is the group gap at X = 4 ?

Check: (a) D = 0 : Y = 20 + 3 X ; D = 1 : Y = 25 + 5 X ; (b) the group difference at X = 0 , which lies outside the observed range of 1 to 10; (c) 5 + 2 ( 4 ) = 13 .

2: partly structured. An analyst centres a covariate at its mean and refits, reporting that the treatment coefficient changed from 1.2 to 4.8.

(a) Did the model change? (b) What do the two numbers mean? (c) Which should be reported?

Check: (a) no, fitted values, residuals and R 2 are identical, and so is the interaction coefficient; only the reference point moved; (b) 1.2 is the treatment effect at the covariate's original zero, 4.8 is the effect at its mean; (c) 4.8 for a single headline figure, since the average case is a point the data actually cover, better still, report the effect across a few substantively chosen covariate values.

3: unstructured. A clinical report states: "In our model including a treatment-by-age interaction, the treatment coefficient was 0.4 ( p = 0.61 ) and the interaction was 0.08 ( p = 0.03 ). We conclude the treatment has no overall effect but works better in older patients."

Assess both conclusions. Patients range from 45 to 80 years, and age was entered uncentred.

Check: the first conclusion is unsupported, 0.4 is the treatment effect at age zero, far outside the 45-to-80 range, so its non-significance says nothing about the effect at any age anyone in the study had. The effect at age a is 0.4 + 0.08 a , giving 0.4 + 3.6 = 4.0 at age 45 and 0.4 + 6.4 = 6.8 at age 80, substantial throughout the observed range, and the opposite of "no overall effect". The second conclusion is better supported, since the interaction is a statement about how the effect varies and does not depend on the reference point, though it should be checked whether the age interaction was pre-specified; if it emerged from exploration, it carries the multiplicity of all the subgroup analyses that might have been run. What to report: the effect at several ages with intervals, or centre age at its mean and report that coefficient as the headline.

What to carry forward

Indicator coefficient. β 1 = E [ Y ∣ D = 1 ] − E [ Y ∣ D = 0 ] . A difference in conditional means against the omitted reference category.

In a two-arm experiment. OLS on an intercept and a treatment indicator returns Y ¯ 1 − Y ¯ 0 exactly. The algebra does not supply the causal reading; the design does.

Interaction model. Y = β 0 + β 1 X + β 2 D + β 3 X D + ε , giving lines β 0 + β 1 X and ( β 0 + β 2 ) + ( β 1 + β 3 ) X .

The key reading. With an interaction present, β 2 is the group difference at X = 0 , not an overall difference. The gap is β 2 + β 3 X .

β 3 . The difference in slopes, equivalently the change in the group gap per unit of X .

Centring. Subtracting a meaningful value from X moves the reference point, changing β 0 and β 2 while leaving the fit and β 3 untouched.

k levels need k − 1 indicators. Including all of them alongside an intercept makes the coefficients undefined.

Subgroup caution. A treatment-covariate interaction not specified in advance carries the multiplicity of every subgroup that could have been examined.

The recurring error. Reading a main effect as an overall effect when an interaction is in the model.

Next step

Practice Indicator Variables and Interactions

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Binary Outcome Models for Experimental Research

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.