Indicator Variables and Interactions

An indicator variable turns a category into a number a regression can use, and its coefficient is a difference in conditional means. An interaction lets a relationship differ by group, which changes what every other coefficient in the model means: once an interaction is present, the main effect is the effect at the reference value, not an overall effect.

Definition

An indicator (dummy) variable takes D i ∈ { 0 , 1 } to represent category membership. In Y i = β 0 + β 1 D i + ε i , the conditional means are E [ Y ∣ D = 0 ] = β 0 and E [ Y ∣ D = 1 ] = β 0 + β 1 , so β 1 = E [ Y ∣ D = 1 ] − E [ Y ∣ D = 0 ] . When D i = W i is treatment assignment in a two-arm experiment, OLS with an intercept gives β ^ 1 = Y ¯ 1 − Y ¯ 0 ; this algebraic equality does not by itself make the contrast causal. With an interaction, Y = β 0 + β 1 X + β 2 D + β 3 X D + ε , the fitted lines are E [ Y ∣ X , D = 0 ] = β 0 + β 1 X and E [ Y ∣ X , D = 1 ] = ( β 0 + β 2 ) + ( β 1 + β 3 ) X . So β 2 is the group difference at X = 0 and β 3 is the difference in slopes. Centring X at a meaningful value makes β 2 interpretable there.

Formal statement

Y = β 0 + β 1 D + ε with β 1 = E [ Y ∣ D = 1 ] − E [ Y ∣ D = 0 ] ; interacted, Y = β 0 + β 1 X + β 2 D + β 3 X D + ε , where β 2 is the group gap at X = 0 and β 3 the slope difference.

Assumptions and scope

  • With an interaction in the model, the main effect of D is the group difference at the reference value X = 0 , not an overall difference. Centring X moves that reference to a meaningful point.

  • An indicator coefficient is a difference in conditional means. It is causal only when the design or an identification argument makes it so; the algebraic equality with Y ¯ 1 − Y ¯ 0 does not supply that.

  • A categorical variable with k levels needs k − 1 indicators alongside an intercept. Including all k makes X ⊤ X singular and the coefficients individually undefined.

  • Every indicator coefficient is a comparison against the omitted reference category, so changing the reference changes each coefficient's meaning without changing the model's fit.

  • An interaction is symmetric: β 3 is equally the slope difference between groups and the change in the group gap per unit of X .

  • Testing an interaction for significance and dropping it when it fails to reject uses the data twice; whether the interaction belongs is a question about the research question, not only about a p-value.

  • A treatment-covariate interaction estimated in an experiment is a subgroup analysis, and its multiplicity must be accounted for if the subgroups were not specified in advance.

Worked material

Example

The same model, two reference points

A trial of a training programme fits, with X = years of experience entered uncentred:

Y = 42.0 + 1.8 X + 2.4 D + 0.9 X D .

The naive reading. "The programme adds 2.4 points."

What β 2 = 2.4 actually says. The group difference at X = 0 , among participants with zero years of experience. If the sample runs from 2 to 30 years, that is an extrapolation to a point no one occupies.

The difference at realistic values. The gap between the groups is 2.4 + 0.9 X :

  • at 5 years: 2.4 + 4.5 = 6.9 points
  • at 15 years: 2.4 + 13.5 = 15.9 points
  • at 25 years: 2.4 + 22.5 = 24.9 points

So the programme's benefit rises steeply with experience, and 2.4 understates it everywhere in the observed range.

What centring changes. Suppose mean experience is 12 years. Refitting with X c = X − 12 gives the same fit and a different β 2 : 2.4 + 0.9 ( 12 ) = 13.2 , the gap at average experience. The fitted lines, residuals and R 2 are identical, only the reference point moved.

β 2 was never "the effect". It was always "the effect at X = 0 ", and centring chooses where that is.

Non-example

Readings the parameterisation does not support

"With an interaction in the model, the main effect is the average effect." It is the effect at the other variable's reference value. Those coincide only if the reference happens to be the mean, which is what centring arranges deliberately.

Reading an uncentred β 2 when zero is outside the data. Age zero, income zero, baseline score zero, if no unit is near that value, the coefficient describes a point the data do not cover.

Including all k indicators for a k -level category alongside an intercept. The design matrix becomes singular and the coefficients are not individually identified.

Comparing indicator coefficients across models with different reference categories. Each is a comparison against whichever category was omitted; changing the reference changes every coefficient without changing the fit.

Dropping an interaction because its p-value exceeded 0.05, then reading the main effect as an overall effect. Testing and then conditioning on the test uses the data twice, and whether effects genuinely vary is a question about the research question, not only a threshold.

Treating a treatment-covariate interaction found after the fact as an established subgroup effect. Unless specified in advance, it carries the multiplicity of all the subgroups that could have been examined.

Contrast

Main effects with and without an interaction term

No interaction in the modelInteraction present
β 2 (group indicator)The group difference, everywhereThe group difference at X = 0
SlopesForced equal across groupsFree to differ, by β 3
Group gapConstant β 2 + β 3 X , varying with X
Effect of centring X Changes β 0 onlyChanges β 2 as well
A single headline numberDefensibleRequires choosing a value of X

Why the misreading persists. Software labels β 2 a "main effect", which in ordinary language suggests the principal or overall effect. In a model with an interaction it means neither.

The diagnostic question. Ask what happens to the interaction term when the other variable is zero. It vanishes, which is exactly why β 2 describes that point and no other.

Symmetry. β 3 is equally "the slope difference between groups" and "the change in the group gap per unit of X ". Both readings are correct and the second is usually more useful when D is a treatment.

What centring does and does not change. It relocates the reference point, changing β 0 and β 2 . It leaves the fitted values, residuals, R 2 , and β 3 untouched. The model is the same; the coefficients answer questions about a different point.

Common errors

Common misconception

In a model containing an interaction, the coefficient on a variable is still its overall effect, averaged across the values of the variable it interacts with.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.