Module 1 of 1 · Lesson 1 of 1

When a Coefficient Is Not an Effect

Why a high R 2 is compatible with a coefficient that is not the effect, and what would identify one.

What you will be able to do

The learner can identify the sources of endogeneity that make a least-squares coefficient differ from the causal effect, compute the direction and size of omitted-variable bias from the bias formula, explain why goodness of fit carries no information about that bias, and state what an instrumental variable would have to satisfy to repair it.

Orientation

Association against causal effect in a fitted coefficient

A fitted regression reports how the outcome differs across units that differ in the regressor. The question usually being asked is different: what would happen to the outcome if the regressor were changed. The two coincide only when nothing else systematically differs alongside the regressor, and in observational data that condition is an assumption rather than a finding.

This unit is about how far apart the two can be, and what the gap is made of.

The gap has a formula for one of its sources. Leaving out a determinant of the outcome that travels with the regressor shifts the coefficient by an amount that can be computed exactly: how much the omitted variable matters, times how much it moves with the included one.

None of the model's own diagnostics can see it. On the worked data later in this unit, a regression omitting a confounder reaches R 2 = 0.992876 , an excellent fit by any conventional reading, while its slope is wrong by 119.3 % . Every residual plot, every standard error, every information criterion is computed from the equation as specified, and an omitted variable is by definition not in it.

The repair is not more controls. Adding a control removes that control's contribution and nothing else. What remains unobserved keeps contributing, and no improvement in fit indicates whether anything does.

The unit ends with instrumental variables, which look like a technical remedy and are mostly an argument: one of the two conditions an instrument must satisfy can be checked in the data, and the one that matters cannot.

Definition

Exogeneity, the bias formula, and the instrument conditions

The canonical statements above give exogeneity, the bias formula and the two instrument conditions. What follows is the status of each claim, since that is what decides how it can be defended.

Exogeneity is an assumption about unobservables. E [ u ∣ x ] = 0 concerns the relationship between the regressor and everything affecting the outcome that is not in the model. Nothing in the data bears on it directly, because u is not observed; what can be done is to name the candidate violations and argue about each.

The bias formula is an algebraic identity, not a modelling assumption.

β ^ short = β ^ long + β 2 δ

holds exactly between two least-squares fits on the same sample, whatever generated the data. It is therefore always available as a decomposition of why a coefficient moved when a control was added. What it does not say is that the long regression is unbiased. It relates two fits to each other, and both may omit something.

ClaimStatusHow it is settled
β ^ short − β ^ long = β 2 δ identityalgebra, exact
E [ u ∣ x ] = 0 assumptionargument about the setting
instrument relevancetestablefirst-stage regression
exclusion restrictionassumptionargument about the setting

The three mechanisms differ in what they do to the coefficient. An omitted determinant moves it by β 2 δ , in either direction depending on the two signs. Simultaneity moves it in a direction determined by the structure of the system. Classical measurement error in the regressor attenuates it toward zero, which means the estimated effect is too small in magnitude rather than too large. A direction worth knowing, since it makes a null result harder to interpret than a positive one.

Relevance and exclusion are asymmetric in kind, not merely in difficulty. Relevance is a statement about two observed variables and has a sample analogue. Exclusion says the instrument has no path to the outcome except through the regressor, which involves the unobserved disturbance and so has no sample analogue at all. Reporting a first-stage F statistic is evidence about the first and silence about the second.

A failed exclusion restriction is not detectable from the output. An instrument correlated with the omitted determinant produces an estimate that is wrong in a way no diagnostic flags, and a strong first stage does not protect against it. The worked example exhibits exactly this, with a first-stage R 2 of 0.986014 and an estimate no closer to the structural value than the biased least-squares one.

Intuition

The source of omitted-variable bias

Least squares finds the line that best tracks the data it was given. If a determinant of the outcome is missing from the equation and travels with the included regressor, the fitting procedure has no way to distinguish the two influences, so it attributes both to the regressor it can see.

That attribution has a size, and the size is a product of two things. How much the omitted variable matters for the outcome, β 2 , and how much it moves with the included regressor, δ . Their product is the whole of the bias:

β ^ short − β ^ long = β 2 δ .

Either factor being zero leaves the coefficient untouched. An omitted variable that matters enormously for y but is unrelated to x does no damage, which is why the question to ask about a missing variable is not "is it important" but "does it travel with the regressor".

The fit statistics are computed inside the equation. R 2 compares the fitted values against the outcome using the regressors that are present. Residual plots show the residuals of the model as specified. Standard errors describe the sampling variability of the coefficients in that specification. Every one of them takes the equation as given, and an omitted variable is exactly what is not in the equation.

So there is no tension between an excellent fit and a badly wrong coefficient. On the worked data the short regression achieves R 2 = 0.992876 while its slope is 119.3 % too large. Both are true at once, because they answer different questions: how well does this equation track the data, and does this coefficient measure the effect.

Adding controls removes one term at a time. Including z subtracts β 2 δ from the bias and nothing else. Anything still unobserved keeps contributing its own product, and no improvement in fit reveals whether it does. This is why a coefficient that shifts substantially when a control is added is informative in a specific way: it shows controls matter in this setting, which is a reason to suspect the ones that remain unavailable.

An instrument works by changing where the variation comes from. Rather than trying to name and measure every confounder, it isolates a slice of the regressor's movement whose origin is known and argued to be unconnected to the outcome except through the regressor. The strength of that slice is measurable. Whether its origin really is unconnected is not, and the worked example shows an instrument with a first-stage R 2 of 0.986014 producing an estimate of 2.451064 against a structural value of 1 , because it carried a relationship to the omitted determinant that no output would reveal.

Example

Five regressions and what each coefficient identifies

Earnings on years of schooling. The classic case. Whatever leads someone to stay in education, family resources, prior attainment, expectations, also bears on earnings, so the omitted determinants travel with the regressor. The coefficient is upward biased if those determinants raise both, and the sign follows from that reasoning rather than from the data.

Hospital admission on health outcome. People admitted to hospital are sicker than those who are not, so a regression of mortality on admission finds a positive coefficient. Reading it as the effect of being admitted inverts the causation: the regressor responds to the same underlying condition driving the outcome. Here the bias is not subtle and the direction is obvious, which is what makes it a useful case for seeing the mechanism.

Quantity sold on price. Price and quantity are determined together by supply and demand, so neither is exogenous to the other. A regression traces out neither curve but a mixture whose composition depends on which side of the market moved more. This is simultaneity, and no set of controls addresses it. The problem is the structure of the system rather than a missing variable.

Consumption on reported income. If income is measured with error, recall error in a survey, say, the coefficient is attenuated toward zero under classical assumptions. The estimate understates the relationship, which makes a small or insignificant result ambiguous between a weak effect and a well-measured one.

Crop yield on rainfall, in an agricultural trial with assigned irrigation. Where the regressor was assigned by the experimenter, exogeneity holds by design and the coefficient does estimate the effect. This is the case that shows the others are about how the data arose rather than about regression as a technique.

---

The first two have omitted determinants, one obvious and one less so. The third has no missing variable at all and is still not identified. The fourth is biased toward zero rather than away from it. The fifth is fine, and the only thing distinguishing it is that somebody controlled the assignment. In every case the diagnosis came from the setting rather than from the output.

Procedure

Working out what a coefficient identifies

Before fitting anything, state the question. Write the quantity of interest as a sentence about an intervention: what would happen to y if x were changed, holding what fixed. A coefficient can then be compared against that sentence rather than being reported as whatever it is.

To assess exogeneity.

  1. List the determinants of y that subject-matter knowledge suggests, whether or not they are in the data.
  2. For each, ask whether it plausibly moves with x . A determinant unrelated to the regressor causes no bias however important it is for the outcome.
  3. Ask whether x could respond to y or to the same disturbance. The simultaneity case.
  4. Ask how x was measured. Classical error in the regressor attenuates the coefficient toward zero.
  5. Record the answers. This list is the identification argument, and it is what a reader needs in order to judge the estimate.

To quantify what one omitted variable does, when it is available.

  1. Fit the short regression, y on x .
  2. Fit the long regression, y on x and z , recording β ^ 2 .
  3. Regress z on x to obtain δ .
  4. Confirm β ^ short − β ^ long = β ^ 2 δ . Disagreement means an arithmetic error, since the relation is an identity.
  5. Sign it before computing it where possible: the direction follows from whether z raises or lowers y , and whether z rises or falls with x .

To use this as evidence about what remains. A coefficient that moves substantially when one control is added shows that controls matter here, which is a reason to expect that unobserved ones do too. A coefficient that barely moves is weak evidence in the other direction and is not a demonstration of exogeneity.

To assess a proposed instrument.

  1. Relevance. Regress x on the instrument and report the first-stage statistic. A weak first stage produces an estimator biased toward least squares with standard errors that understate the uncertainty.
  2. Exclusion. Write the argument: where does the instrument's variation come from, and why can it not reach y except through x ? This is prose, not a statistic, and it is the part a reader should weigh most heavily.
  3. Look for the clear-cut failures. Is the instrument correlated with any determinant already on the list from the exogeneity assessment? If so, the restriction fails and the strength of the first stage is irrelevant.
  4. State the population. The estimate identifies an effect for units whose x responds to the instrument, which may not be the population the question is about.

Checks. Confirm that adding a control changes the coefficient by exactly the amount the bias formula predicts, which validates the arithmetic. And before reporting any estimate as an effect, write down what would have to be true for that reading to hold; if the list cannot be written, the estimate is a description of an association and should be reported as one.

Worked example

An excellent fit with a coefficient wrong by 119%

Twelve observations. The outcome is generated as y = 2 + 1.0 x + 3.0 z + e , so the structural coefficient on x is 1.0 , and z is correlated with x .

x 123456789101112
z 223344556677

Step 1: the short regression, omitting z .

β ^ short = S x y S x x = 1401 572 = 2.449301 .

Step 2: the long regression, including z .

β ^ long = 67 60 = 1.116667 , β ^ 2 = 5717 2100 = 2.722381 .

Step 3: the bias formula. Regressing z on x gives

δ = S x z S x x = 70 143 = 0.489510 ,

so the predicted bias is

β ^ 2 δ = 2.722381 × 0.489510 = 1.332634 .

And the observed difference:

β ^ short − β ^ long = 2.449301 − 1.116667 = 1.332634 .

Exactly equal. Computed in rational arithmetic the two agree with no rounding at all, because the relation is an algebraic identity between the two fits rather than an approximation.

Step 4: what the short regression says about itself.

short (omits z )long (includes z )
coefficient on x 2.449301 1.116667
R 2 0.992876 0.999175

The short regression explains 99.3% of the variation in y . By every conventional reading it is an excellent model, and its coefficient on x is 119.3 % too large. The fit statistic is not failing to warn; it is answering a different question, how well this equation tracks these data, and answering it correctly.

A caution about the long regression. Its coefficients are 1.1167 and 2.7224 against structural values of 1.0 and 3.0 . They do not recover the truth exactly, because in a sample of twelve the disturbance is not orthogonal to the regressors. The identity in step 3 is unaffected: it relates the two fits to each other, and says nothing about either being unbiased for a structural parameter.

---

Step 5: an instrument with a strong first stage that repairs nothing. Take w correlated with x :

first-stage slope = 1.5 , first-stage  R 2 = 0.986014 .

By the usual rule of thumb this is a strong instrument. The estimate it produces is

β ^ IV = S w y S w x = 576 235 = 2.451064 ,

against a structural value of 1.0 . That is no better than the biased least-squares estimate of 2.449301 it was introduced to repair.

The reason is visible in the construction and nowhere in the output: w was built to move with x , and x moves with z , so w carries a relationship to the omitted determinant. The exclusion restriction fails. No statistic reported by the procedure detects this, and the strong first stage is silent about it, because relevance and exclusion are different conditions.

What this example is not. It is not a demonstration that instrumental variables do not work. It is a demonstration that the checkable condition is the less important one, and that an instrument's credibility rests on an argument about where its variation comes from, which is why published applications spend their space on that argument rather than on the first stage.

Contrast

Pairs that differ in one respect

The same data, with and without the confounder.

short regressionlong regression
coefficient on x 2.449301 1.116667
R 2 0.992876 0.999175
specification erroromits z none stated

The difference in coefficients is 1.332634 , which equals β ^ 2 δ = 2.722381 × 0.489510 exactly. The difference in R 2 is 0.0063 . One of these two comparisons tells you the coefficient was wrong by more than a factor of two; the other is the one usually reported.

Fit against bias, as questions.

Fit asks whether the equation as specified tracks the data. Bias asks whether the coefficient measures the effect. The first is answered inside the model and the second cannot be, which is why improving one carries no information about the other. A specification search that maximises R 2 optimises the question nobody asked.

Relevance against exclusion.

relevanceexclusion
statementinstrument correlates with x instrument reaches y only through x
involvestwo observed variablesthe unobserved disturbance
sample analogueyes, the first stagenone
how it is defendeda statistican argument

The worked instrument has first-stage R 2 = 0.986014 , which is strong by any conventional threshold, and produces 2.451064 against a structural 1 . It satisfies the checkable condition completely and fails the other one, and the output looks the same either way.

Omitted variables against simultaneity.

Both break exogeneity and they call for different responses. An omitted variable can in principle be measured and included, so the repair is a better dataset. Simultaneity is a property of how the system determines its variables, so no control fixes it and the repair must come from a design. An instrument, or variation whose direction of causation is known.

Measurement error in the regressor against error in the outcome.

Error in x attenuates the coefficient toward zero. Error in y enlarges the standard errors and leaves the coefficient unbiased. They look similar in a data dictionary and have opposite consequences for how a null result should be read.

Warning

Reassurances that are not reassurances

A high R 2 . Computed from the included regressors, so it cannot register an omitted one. The worked short regression reaches 0.992876 with a coefficient 119.3 % too large. A model can track the data almost perfectly and answer the wrong question.

Tight standard errors. They describe how much the coefficient would vary across samples under the specification as given. A precisely estimated biased coefficient is precisely wrong, and narrow intervals make it look more credible rather than less.

A large sample. More data shrinks the standard error and leaves the bias where it is. Asymptotically the estimator converges to the wrong number with increasing confidence.

Passing residual diagnostics. Tests for heteroscedasticity, normality and functional form all take the equation as specified. None of them can detect a determinant that is not in it.

Adding controls until the coefficient stabilises. A coefficient that stops moving may have converged on the effect, or the remaining confounders may simply resemble each other. Stability across the controls that happen to be in the dataset is not evidence about the ones that are not.

---

And two specific to instruments.

A strong first stage. It establishes relevance and nothing else. The worked instrument has R 2 = 0.986014 in the first stage and returns 2.451064 where the structural coefficient is 1 . No improvement on the biased least-squares estimate of 2.449301 it was meant to repair, because it carries a relationship to the omitted determinant. The two conditions are independent, and only one has a statistic.

An overidentification test. With more instruments than endogenous regressors, such a test can reject when the instruments disagree. Failing to reject is weak evidence at best: instruments sharing the same flaw agree with each other, and the test is silent when all of them are invalid in the same way.

---

What does count. A stated identification argument: where the variation in the regressor comes from, which determinants of the outcome could travel with it, and why the proposed instrument's origin is unconnected to them. That argument can be wrong, and it can be examined, which is more than can be said for a fit statistic.

Application

Where the identification argument is the contribution

Returns to schooling. The regression is trivial to run and has been for a century; the literature is about identification. Compulsory-schooling laws, distance to the nearest college and quarter of birth have all been used as instruments, and each is argued over on the exclusion restriction rather than on its first stage. The estimates that survive are the ones whose arguments survive.

Minimum wage and employment. Comparing jurisdictions that raised a minimum wage against those that did not confounds the policy with whatever led to its adoption, since places raising wages differ economically from places that do not. The designs that carry weight compare adjacent counties across a state border, where the argument is that local conditions are similar and the policy difference is administrative.

Advertising and sales. Firms advertise more when they expect to sell more, so the regressor responds to the same expectations that move the outcome. Randomised holdout regions exist because the observational coefficient conflates the effect of advertising with the forecast that prompted it, and no control for past sales resolves that.

Class size and attainment. Pupils are not assigned to classes at random: schools allocate by ability, resources and parental pressure. The well-known studies exploit administrative rules, a maximum class size that forces a split at a threshold, because the rule generates variation whose origin is documented and arguably unrelated to pupil characteristics.

Credit scoring and default. A model predicting default from observed borrower characteristics can be excellent for ranking applicants while its coefficients are not effects, since lenders' past decisions shaped who appears in the data at all. Using such coefficients to argue that changing a characteristic would change default risk is the error this unit is about, and it is common precisely because the predictive performance is genuinely good.

---

The recurring shape. In each case the estimation is routine and the contribution is the argument about where the variation came from. That is why published work in this area spends its length on the design and its threats rather than on the regression table, and why a reported coefficient without such an argument should be read as a description of an association.

Next step

Practice When a Coefficient Is Not an Effect

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lesson

This is the last lesson in Econometrics. Review the course map to see what is left.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.