When a Coefficient Is Not an Effect

Why a least-squares coefficient estimated from observational data need not be the causal effect, the three mechanisms that break the exogeneity assumption, the formula that gives the size and direction of omitted-variable bias, and what an instrumental variable would have to satisfy to repair it.

Definition

Least squares estimates the coefficients of the conditional expectation of y given the included regressors. It estimates a causal effect only under the additional assumption that the regressor is uncorrelated with everything else affecting the outcome, written E [ u ∣ x ] = 0 and called exogeneity.

Three mechanisms break it.

Omitted variables. A determinant z of y that is correlated with x and left out of the model. Fitting the short regression y on x alone gives

β ^ short = β ^ long + β 2 δ ,

where β 2 is z 's coefficient in the long regression and δ is the slope from regressing z on x . The product β 2 δ is the omitted-variable bias, and its sign follows from the signs of its two factors.

Simultaneity. y and x are determined together, so x responds to the same disturbance that moves y . Price and quantity in a market are the standard case.

Measurement error in the regressor. If x is observed with independent noise, the estimated coefficient is pulled toward zero, an effect called attenuation.

An instrumental variable w repairs the problem when it satisfies two conditions. Relevance: w is correlated with x , which is checkable from the data through the first-stage regression. Exclusion: w affects y only through x , and in particular is uncorrelated with the omitted determinants, which is not checkable from the data. The estimator is

β ^ IV = S w y S w x .

Assumptions and scope

  • The bias formula β ^ short − β ^ long = β 2 δ is an algebraic identity between two fits on the same sample. It holds exactly, whatever the data-generating process, and does not by itself say that the long regression recovers a causal effect.

  • Including a control removes that control's contribution to the bias. It says nothing about determinants still omitted, so a coefficient that moves when a control is added is evidence of confounding rather than evidence that the confounding has been eliminated.

  • Measurement error in the regressor attenuates the coefficient toward zero under classical assumptions. Error in the outcome inflates the standard errors without biasing the coefficient, so the two are not interchangeable.

  • Relevance is checkable and exclusion is not. A first-stage F statistic or R 2 bears only on relevance, and reporting one is not evidence about the exclusion restriction.

  • A weak instrument produces an estimator that is biased toward the least-squares estimate and whose conventional standard errors understate the uncertainty, so weak relevance is a problem distinct from a failed exclusion restriction.

  • An instrumental variables estimate identifies an effect for the subpopulation whose regressor value responds to the instrument, which need not be the population of interest.

Worked material

Example

Five regressions and what each coefficient identifies

Earnings on years of schooling. The classic case. Whatever leads someone to stay in education, family resources, prior attainment, expectations, also bears on earnings, so the omitted determinants travel with the regressor. The coefficient is upward biased if those determinants raise both, and the sign follows from that reasoning rather than from the data.

Hospital admission on health outcome. People admitted to hospital are sicker than those who are not, so a regression of mortality on admission finds a positive coefficient. Reading it as the effect of being admitted inverts the causation: the regressor responds to the same underlying condition driving the outcome. Here the bias is not subtle and the direction is obvious, which is what makes it a useful case for seeing the mechanism.

Quantity sold on price. Price and quantity are determined together by supply and demand, so neither is exogenous to the other. A regression traces out neither curve but a mixture whose composition depends on which side of the market moved more. This is simultaneity, and no set of controls addresses it. The problem is the structure of the system rather than a missing variable.

Consumption on reported income. If income is measured with error, recall error in a survey, say, the coefficient is attenuated toward zero under classical assumptions. The estimate understates the relationship, which makes a small or insignificant result ambiguous between a weak effect and a well-measured one.

Crop yield on rainfall, in an agricultural trial with assigned irrigation. Where the regressor was assigned by the experimenter, exogeneity holds by design and the coefficient does estimate the effect. This is the case that shows the others are about how the data arose rather than about regression as a technique.

---

The first two have omitted determinants, one obvious and one less so. The third has no missing variable at all and is still not identified. The fourth is biased toward zero rather than away from it. The fifth is fine, and the only thing distinguishing it is that somebody controlled the assignment. In every case the diagnosis came from the setting rather than from the output.

Contrast

Pairs that differ in one respect

The same data, with and without the confounder.

short regressionlong regression
coefficient on x 2.449301 1.116667
R 2 0.992876 0.999175
specification erroromits z none stated

The difference in coefficients is 1.332634 , which equals β ^ 2 δ = 2.722381 × 0.489510 exactly. The difference in R 2 is 0.0063 . One of these two comparisons tells you the coefficient was wrong by more than a factor of two; the other is the one usually reported.

Fit against bias, as questions.

Fit asks whether the equation as specified tracks the data. Bias asks whether the coefficient measures the effect. The first is answered inside the model and the second cannot be, which is why improving one carries no information about the other. A specification search that maximises R 2 optimises the question nobody asked.

Relevance against exclusion.

relevanceexclusion
statementinstrument correlates with x instrument reaches y only through x
involvestwo observed variablesthe unobserved disturbance
sample analogueyes, the first stagenone
how it is defendeda statistican argument

The worked instrument has first-stage R 2 = 0.986014 , which is strong by any conventional threshold, and produces 2.451064 against a structural 1 . It satisfies the checkable condition completely and fails the other one, and the output looks the same either way.

Omitted variables against simultaneity.

Both break exogeneity and they call for different responses. An omitted variable can in principle be measured and included, so the repair is a better dataset. Simultaneity is a property of how the system determines its variables, so no control fixes it and the repair must come from a design. An instrument, or variation whose direction of causation is known.

Measurement error in the regressor against error in the outcome.

Error in x attenuates the coefficient toward zero. Error in y enlarges the standard errors and leaves the coefficient unbiased. They look similar in a data dictionary and have opposite consequences for how a null result should be read.

Common errors

Common misconception

That a regression with a high R 2 , small residuals and tight standard errors is unlikely to suffer from omitted-variable bias. Every one of those quantities is computed from the fitted model using the included regressors, so none of them can register the influence of a determinant that was left out. On a worked dataset the short regression omitting a confounder attains R 2 = 0.992876 while its slope is 2.449301 against 1.116667 once the confounder is included, wrong by 119.3 % , with the entire gap accounted for by the bias formula β 2 δ = 1.332634 . Fit describes how well the equation as specified tracks the data; bias describes whether the coefficient answers the question asked, and no diagnostic inside the regression output speaks to the second.

Common misconception

That a strong first stage, a high F statistic or R 2 when the regressor is regressed on the instrument, establishes that the instrument is valid. The first stage tests relevance only, which is one of two conditions. The other, that the instrument affects the outcome solely through the regressor, cannot be checked against the data at all and must be argued from knowledge of how the variables arise. On a worked dataset an instrument with first-stage R 2 = 0.986014 yields an instrumental-variables estimate of 2.451064 where the structural coefficient is 1 , because the instrument is itself correlated with the omitted determinant. It is no better than the biased least-squares estimate of 2.449301 it was introduced to repair, and nothing in the regression output signals the failure.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.