Module 1 of 1 · Lesson 1 of 1
When a Coefficient Is Not an Effect
Why a high
What you will be able to do
The learner can identify the sources of endogeneity that make a least-squares coefficient differ from the causal effect, compute the direction and size of omitted-variable bias from the bias formula, explain why goodness of fit carries no information about that bias, and state what an instrumental variable would have to satisfy to repair it.
Orientation
Association against causal effect in a fitted coefficient
A fitted regression reports how the outcome differs across units that differ in the regressor. The question usually being asked is different: what would happen to the outcome if the regressor were changed. The two coincide only when nothing else systematically differs alongside the regressor, and in observational data that condition is an assumption rather than a finding.
This unit is about how far apart the two can be, and what the gap is made of.
The gap has a formula for one of its sources. Leaving out a determinant of the outcome that travels with the regressor shifts the coefficient by an amount that can be computed exactly: how much the omitted variable matters, times how much it moves with the included one.
None of the model's own diagnostics can see it. On the worked data later in this unit, a regression omitting a confounder reaches
The repair is not more controls. Adding a control removes that control's contribution and nothing else. What remains unobserved keeps contributing, and no improvement in fit indicates whether anything does.
The unit ends with instrumental variables, which look like a technical remedy and are mostly an argument: one of the two conditions an instrument must satisfy can be checked in the data, and the one that matters cannot.
Definition
Exogeneity, the bias formula, and the instrument conditions
The canonical statements above give exogeneity, the bias formula and the two instrument conditions. What follows is the status of each claim, since that is what decides how it can be defended.
Exogeneity is an assumption about unobservables.
The bias formula is an algebraic identity, not a modelling assumption.
holds exactly between two least-squares fits on the same sample, whatever generated the data. It is therefore always available as a decomposition of why a coefficient moved when a control was added. What it does not say is that the long regression is unbiased. It relates two fits to each other, and both may omit something.
| Claim | Status | How it is settled |
|---|---|---|
| identity | algebra, exact | |
| assumption | argument about the setting | |
| instrument relevance | testable | first-stage regression |
| exclusion restriction | assumption | argument about the setting |
The three mechanisms differ in what they do to the coefficient. An omitted determinant moves it by
Relevance and exclusion are asymmetric in kind, not merely in difficulty. Relevance is a statement about two observed variables and has a sample analogue. Exclusion says the instrument has no path to the outcome except through the regressor, which involves the unobserved disturbance and so has no sample analogue at all. Reporting a first-stage
A failed exclusion restriction is not detectable from the output. An instrument correlated with the omitted determinant produces an estimate that is wrong in a way no diagnostic flags, and a strong first stage does not protect against it. The worked example exhibits exactly this, with a first-stage
Intuition
The source of omitted-variable bias
Least squares finds the line that best tracks the data it was given. If a determinant of the outcome is missing from the equation and travels with the included regressor, the fitting procedure has no way to distinguish the two influences, so it attributes both to the regressor it can see.
That attribution has a size, and the size is a product of two things. How much the omitted variable matters for the outcome,
Either factor being zero leaves the coefficient untouched. An omitted variable that matters enormously for
The fit statistics are computed inside the equation.
So there is no tension between an excellent fit and a badly wrong coefficient. On the worked data the short regression achieves
Adding controls removes one term at a time. Including
An instrument works by changing where the variation comes from. Rather than trying to name and measure every confounder, it isolates a slice of the regressor's movement whose origin is known and argued to be unconnected to the outcome except through the regressor. The strength of that slice is measurable. Whether its origin really is unconnected is not, and the worked example shows an instrument with a first-stage
Example
Five regressions and what each coefficient identifies
Earnings on years of schooling. The classic case. Whatever leads someone to stay in education, family resources, prior attainment, expectations, also bears on earnings, so the omitted determinants travel with the regressor. The coefficient is upward biased if those determinants raise both, and the sign follows from that reasoning rather than from the data.
Hospital admission on health outcome. People admitted to hospital are sicker than those who are not, so a regression of mortality on admission finds a positive coefficient. Reading it as the effect of being admitted inverts the causation: the regressor responds to the same underlying condition driving the outcome. Here the bias is not subtle and the direction is obvious, which is what makes it a useful case for seeing the mechanism.
Quantity sold on price. Price and quantity are determined together by supply and demand, so neither is exogenous to the other. A regression traces out neither curve but a mixture whose composition depends on which side of the market moved more. This is simultaneity, and no set of controls addresses it. The problem is the structure of the system rather than a missing variable.
Consumption on reported income. If income is measured with error, recall error in a survey, say, the coefficient is attenuated toward zero under classical assumptions. The estimate understates the relationship, which makes a small or insignificant result ambiguous between a weak effect and a well-measured one.
Crop yield on rainfall, in an agricultural trial with assigned irrigation. Where the regressor was assigned by the experimenter, exogeneity holds by design and the coefficient does estimate the effect. This is the case that shows the others are about how the data arose rather than about regression as a technique.
---
The first two have omitted determinants, one obvious and one less so. The third has no missing variable at all and is still not identified. The fourth is biased toward zero rather than away from it. The fifth is fine, and the only thing distinguishing it is that somebody controlled the assignment. In every case the diagnosis came from the setting rather than from the output.
Procedure
Working out what a coefficient identifies
Before fitting anything, state the question. Write the quantity of interest as a sentence about an intervention: what would happen to
To assess exogeneity.
- List the determinants of
that subject-matter knowledge suggests, whether or not they are in the data. - For each, ask whether it plausibly moves with
. A determinant unrelated to the regressor causes no bias however important it is for the outcome. - Ask whether
could respond to or to the same disturbance. The simultaneity case. - Ask how
was measured. Classical error in the regressor attenuates the coefficient toward zero. - Record the answers. This list is the identification argument, and it is what a reader needs in order to judge the estimate.
To quantify what one omitted variable does, when it is available.
- Fit the short regression,
on . - Fit the long regression,
on and , recording . - Regress
on to obtain . - Confirm
. Disagreement means an arithmetic error, since the relation is an identity. - Sign it before computing it where possible: the direction follows from whether
raises or lowers , and whether rises or falls with .
To use this as evidence about what remains. A coefficient that moves substantially when one control is added shows that controls matter here, which is a reason to expect that unobserved ones do too. A coefficient that barely moves is weak evidence in the other direction and is not a demonstration of exogeneity.
To assess a proposed instrument.
- Relevance. Regress
on the instrument and report the first-stage statistic. A weak first stage produces an estimator biased toward least squares with standard errors that understate the uncertainty. - Exclusion. Write the argument: where does the instrument's variation come from, and why can it not reach
except through ? This is prose, not a statistic, and it is the part a reader should weigh most heavily. - Look for the clear-cut failures. Is the instrument correlated with any determinant already on the list from the exogeneity assessment? If so, the restriction fails and the strength of the first stage is irrelevant.
- State the population. The estimate identifies an effect for units whose
responds to the instrument, which may not be the population the question is about.
Checks. Confirm that adding a control changes the coefficient by exactly the amount the bias formula predicts, which validates the arithmetic. And before reporting any estimate as an effect, write down what would have to be true for that reading to hold; if the list cannot be written, the estimate is a description of an association and should be reported as one.
Worked example
An excellent fit with a coefficient wrong by 119%
Twelve observations. The outcome is generated as
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2 | 2 | 3 | 3 | 4 | 4 | 5 | 5 | 6 | 6 | 7 | 7 |
Step 1: the short regression, omitting
Step 2: the long regression, including
Step 3: the bias formula. Regressing
so the predicted bias is
And the observed difference:
Exactly equal. Computed in rational arithmetic the two agree with no rounding at all, because the relation is an algebraic identity between the two fits rather than an approximation.
Step 4: what the short regression says about itself.
| short (omits | long (includes | |
|---|---|---|
| coefficient on | ||
The short regression explains 99.3% of the variation in
A caution about the long regression. Its coefficients are
---
Step 5: an instrument with a strong first stage that repairs nothing. Take
By the usual rule of thumb this is a strong instrument. The estimate it produces is
against a structural value of
The reason is visible in the construction and nowhere in the output:
What this example is not. It is not a demonstration that instrumental variables do not work. It is a demonstration that the checkable condition is the less important one, and that an instrument's credibility rests on an argument about where its variation comes from, which is why published applications spend their space on that argument rather than on the first stage.
Contrast
Pairs that differ in one respect
The same data, with and without the confounder.
| short regression | long regression | |
|---|---|---|
| coefficient on | ||
| specification error | omits | none stated |
The difference in coefficients is
Fit against bias, as questions.
Fit asks whether the equation as specified tracks the data. Bias asks whether the coefficient measures the effect. The first is answered inside the model and the second cannot be, which is why improving one carries no information about the other. A specification search that maximises
Relevance against exclusion.
| relevance | exclusion | |
|---|---|---|
| statement | instrument correlates with | instrument reaches |
| involves | two observed variables | the unobserved disturbance |
| sample analogue | yes, the first stage | none |
| how it is defended | a statistic | an argument |
The worked instrument has first-stage
Omitted variables against simultaneity.
Both break exogeneity and they call for different responses. An omitted variable can in principle be measured and included, so the repair is a better dataset. Simultaneity is a property of how the system determines its variables, so no control fixes it and the repair must come from a design. An instrument, or variation whose direction of causation is known.
Measurement error in the regressor against error in the outcome.
Error in
Warning
Reassurances that are not reassurances
A high
Tight standard errors. They describe how much the coefficient would vary across samples under the specification as given. A precisely estimated biased coefficient is precisely wrong, and narrow intervals make it look more credible rather than less.
A large sample. More data shrinks the standard error and leaves the bias where it is. Asymptotically the estimator converges to the wrong number with increasing confidence.
Passing residual diagnostics. Tests for heteroscedasticity, normality and functional form all take the equation as specified. None of them can detect a determinant that is not in it.
Adding controls until the coefficient stabilises. A coefficient that stops moving may have converged on the effect, or the remaining confounders may simply resemble each other. Stability across the controls that happen to be in the dataset is not evidence about the ones that are not.
---
And two specific to instruments.
A strong first stage. It establishes relevance and nothing else. The worked instrument has
An overidentification test. With more instruments than endogenous regressors, such a test can reject when the instruments disagree. Failing to reject is weak evidence at best: instruments sharing the same flaw agree with each other, and the test is silent when all of them are invalid in the same way.
---
What does count. A stated identification argument: where the variation in the regressor comes from, which determinants of the outcome could travel with it, and why the proposed instrument's origin is unconnected to them. That argument can be wrong, and it can be examined, which is more than can be said for a fit statistic.
Application
Where the identification argument is the contribution
Returns to schooling. The regression is trivial to run and has been for a century; the literature is about identification. Compulsory-schooling laws, distance to the nearest college and quarter of birth have all been used as instruments, and each is argued over on the exclusion restriction rather than on its first stage. The estimates that survive are the ones whose arguments survive.
Minimum wage and employment. Comparing jurisdictions that raised a minimum wage against those that did not confounds the policy with whatever led to its adoption, since places raising wages differ economically from places that do not. The designs that carry weight compare adjacent counties across a state border, where the argument is that local conditions are similar and the policy difference is administrative.
Advertising and sales. Firms advertise more when they expect to sell more, so the regressor responds to the same expectations that move the outcome. Randomised holdout regions exist because the observational coefficient conflates the effect of advertising with the forecast that prompted it, and no control for past sales resolves that.
Class size and attainment. Pupils are not assigned to classes at random: schools allocate by ability, resources and parental pressure. The well-known studies exploit administrative rules, a maximum class size that forces a split at a threshold, because the rule generates variation whose origin is documented and arguably unrelated to pupil characteristics.
Credit scoring and default. A model predicting default from observed borrower characteristics can be excellent for ranking applicants while its coefficients are not effects, since lenders' past decisions shaped who appears in the data at all. Using such coefficients to argue that changing a characteristic would change default risk is the error this unit is about, and it is common precisely because the predictive performance is genuinely good.
---
The recurring shape. In each case the estimation is routine and the contribution is the argument about where the variation came from. That is why published work in this area spends its length on the design and its threats rather than on the regression table, and why a reported coefficient without such an argument should be read as a description of an association.