Choosing a Regression Form
What the response variable and the shape of the relationship each rule out, how polynomial terms curve a fit while leaving it linear in its coefficients, why exponential and Poisson models are read on a multiplicative scale, and what a fitted form does beyond the range it was fitted on.
Definition
A regression model states how the expected value of a response depends on predictors. The forms below differ in what they assume about the response and about the shape of that dependence.
Polynomial regression adds powers of a predictor as further columns:
Despite the curve it produces, this is still a linear model: linear means linear in the coefficients
Exponential regression applies when the response is multiplied rather than incremented as a predictor advances. Fitting
so
Poisson regression applies when the response is a count. It models
so
Parametric and nonparametric. All three fix a functional form in advance and estimate finitely many coefficients, which is what parametric means. A nonparametric fit lets the data determine the shape, at the cost of more data and no coefficients to report.
Assumptions and scope
Ordinary least squares assumes the error variance is constant across observations. Count responses violate this systematically, since their variance rises with their mean.
Fitting
by least squares minimises squared error on the log scale, not on the original scale. The result is not the same as fitting an exponential curve directly, and the difference matters when a few large values dominate. is undefined at . A dataset with zero counts cannot be log-transformed without an arbitrary offset, which is one reason Poisson regression models the log of the mean rather than the log of the response.A polynomial of degree
has turning points available to it and behaves likefar from the data. Every polynomial fit diverges outside the observed range, and higher degree diverges faster. Choosing a form by looking at the data and then reporting that form's fit as if it had been specified in advance is the selection bias of the previous unit, at the level of the model family.
Worked material
Example
Five responses, and what each rules out
Each case names a response and a predictor. The question is which forms remain available once the response is taken seriously.
Number of emergency admissions per day, against the day's mean temperature. A count, bounded below at zero, typically small. A straight line predicts negative admissions at some temperatures, and the spread around the fit is visibly wider on busy days than on quiet ones. Poisson regression, with
Bacterial population, against hours since inoculation. Positive, and multiplied rather than incremented while growth is unconstrained. Successive ratios roughly constant is the check. Exponential regression on
Braking distance, against speed. Continuous, positive, and curved for a reason that is known in advance: kinetic energy rises with the square of speed. Polynomial regression with an
Proportion of a cohort still subscribed, against months elapsed. Bounded in
Household electricity use, against outdoor temperature. Continuous and non-monotone: consumption falls as temperature rises toward a comfortable range, then climbs again as cooling starts. No single-turning-point form fits. A quadratic can represent one turning point and will place it somewhere between the two real ones; the available options are a higher-degree polynomial chosen on held-out error, or a nonparametric fit that lets the data set the shape.
---
In each case the response's nature narrowed the field before any curve was drawn, and in two of them, the proportion and the non-monotone case, it ruled out every form this unit offers. Knowing that a form is unavailable is as useful as knowing which to pick, and it is only visible by asking what the response can be.
Contrast
Pairs that look alike and commit to different things
| log-scale fit | linear fit | |
|---|---|---|
| claim | ||
| coefficient reads as | a factor or percentage | an amount in the units of |
| minimises | squared error in | squared error in |
| undefined | permitted |
The third row is the one most often overlooked. Least squares on
Poisson regression against least squares on
Both handle a count response and produce a multiplicative reading. The second requires an arbitrary offset to cope with zero counts, and the choice of offset changes the estimate, adding 1 rather than 0.5 gives a different slope with no principle deciding between them. Poisson regression models the log of the mean, which is defined when an observed count is zero, so no offset is required.
A degree-2 fit against a degree-6 fit on the same eight points.
| quadratic | degree 6 | |
|---|---|---|
| smaller still | ||
| prediction at | ||
| direction past the data | rising | falling |
The more flexible fit is better by the criterion computed on the training data and disagrees by roughly a quarter two units beyond it, in the opposite direction. Training fit and extrapolation behaviour are not merely different questions; here they rank the two models oppositely.
Choosing a form from the mechanism against choosing it from the plot.
Braking distance takes an
In each, two procedures produce comparable-looking output and commit to different claims: about scale, about what is defined at zero, about behaviour outside the data, or about what an error figure estimates. None of the differences is visible in the fitted line alone.
Common errors
Common misconception
That adding
Common misconception
That a coefficient from a model fitted on the log scale reports how much the response changes per unit of the predictor. A slope of
Related units
Requires
Connected
- The Natural Logarithm (used by)