Module 2 of 8 · Lesson 1 of 1
Choosing a Regression Form
How the response type and the shape of the relationship restrict the model form.
What you will be able to do
The learner can choose a regression form appropriate to the response variable and the shape of the relationship, fit and interpret polynomial, exponential and Poisson models, and say what each assumes about the response that a linear model does not.
Orientation
Two questions that precede any fitting
A scatter plot suggests a shape, and it is tempting to choose a model from the shape alone. Two questions come first, and they often decide the matter before the picture is consulted.
What kind of quantity is the response? A count of visits cannot be negative and has no ceiling. A proportion lies between 0 and 1. A waiting time is positive and usually skewed. Ordinary least squares assumes none of this, so a straight line fitted to counts will predict negative visits for some inputs, and a fitted proportion can exceed 1. Those predictions are not merely inaccurate; they are outside the set of values the response can take.
Does a unit of the predictor add to the response, or multiply it? A fixed increase per unit is a straight line. A fixed factor per unit is exponential, straight only after taking logarithms, and its coefficient must then be reported as a ratio rather than as an amount.
This unit covers the forms those two questions select between: polynomial terms for a curved relationship on a continuous response, exponential regression for multiplicative growth, and Poisson regression for counts. It also covers what each form does outside the range it was fitted on, which is where the differences between them are largest and the data say least.
Definition
What each form assumes, and where the assumption bites
The canonical statements above give each form. What follows is what each one commits to beyond its formula.
Polynomial regression commits to a global shape. The fitted curve is one polynomial across the whole range, so the coefficients are determined by every observation at once. Data at one end influence the fit at the other, which is why a single outlying point can move the curve far from it. A degree-
Exponential regression commits to a constant proportional rate. The claim
Poisson regression commits to a variance structure. Modelling
| Form | Response it suits | Coefficient reads as | Fails when |
|---|---|---|---|
| linear | continuous, unbounded | amount added per unit | response is bounded or a count |
| polynomial | continuous, curved | no single interpretation | extrapolating, or degree chosen on training error |
| exponential | positive, multiplicative | factor per unit | response reaches zero |
| Poisson | counts | rate ratio | variance exceeds the mean |
What "parametric" supplies and costs. Each form fixes the shape in advance and estimates a handful of coefficients. That is what makes the result reportable, a rate ratio is a sentence, a fitted curve from a nonparametric smoother is not, and it is also what makes the result wrong in a specific way if the chosen shape is not the true one. The bias term of the previous unit is exactly this cost.
Intuition
Why a curved fit is still a linear model
The word linear in linear regression describes the coefficients, not the picture. This is worth pinning down because the picture is what a reader sees first.
Least squares solves for the
The test for whether a model is linear. Look at where the coefficients sit. In
Why exponential regression escapes that. Taking logarithms of
The cost of the conversion. A slope of
Counts are different in kind, not degree. For a count response, the problem is not the shape of the relationship but the spread around it. Observations with a small expected count vary little; observations with a large one vary a great deal. Least squares weights every observation equally because it assumes they vary equally, so it gives the noisiest observations the same influence as the quietest and then reports intervals computed from a single pooled variance that describes neither.
Example
Five responses, and what each rules out
Each case names a response and a predictor. The question is which forms remain available once the response is taken seriously.
Number of emergency admissions per day, against the day's mean temperature. A count, bounded below at zero, typically small. A straight line predicts negative admissions at some temperatures, and the spread around the fit is visibly wider on busy days than on quiet ones. Poisson regression, with
Bacterial population, against hours since inoculation. Positive, and multiplied rather than incremented while growth is unconstrained. Successive ratios roughly constant is the check. Exponential regression on
Braking distance, against speed. Continuous, positive, and curved for a reason that is known in advance: kinetic energy rises with the square of speed. Polynomial regression with an
Proportion of a cohort still subscribed, against months elapsed. Bounded in
Household electricity use, against outdoor temperature. Continuous and non-monotone: consumption falls as temperature rises toward a comfortable range, then climbs again as cooling starts. No single-turning-point form fits. A quadratic can represent one turning point and will place it somewhere between the two real ones; the available options are a higher-degree polynomial chosen on held-out error, or a nonparametric fit that lets the data set the shape.
---
In each case the response's nature narrowed the field before any curve was drawn, and in two of them, the proportion and the non-monotone case, it ruled out every form this unit offers. Knowing that a form is unavailable is as useful as knowing which to pick, and it is only visible by asking what the response can be.
Procedure
Selecting a form, fitting it, and reporting the coefficient
To choose a form.
- Name the response. Continuous and unbounded, positive and continuous, a count, a proportion, or a duration.
- Rule out on that basis. A count or a proportion excludes ordinary least squares on the raw response, whatever the scatter plot shows.
- Ask whether a unit of the predictor adds or multiplies. Constant successive differences indicate a line; constant successive ratios indicate an exponential form.
- Check for turning points. A relationship that changes direction
times needs a polynomial of degree at least , or a nonparametric fit. - Prefer a form the subject matter suggests over one found by trying several. Where several are genuinely plausible, choose by held-out error, and use a nested scheme if that same data must also report performance.
To fit a polynomial.
- Add columns
to the design matrix alongside . - Fit by ordinary least squares. The model is linear in its coefficients, so no special method is needed.
- Choose
by cross-validated error, not by , which falls with every added column. - Report the range of
the fit was estimated on.
Do not interpret individual polynomial coefficients. The slope at a point is
To fit an exponential form.
- Confirm every
. A zero makesundefined, and adding a constant to permit the transformation changes the model being fitted. - Regress
on by least squares, giving . - Convert: the multiplier is
and the per-unit factor is . - Report on the response scale. State
as a factor, as a percentage change , or as a doubling interval. Never quote alone as the effect.
To fit a count response.
- Use Poisson regression, which models
and so keeps the fitted mean positive. - Report
as a rate ratio: the factor by which the expected count is multiplied per unit of the predictor. - Check the variance assumption by comparing the variance of the residuals against the fitted means. Variance substantially exceeding the mean indicates overdispersion, which leaves point predictions usable and makes the reported intervals too narrow.
Checks that apply to all of them.
Plot the residuals against the fitted values. A visible curve means the chosen shape is wrong; a widening funnel means the variance assumption is wrong. These are different defects with different repairs, and the residual plot distinguishes them where a single error figure does not.
Worked example
Three fits, three scales to read them on
(a) A quadratic, fitted by ordinary least squares.
Eight observations,
The differences between successive values are
Build the design matrix with columns
Three coefficients from one least-squares solve, exactly as for a straight line. The negative linear coefficient alongside a positive quadratic one signals that neither coefficient has a standalone reading here, because the fitted slope at any point,
(b) An exponential, fitted on the log scale.
Six observations,
Successive ratios are
Convert back:
Reading it. The slope
(c) Counts, and why least squares reports the wrong uncertainty.
Eight observations in each of three exposure groups:
| Group | Counts | Mean | Sample variance | Variance / mean |
|---|---|---|---|---|
| low | 0, 1, 0, 2, 1, 0, 1, 1 | |||
| medium | 3, 5, 2, 4, 6, 3, 4, 5 | |||
| high | 12, 18, 9, 15, 21, 14, 11, 16 |
What this shows and what it does not. The variance rises steeply with the mean: from
The ratios do not demonstrate
The practical consequence. A least-squares fit to these counts would produce a fitted line not far from the right place, and standard errors computed from a pooled variance, too wide where counts are small, too narrow where they are large. The point estimate survives the wrong assumption; the uncertainty does not.
Contrast
Pairs that look alike and commit to different things
| log-scale fit | linear fit | |
|---|---|---|
| claim | ||
| coefficient reads as | a factor or percentage | an amount in the units of |
| minimises | squared error in | squared error in |
| undefined | permitted |
The third row is the one most often overlooked. Least squares on
Poisson regression against least squares on
Both handle a count response and produce a multiplicative reading. The second requires an arbitrary offset to cope with zero counts, and the choice of offset changes the estimate, adding 1 rather than 0.5 gives a different slope with no principle deciding between them. Poisson regression models the log of the mean, which is defined when an observed count is zero, so no offset is required.
A degree-2 fit against a degree-6 fit on the same eight points.
| quadratic | degree 6 | |
|---|---|---|
| smaller still | ||
| prediction at | ||
| direction past the data | rising | falling |
The more flexible fit is better by the criterion computed on the training data and disagrees by roughly a quarter two units beyond it, in the opposite direction. Training fit and extrapolation behaviour are not merely different questions; here they rank the two models oppositely.
Choosing a form from the mechanism against choosing it from the plot.
Braking distance takes an
In each, two procedures produce comparable-looking output and commit to different claims: about scale, about what is defined at zero, about behaviour outside the data, or about what an error figure estimates. None of the differences is visible in the fitted line alone.
Warning
Every fitted form leaves the data eventually
A fitted curve is supported by the range the observations cover. Outside it the form itself decides what happens, and the forms disagree most exactly where there is nothing to adjudicate between them.
A worked comparison. Take the eight observations of the worked example, which run to
| Fit | Prediction at |
|---|---|
| quadratic | |
| degree 6 |
The two disagree by about 24, roughly a quarter of the value, two units past the last observation. They disagree not because one is wrong on the data, both fit it well, but because extrapolation is decided by the functional form, and the flexible fit has curvature near the boundary that continues in a direction the data never tested. Here the degree-6 fit turns downward while the quadratic keeps rising.
Why higher degree is worse rather than better. Far from the data a degree-
The same caution for the other forms, with different shapes.
- An exponential fit continues multiplying. A model of early growth extrapolated far enough predicts values exceeding any physical bound, because nothing in the form knows about saturation.
- A Poisson fit predicts a positive mean everywhere, which is the right kind of answer, and says nothing about whether the rate relationship persists at exposures far from those observed.
- A linear fit is the most modest extrapolator, which is not a merit in itself: it is simply wrong more slowly.
What to do. State the range over which the fit was estimated whenever a fitted model is reported, and treat any prediction outside it as a claim about the functional form rather than a finding from the data. When an out-of-range prediction genuinely is the goal, that is an argument for a form chosen on subject-matter grounds, a saturating curve where saturation is expected, rather than for the form that fitted best inside the range.
And the selection point from the previous unit applies here too. Choosing the degree by comparing fits on the same data that then reports the error gives an optimistic figure. Degree is a flexibility parameter, so it is chosen by held-out error, nested if the same data must also report performance.
Application
Where the choice of form is the finding
Epidemic reporting. Early case counts are modelled on the log scale, and the reported quantity is a growth factor or doubling time rather than cases per day. The choice is what makes the numbers comparable across places with different population sizes and reporting dates, and it is also what makes the model wrong once transmission slows: an exponential form cannot represent a peak, so continuing to fit one past the turn produces confident forecasts in the wrong direction.
Insurance pricing. Claim counts per policy-year are modelled by Poisson regression, and coefficients are quoted as rate ratios: a factor applied to the base rate for a given characteristic. The multiplicative structure is not an approximation adopted for convenience, it matches how premiums are constructed, as a base rate multiplied by factors, so the model's parameters and the pricing table are the same objects.
Dose-response work. The response is bounded below at zero and often saturates, so a form is chosen for those properties and not for fit. A polynomial that happened to fit the observed doses better would predict falling response at high doses, an artefact of the leading term rather than a pharmacological claim.
Energy demand against temperature. The relationship turns, because heating and cooling both raise consumption. Fitting a single quadratic places one turning point between the two real ones and misestimates demand at both extremes. This is a case where no form in this unit applies, and the choice is between a higher-degree polynomial selected on held-out error and a nonparametric smoother.
Physical calibration. A sensor's response is fitted by a low-degree polynomial over a stated operating range, and the range is published with the coefficients. The convention exists because the calibration is used by people who did not collect the data, and a fitted polynomial carries no internal record of where it was supported.
---
The common thread. In each case the form was selected from what the response is and what is known about the mechanism, and the coefficient was reported on the scale that matches how the answer will be used: a doubling time, a rate ratio, a calibrated value with its range. Where the form was chosen to fit rather than to match, the failures show up outside the data, at high doses, past the epidemic peak, beyond the calibrated range, which is where a fitted curve is least able to warn anyone.