Choosing a Regression Form

What you will be able to do

The learner can choose a regression form appropriate to the response variable and the shape of the relationship, fit and interpret polynomial, exponential and Poisson models, and say what each assumes about the response that a linear model does not.

Orientation

Two questions that precede any fitting

A scatter plot suggests a shape, and it is tempting to choose a model from the shape alone. Two questions come first, and they often decide the matter before the picture is consulted.

What kind of quantity is the response? A count of visits cannot be negative and has no ceiling. A proportion lies between 0 and 1. A waiting time is positive and usually skewed. Ordinary least squares assumes none of this, so a straight line fitted to counts will predict negative visits for some inputs, and a fitted proportion can exceed 1. Those predictions are not merely inaccurate; they are outside the set of values the response can take.

Does a unit of the predictor add to the response, or multiply it? A fixed increase per unit is a straight line. A fixed factor per unit is exponential, straight only after taking logarithms, and its coefficient must then be reported as a ratio rather than as an amount.

This unit covers the forms those two questions select between: polynomial terms for a curved relationship on a continuous response, exponential regression for multiplicative growth, and Poisson regression for counts. It also covers what each form does outside the range it was fitted on, which is where the differences between them are largest and the data say least.

Definition

What each form assumes, and where the assumption bites

The canonical statements above give each form. What follows is what each one commits to beyond its formula.

Polynomial regression commits to a global shape. The fitted curve is one polynomial across the whole range, so the coefficients are determined by every observation at once. Data at one end influence the fit at the other, which is why a single outlying point can move the curve far from it. A degree- d polynomial also has exactly d − 1 turning points available, so it cannot represent a relationship that flattens and then flattens again without spending degree on it.

Exponential regression commits to a constant proportional rate. The claim y = A r x says the ratio between successive values is fixed. Fitting it by least squares on log ⁡ y makes a further commitment that is easy to miss: it minimises squared error on the log scale, so a factor-of-two error on a small value counts as heavily as a factor-of-two error on a large one. Fitting the same curve directly on the original scale minimises something different and gives different coefficients. Neither is wrong; they answer different questions, and the log-scale fit is the one with a closed-form solution.

Poisson regression commits to a variance structure. Modelling log ⁡ E [ y ∣ x ] rather than E [ y ∣ x ] guarantees the fitted mean stays positive, which matters for counts. The distributional assumption Var ⁡ ( y ) = E [ y ] is what makes the reported standard errors correct. Real count data frequently have variance exceeding the mean, and the model then understates uncertainty while still producing sensible point predictions. A failure that shows in the intervals rather than in the fitted line.

FormResponse it suitsCoefficient reads asFails when
linearcontinuous, unboundedamount added per unitresponse is bounded or a count
polynomialcontinuous, curvedno single interpretationextrapolating, or degree chosen on training error
exponentialpositive, multiplicativefactor per unitresponse reaches zero
Poissoncountsrate ratiovariance exceeds the mean

What "parametric" supplies and costs. Each form fixes the shape in advance and estimates a handful of coefficients. That is what makes the result reportable, a rate ratio is a sentence, a fitted curve from a nonparametric smoother is not, and it is also what makes the result wrong in a specific way if the chosen shape is not the true one. The bias term of the previous unit is exactly this cost.

Intuition

Why a curved fit is still a linear model

The word linear in linear regression describes the coefficients, not the picture. This is worth pinning down because the picture is what a reader sees first.

Least squares solves for the β j that minimise ∑ i ( y i − ∑ j β j z i j ) 2 , where z i j is the value of the j th column for observation i . Nothing in that expression cares what the columns contain. Fill them with x , x 2 and x 3 and the solution is found the same way, by the same formula, in the same single step. The curve appears because the columns are curved functions of x , not because the estimation changed.

The test for whether a model is linear. Look at where the coefficients sit. In y = β 0 + β 1 x + β 2 x 2 each β multiplies something and is added: linear. In y = β 0 e β 1 x the coefficient β 1 sits inside a function, so no rearrangement produces a linear system, and fitting it requires iteration rather than a formula. That is a genuinely nonlinear model, and there are few of them in routine work.

Why exponential regression escapes that. Taking logarithms of y = A r x gives log ⁡ y = log ⁡ A + x log ⁡ r , which is linear in log ⁡ A and log ⁡ r . The nonlinearity is absorbed by transforming the response, and least squares applies to the transformed problem. This is why the fitted quantity is a slope on the log scale and why reporting it requires converting back.

The cost of the conversion. A slope of 0.6931 on the log scale is e 0.6931 = 2.000 on the response scale: a doubling. Reported as "0.69" the effect sounds small. The number is the same; the scale it belongs to is what makes it legible, and a coefficient quoted without its scale is the most common way a correct fit becomes a wrong sentence.

Counts are different in kind, not degree. For a count response, the problem is not the shape of the relationship but the spread around it. Observations with a small expected count vary little; observations with a large one vary a great deal. Least squares weights every observation equally because it assumes they vary equally, so it gives the noisiest observations the same influence as the quietest and then reports intervals computed from a single pooled variance that describes neither.

Example

Five responses, and what each rules out

Each case names a response and a predictor. The question is which forms remain available once the response is taken seriously.

Number of emergency admissions per day, against the day's mean temperature. A count, bounded below at zero, typically small. A straight line predicts negative admissions at some temperatures, and the spread around the fit is visibly wider on busy days than on quiet ones. Poisson regression, with e β read as the multiplicative change in the daily rate per degree.

Bacterial population, against hours since inoculation. Positive, and multiplied rather than incremented while growth is unconstrained. Successive ratios roughly constant is the check. Exponential regression on log ⁡ y , with the coefficient reported as a factor per hour or a doubling time. The form fails once the culture saturates, which is a statement about the biology rather than about the fit.

Braking distance, against speed. Continuous, positive, and curved for a reason that is known in advance: kinetic energy rises with the square of speed. Polynomial regression with an x 2 term, chosen from the mechanism rather than from the scatter plot. That distinction matters. The same curve chosen by trying degrees until one fits well is the selection problem of the previous unit.

Proportion of a cohort still subscribed, against months elapsed. Bounded in [ 0 , 1 ] . Every form in this unit will eventually predict outside that interval, the linear one soonest. This case needs logistic regression, which is not covered here; recognising the bound is what shows the available forms are inadequate.

Household electricity use, against outdoor temperature. Continuous and non-monotone: consumption falls as temperature rises toward a comfortable range, then climbs again as cooling starts. No single-turning-point form fits. A quadratic can represent one turning point and will place it somewhere between the two real ones; the available options are a higher-degree polynomial chosen on held-out error, or a nonparametric fit that lets the data set the shape.

---

In each case the response's nature narrowed the field before any curve was drawn, and in two of them, the proportion and the non-monotone case, it ruled out every form this unit offers. Knowing that a form is unavailable is as useful as knowing which to pick, and it is only visible by asking what the response can be.

Procedure

Selecting a form, fitting it, and reporting the coefficient

To choose a form.

  1. Name the response. Continuous and unbounded, positive and continuous, a count, a proportion, or a duration.
  2. Rule out on that basis. A count or a proportion excludes ordinary least squares on the raw response, whatever the scatter plot shows.
  3. Ask whether a unit of the predictor adds or multiplies. Constant successive differences indicate a line; constant successive ratios indicate an exponential form.
  4. Check for turning points. A relationship that changes direction t times needs a polynomial of degree at least t + 1 , or a nonparametric fit.
  5. Prefer a form the subject matter suggests over one found by trying several. Where several are genuinely plausible, choose by held-out error, and use a nested scheme if that same data must also report performance.

To fit a polynomial.

  1. Add columns x 2 , … , x d to the design matrix alongside x .
  2. Fit by ordinary least squares. The model is linear in its coefficients, so no special method is needed.
  3. Choose d by cross-validated error, not by RSS , which falls with every added column.
  4. Report the range of x the fit was estimated on.

Do not interpret individual polynomial coefficients. The slope at a point is β 1 + 2 β 2 x + … , so it depends on where you stand; report fitted values or slopes at stated values of x instead.

To fit an exponential form.

  1. Confirm every y > 0 . A zero makes log ⁡ y undefined, and adding a constant to permit the transformation changes the model being fitted.
  2. Regress log ⁡ y on x by least squares, giving log ⁡ y = a + b x .
  3. Convert: the multiplier is A = e a and the per-unit factor is r = e b .
  4. Report on the response scale. State r as a factor, as a percentage change 100 ( r − 1 ) % , or as a doubling interval log ⁡ 2 / b . Never quote b alone as the effect.

To fit a count response.

  1. Use Poisson regression, which models log ⁡ E [ y ∣ x ] and so keeps the fitted mean positive.
  2. Report e β as a rate ratio: the factor by which the expected count is multiplied per unit of the predictor.
  3. Check the variance assumption by comparing the variance of the residuals against the fitted means. Variance substantially exceeding the mean indicates overdispersion, which leaves point predictions usable and makes the reported intervals too narrow.

Checks that apply to all of them.

Plot the residuals against the fitted values. A visible curve means the chosen shape is wrong; a widening funnel means the variance assumption is wrong. These are different defects with different repairs, and the residual plot distinguishes them where a single error figure does not.

Worked example

Three fits, three scales to read them on

(a) A quadratic, fitted by ordinary least squares.

Eight observations, x = 1 , … , 8 :

y = 2.0 ,   4.6 ,   9.1 ,   15.8 ,   24.5 ,   35.0 ,   47.9 ,   62.6 .

The differences between successive values are 2.6 , 4.5 , 6.7 , 8.7 , 10.5 , 12.9 , 14.7 , not constant, so no straight line fits, while the differences of those differences are close to constant at about 2. A constant second difference is the signature of a quadratic.

Build the design matrix with columns 1 , x , x 2 and solve the normal equations:

y ^ = 1.4946 − 0.4994 x + 1.0173 x 2 , RSS = 0.0272 .

Three coefficients from one least-squares solve, exactly as for a straight line. The negative linear coefficient alongside a positive quadratic one signals that neither coefficient has a standalone reading here, because the fitted slope at any point, − 0.4994 + 2 ( 1.0173 ) x , depends on where you stand.

(b) An exponential, fitted on the log scale.

Six observations, x = 0 , … , 5 :

y = 3.0 ,   6.1 ,   11.8 ,   24.5 ,   48.2 ,   97.0 .

Successive ratios are 2.03 , 1.93 , 2.08 , 1.97 , 2.01 , near-constant, which is the multiplicative signature. Take logarithms and fit a straight line:

log ⁡ y = 1.100699 + 0.694637 x .

Convert back:

y = e 1.100699 ( e 0.694637 ) x = 3.0063 × ( 2.00298 ) x .

Reading it. The slope 0.694637 is not an increase of 0.69 in y . Exponentiated it gives 2.0030 : each unit of x multiplies y by about 2 , a rise of 100.30 % . The doubling interval is log ⁡ 2 / 0.694637 = 0.9979 units. Three ways of saying one thing, none of which is " y rises by 0.69 ".

(c) Counts, and why least squares reports the wrong uncertainty.

Eight observations in each of three exposure groups:

GroupCountsMeanSample varianceVariance / mean
low0, 1, 0, 2, 1, 0, 1, 1 0.750 0.500 0.667
medium3, 5, 2, 4, 6, 3, 4, 5 4.000 1.714 0.429
high12, 18, 9, 15, 21, 14, 11, 16 14.500 15.143 1.044

What this shows and what it does not. The variance rises steeply with the mean: from 0.500 to 15.143 , a factor of about 30 , as the mean rises by a factor of about 19 . That is the finding, and it is enough to defeat ordinary least squares, which assumes one constant variance for all three groups and would pool them into a single figure describing none.

The ratios do not demonstrate Var ⁡ ( y ) = E [ y ] . Two of the three sit well below 1. With eight observations per group a sample variance is itself highly variable, so these ratios are consistent with the Poisson relationship without establishing it. Confirming that would need far more data per group, or a formal test.

The practical consequence. A least-squares fit to these counts would produce a fitted line not far from the right place, and standard errors computed from a pooled variance, too wide where counts are small, too narrow where they are large. The point estimate survives the wrong assumption; the uncertainty does not.

Contrast

Pairs that look alike and commit to different things

log ⁡ y = a + b x against y = β 0 + β 1 x .

log-scale fitlinear fit
claim y multiplied by e b per unit y increased by β 1 per unit
coefficient reads asa factor or percentagean amount in the units of y
minimisessquared error in log ⁡ y squared error in y
y = 0 undefinedpermitted

The third row is the one most often overlooked. Least squares on log ⁡ y treats a factor-of-two error on a value of 3 as equal to a factor-of-two error on a value of 300. Fitting the same exponential curve directly on the original scale weights the large values far more heavily and gives different coefficients. Both are defensible; they are not the same fit.

Poisson regression against least squares on log ⁡ ( y + 1 ) .

Both handle a count response and produce a multiplicative reading. The second requires an arbitrary offset to cope with zero counts, and the choice of offset changes the estimate, adding 1 rather than 0.5 gives a different slope with no principle deciding between them. Poisson regression models the log of the mean, which is defined when an observed count is zero, so no offset is required.

A degree-2 fit against a degree-6 fit on the same eight points.

quadraticdegree 6
RSS on the data 0.0272 smaller still
prediction at x = 10 98.23 74.58
direction past the datarisingfalling

The more flexible fit is better by the criterion computed on the training data and disagrees by roughly a quarter two units beyond it, in the opposite direction. Training fit and extrapolation behaviour are not merely different questions; here they rank the two models oppositely.

Choosing a form from the mechanism against choosing it from the plot.

Braking distance takes an x 2 term because kinetic energy goes as the square of speed. A reason available before the data. The same term chosen by fitting degrees 1 through 6 and keeping the best is a selection, and the reported error from that search is optimistic by the amount the previous unit describes. The fitted curve can be identical; what differs is what the error figure means.

In each, two procedures produce comparable-looking output and commit to different claims: about scale, about what is defined at zero, about behaviour outside the data, or about what an error figure estimates. None of the differences is visible in the fitted line alone.

Warning

Every fitted form leaves the data eventually

A fitted curve is supported by the range the observations cover. Outside it the form itself decides what happens, and the forms disagree most exactly where there is nothing to adjudicate between them.

A worked comparison. Take the eight observations of the worked example, which run to x = 8 . Fit a quadratic and a degree-6 polynomial to the same points. Both pass close to every observation. Ask each for a prediction at x = 10 :

FitPrediction at x = 10
quadratic 98.23
degree 6 74.58

The two disagree by about 24, roughly a quarter of the value, two units past the last observation. They disagree not because one is wrong on the data, both fit it well, but because extrapolation is decided by the functional form, and the flexible fit has curvature near the boundary that continues in a direction the data never tested. Here the degree-6 fit turns downward while the quadratic keeps rising.

Why higher degree is worse rather than better. Far from the data a degree- d polynomial behaves like its leading term, x d . Higher degree therefore diverges faster, and the coefficient of that leading term is estimated from whatever wiggle the fit needed near the edges of the observed range. Flexibility that helped inside the data becomes leverage outside it.

The same caution for the other forms, with different shapes.

  • An exponential fit continues multiplying. A model of early growth extrapolated far enough predicts values exceeding any physical bound, because nothing in the form knows about saturation.
  • A Poisson fit predicts a positive mean everywhere, which is the right kind of answer, and says nothing about whether the rate relationship persists at exposures far from those observed.
  • A linear fit is the most modest extrapolator, which is not a merit in itself: it is simply wrong more slowly.

What to do. State the range over which the fit was estimated whenever a fitted model is reported, and treat any prediction outside it as a claim about the functional form rather than a finding from the data. When an out-of-range prediction genuinely is the goal, that is an argument for a form chosen on subject-matter grounds, a saturating curve where saturation is expected, rather than for the form that fitted best inside the range.

And the selection point from the previous unit applies here too. Choosing the degree by comparing fits on the same data that then reports the error gives an optimistic figure. Degree is a flexibility parameter, so it is chosen by held-out error, nested if the same data must also report performance.

Application

Where the choice of form is the finding

Epidemic reporting. Early case counts are modelled on the log scale, and the reported quantity is a growth factor or doubling time rather than cases per day. The choice is what makes the numbers comparable across places with different population sizes and reporting dates, and it is also what makes the model wrong once transmission slows: an exponential form cannot represent a peak, so continuing to fit one past the turn produces confident forecasts in the wrong direction.

Insurance pricing. Claim counts per policy-year are modelled by Poisson regression, and coefficients are quoted as rate ratios: a factor applied to the base rate for a given characteristic. The multiplicative structure is not an approximation adopted for convenience, it matches how premiums are constructed, as a base rate multiplied by factors, so the model's parameters and the pricing table are the same objects.

Dose-response work. The response is bounded below at zero and often saturates, so a form is chosen for those properties and not for fit. A polynomial that happened to fit the observed doses better would predict falling response at high doses, an artefact of the leading term rather than a pharmacological claim.

Energy demand against temperature. The relationship turns, because heating and cooling both raise consumption. Fitting a single quadratic places one turning point between the two real ones and misestimates demand at both extremes. This is a case where no form in this unit applies, and the choice is between a higher-degree polynomial selected on held-out error and a nonparametric smoother.

Physical calibration. A sensor's response is fitted by a low-degree polynomial over a stated operating range, and the range is published with the coefficients. The convention exists because the calibration is used by people who did not collect the data, and a fitted polynomial carries no internal record of where it was supported.

---

The common thread. In each case the form was selected from what the response is and what is known about the mechanism, and the coefficient was reported on the scale that matches how the answer will be used: a doubling time, a rate ratio, a calibrated value with its range. Where the form was chosen to fit rather than to match, the failures show up outside the data, at high doses, past the epidemic peak, beyond the calibrated range, which is where a fitted curve is least able to warn anyone.

Next step

Practice Choosing a Regression Form

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.