Choosing a Regression Form

What the response variable and the shape of the relationship each rule out, how polynomial terms curve a fit while leaving it linear in its coefficients, why exponential and Poisson models are read on a multiplicative scale, and what a fitted form does beyond the range it was fitted on.

Definition

A regression model states how the expected value of a response depends on predictors. The forms below differ in what they assume about the response and about the shape of that dependence.

Polynomial regression adds powers of a predictor as further columns:

E [ y ∣ x ] = β 0 + β 1 x + β 2 x 2 + ⋯ + β d x d .

Despite the curve it produces, this is still a linear model: linear means linear in the coefficients β j , which enter to the first power. It is fitted by ordinary least squares on the expanded design matrix.

Exponential regression applies when the response is multiplied rather than incremented as a predictor advances. Fitting log ⁡ y by least squares gives

log ⁡ y = a + b x ⟺ y = e a ( e b ) x ,

so e b is the factor by which y is multiplied per unit of x , not an amount added.

Poisson regression applies when the response is a count. It models

log ⁡ E [ y ∣ x ] = β 0 + β 1 x ,

so e β 1 is a rate ratio. The Poisson distribution has Var ⁡ ( y ) = E [ y ] , so the variance is not constant across observations.

Parametric and nonparametric. All three fix a functional form in advance and estimate finitely many coefficients, which is what parametric means. A nonparametric fit lets the data determine the shape, at the cost of more data and no coefficients to report.

Assumptions and scope

  • Ordinary least squares assumes the error variance is constant across observations. Count responses violate this systematically, since their variance rises with their mean.

  • Fitting log ⁡ y by least squares minimises squared error on the log scale, not on the original scale. The result is not the same as fitting an exponential curve directly, and the difference matters when a few large values dominate.

  • log ⁡ y is undefined at y = 0 . A dataset with zero counts cannot be log-transformed without an arbitrary offset, which is one reason Poisson regression models the log of the mean rather than the log of the response.

  • A polynomial of degree d has d − 1 turning points available to it and behaves like x d far from the data. Every polynomial fit diverges outside the observed range, and higher degree diverges faster.

  • Choosing a form by looking at the data and then reporting that form's fit as if it had been specified in advance is the selection bias of the previous unit, at the level of the model family.

Worked material

Example

Five responses, and what each rules out

Each case names a response and a predictor. The question is which forms remain available once the response is taken seriously.

Number of emergency admissions per day, against the day's mean temperature. A count, bounded below at zero, typically small. A straight line predicts negative admissions at some temperatures, and the spread around the fit is visibly wider on busy days than on quiet ones. Poisson regression, with e β read as the multiplicative change in the daily rate per degree.

Bacterial population, against hours since inoculation. Positive, and multiplied rather than incremented while growth is unconstrained. Successive ratios roughly constant is the check. Exponential regression on log ⁡ y , with the coefficient reported as a factor per hour or a doubling time. The form fails once the culture saturates, which is a statement about the biology rather than about the fit.

Braking distance, against speed. Continuous, positive, and curved for a reason that is known in advance: kinetic energy rises with the square of speed. Polynomial regression with an x 2 term, chosen from the mechanism rather than from the scatter plot. That distinction matters. The same curve chosen by trying degrees until one fits well is the selection problem of the previous unit.

Proportion of a cohort still subscribed, against months elapsed. Bounded in [ 0 , 1 ] . Every form in this unit will eventually predict outside that interval, the linear one soonest. This case needs logistic regression, which is not covered here; recognising the bound is what shows the available forms are inadequate.

Household electricity use, against outdoor temperature. Continuous and non-monotone: consumption falls as temperature rises toward a comfortable range, then climbs again as cooling starts. No single-turning-point form fits. A quadratic can represent one turning point and will place it somewhere between the two real ones; the available options are a higher-degree polynomial chosen on held-out error, or a nonparametric fit that lets the data set the shape.

---

In each case the response's nature narrowed the field before any curve was drawn, and in two of them, the proportion and the non-monotone case, it ruled out every form this unit offers. Knowing that a form is unavailable is as useful as knowing which to pick, and it is only visible by asking what the response can be.

Contrast

Pairs that look alike and commit to different things

log ⁡ y = a + b x against y = β 0 + β 1 x .

log-scale fitlinear fit
claim y multiplied by e b per unit y increased by β 1 per unit
coefficient reads asa factor or percentagean amount in the units of y
minimisessquared error in log ⁡ y squared error in y
y = 0 undefinedpermitted

The third row is the one most often overlooked. Least squares on log ⁡ y treats a factor-of-two error on a value of 3 as equal to a factor-of-two error on a value of 300. Fitting the same exponential curve directly on the original scale weights the large values far more heavily and gives different coefficients. Both are defensible; they are not the same fit.

Poisson regression against least squares on log ⁡ ( y + 1 ) .

Both handle a count response and produce a multiplicative reading. The second requires an arbitrary offset to cope with zero counts, and the choice of offset changes the estimate, adding 1 rather than 0.5 gives a different slope with no principle deciding between them. Poisson regression models the log of the mean, which is defined when an observed count is zero, so no offset is required.

A degree-2 fit against a degree-6 fit on the same eight points.

quadraticdegree 6
RSS on the data 0.0272 smaller still
prediction at x = 10 98.23 74.58
direction past the datarisingfalling

The more flexible fit is better by the criterion computed on the training data and disagrees by roughly a quarter two units beyond it, in the opposite direction. Training fit and extrapolation behaviour are not merely different questions; here they rank the two models oppositely.

Choosing a form from the mechanism against choosing it from the plot.

Braking distance takes an x 2 term because kinetic energy goes as the square of speed. A reason available before the data. The same term chosen by fitting degrees 1 through 6 and keeping the best is a selection, and the reported error from that search is optimistic by the amount the previous unit describes. The fitted curve can be identical; what differs is what the error figure means.

In each, two procedures produce comparable-looking output and commit to different claims: about scale, about what is defined at zero, about behaviour outside the data, or about what an error figure estimates. None of the differences is visible in the fitted line alone.

Common errors

Common misconception

That adding x 2 or x 3 to a regression makes it a nonlinear model requiring a different fitting method. Linear regression means linear in the coefficients, not linear in the predictors. A polynomial model enters each power as an additional column of the design matrix and is fitted by the same least-squares solution, with the same closed form and the same standard errors. The curve in the picture comes from the predictor values, not from any change to the estimation. Genuinely nonlinear models are those where a coefficient appears inside a function, such as y = β 0 e β 1 x fitted directly, which has no closed-form least-squares solution.

Common misconception

That a coefficient from a model fitted on the log scale reports how much the response changes per unit of the predictor. A slope of 0.7 in log ⁡ y = a + b x does not mean y rises by 0.7 ; it means y is multiplied by e 0.7 ≈ 2.01 for each unit of x , roughly a doubling. The same applies to a Poisson regression coefficient, where e β is a rate ratio. Reading the log-scale value additively understates large effects and misstates the units, and the error is invisible because the reported number is plausible on its face.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.