Practice: Choosing a Regression Form

Classification

Three of these can be fitted by ordinary least squares in one step, because they are linear in their coefficients. Which one cannot?

Direct application

A bacterial count is fitted by least squares on the log scale, giving log ⁡ y = 1.100699 + 0.694637 x with x in hours. By what factor is the count multiplied per hour? Give the factor to four decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

The fitted equation is on the log scale. Undo the logarithm to return to the scale of the count.

Hint 2: Next step

The per-unit factor is e b , where b is the slope.

Error diagnosis

Daily counts of equipment failures are fitted by ordinary least squares against machine age. The fitted line passes sensibly through the data, but a plot of residuals against fitted values shows a widening funnel: residuals near a fitted value of 1 are tiny, and residuals near a fitted value of 15 are large in both directions. What does this indicate, and what follows?

Method selection

A clinic records, for each of 400 patients, the number of appointments missed in a year and the distance they live from the clinic. Most patients miss none or one; a few miss more than ten. Which form should be fitted, and on what grounds?

Interpretation

Eight observations run from x = 1 to x = 8 . A quadratic and a degree-6 polynomial are both fitted, and both pass close to every point; the degree-6 fit has the smaller residual sum of squares. At x = 10 the quadratic predicts 98.23 and the degree-6 predicts 74.58 . What should be concluded?

Transfer · Evaluation

During the first three weeks of an outbreak, daily case counts are fitted on the log scale and the growth factor is reported as 1.23 per day. A briefing extrapolates this to twelve weeks and projects more cases than the region has residents. Which statement identifies the error most precisely?

Construction · Evaluation · Explanation

A utility records, for each of 900 installed meters over one year, the number of repair callouts and the meter's age in years. Ages run from 0 to 14. Most meters had no callouts; a few had more than eight. A colleague has fitted callouts = β 0 + β 1 ⋅ age by ordinary least squares, reports β 1 = 0.41 as "0.41 more callouts per year of age", and projects the callout rate for meters aged 25 years, which the utility is considering keeping in service.

Write an assessment. Address all of the following.

  1. The response. Say what kind of quantity the callout count is and which forms its nature rules out, naming the specific assumption each violated form makes.
  2. The fitted model. State what is wrong with the colleague's fit, including what it predicts for young meters and what its reported standard errors describe.
  3. The form you would use. Name it, say what it models, and say how you would report its coefficient in a sentence an engineer could act on.
  4. A curved relationship. Suppose callouts rise slowly to age 8 and then faster. Say how you would accommodate that, whether the resulting model is still a linear model, and how you would choose any flexibility parameter it introduces.
  5. The projection to age 25. Say what can and cannot be claimed, and why the answer does not depend on which of the forms was fitted.

Write your answer, then compare it with the worked solution.

3 hints available, least help first.

Hint 1: Retrieval cue

Start from what the response can be, not from the shape of the scatter plot.

Hint 2: Concept cue

Two separate defects afflict the colleague's fit: what it can predict, and what its uncertainty describes.

Hint 3: Strategy cue

For the projection, ask what evidence exists between ages 14 and 25, and what is producing the number in its absence.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

1. The response. Callouts are a count: non-negative integers, no upper bound, concentrated near zero with a long right tail. Two consequences. Ordinary least squares on the raw count assumes the response is continuous and unbounded in both directions, and that its variance is the same for every observation. Counts satisfy neither. Least squares on log ⁡ ( callouts ) is unavailable because most meters recorded zero and log ⁡ 0 is undefined. An offset such as log ⁡ ( y + 1 ) would permit the fit, but the estimate then depends on whether 1 or 0.5 was added, with no principle to settle it. 2. The fitted model. The fit will predict negative callouts for young meters: at age 0 the intercept is whatever least squares chose, and with most observations at zero and a positive slope the line sits below zero somewhere in the lower age range. A negative count is not an imprecise prediction but an impossible one. The reported standard errors are computed from a single pooled residual variance. Because a count's variance rises with its mean, that figure is too large where callouts are rare and too small where they are common, so the uncertainty is misstated in both directions rather than uniformly. The sentence "0.41 more callouts per year of age" is a defensible reading of that fit; the problem is the fit, not the interpretation. 3. The form I would use. Poisson regression, modelling log ⁡ E [ callouts ∣ age ] = β 0 + β 1 age . Modelling the log of the mean keeps every fitted mean positive, and the Poisson assumption Var ⁡ ( y ) = E [ y ] matches the variance structure the data show. The coefficient is reported by exponentiating: if β 1 = 0.12 then e 0.12 = 1.127 , so the expected callout rate rises by about 12.7% per year of meter age, or equivalently is multiplied by 1.127. Quoting 0.12 itself would state a log-scale slope and would be read as an additive change it is not. I would also check for overdispersion by comparing residual variance against fitted means. If variance runs well above the mean, the point estimates stand but the intervals are too narrow, and a model allowing extra dispersion is needed. 4. A curved relationship. Add a quadratic term in age, fitting log ⁡ E [ y ] = β 0 + β 1 a + β 2 a 2 . This is still a linear model in the sense that matters: the coefficients enter to the first power and multiply columns, so the estimation is unchanged and only the design matrix grows. Linearity refers to the coefficients, not to the predictors. The degree is a flexibility parameter, so it is chosen by held-out error rather than by residual deviance, which falls with every added column. If the same data must also report performance, the degree search belongs inside a nested scheme so the reported error is not the minimum of the values the search examined. 5. The projection to age 25. The data stop at 14. A prediction at 25 is nearly twice the observed range, and no observation bears on it. Whatever the fitted form produces there is a statement about the form: a Poisson fit with a positive slope keeps multiplying, a quadratic term makes it accelerate or turn depending on the sign of β 2 , and a linear least-squares fit rises steadily. These disagree, and the data cannot adjudicate. This is why the answer does not depend on which form was fitted. Choosing the right form fixes points 1 through 4 and changes nothing about point 5. What can be claimed is the rate relationship over ages 0 to 14, reported with that range stated. Deciding about 25-year-old meters needs either observations of meters that old or an engineering model of the failure mechanism, not an extrapolated curve.

A complete answer does each of these:

  • reads response type
  • fits polynomial terms
  • interprets multiplicative scale
  • checks variance assumption
  • bounds extrapolation
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.