Practice: Trees and Ensembles of Them

Recognition · Error diagnosis

The variance of an average of B tree predictions, each of variance σ 2 and pairwise correlation ρ , is ρ σ 2 + ( 1 − ρ ) σ 2 / B .

With ρ = 0.2 , a learner argues that growing the forest from 100 to 5000 trees will drive the variance towards zero.

Which response identifies the error?

2 hints available, least help first.

Hint 1: Retrieval cue

Which of the two terms contains B , and what happens to the other as B grows?

Hint 2: Concept cue

Evaluate 0.2 + 0.8 / B at B = 100 and at B = 5000 , and compare with 0.2 .

Direct application

Ten observations, x = 1 , … , 10 , with responses

y = 2.1 ,   1.9 ,   2.4 ,   2.2 ,   2.6 ,   7.8 ,   8.1 ,   7.6 ,   8.4 ,   8.0 .

A regression tree considers splitting at x < 5.5 . Compute the total residual sum of squares of the two resulting groups, using each group's own mean. Give your answer to four decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

The split puts the first five observations in one group and the last five in the other.

Hint 2: Next step

Compute each group's mean, then sum the squared deviations within each group and add the two totals.

Direct application · Prediction

A regression tree has one split, at x < 5.5 . Its left leaf holds five training observations with mean 2.240 ; its right leaf holds five with mean 7.980 . What does the tree predict for a new observation at x = 9.3 ?

Enter the value. It is checked against the answer and the precision this task asks for.

1 hint available, least help first.

Hint 1: Retrieval cue

Follow the input through the split, then report what that leaf holds.

Classification

A deep regression tree is fitted, then refitted on a second sample from the same population. The two trees share no split, though their predictions are broadly similar. Which description is correct?

Direct application

An ensemble averages B = 100 trees, each with prediction variance σ 2 and pairwise correlation ρ = 0.2 . Using ρ σ 2 + ( 1 − ρ ) σ 2 / B , what is the variance of the average, as a multiple of σ 2 ? Give your answer to four decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

Substitute directly; the two terms are ρ and ( 1 − ρ ) / B .

Hint 2: Next step

0.2 + 0.8 / 100 .

Interpretation

A tree's root splits on income at 41,500 dollars, scoring 12.40 . The runner-up candidate splits on years of tenure, scoring 12.47 . A report states that income is the most important predictor because it was chosen first. What should be said instead?

Transfer · Evaluation

A firm averages the quarterly forecasts of twelve analysts and finds the average far more accurate than any individual. It expands the panel to sixty analysts, all trained in the same programme and reading the same market reports, and the average barely improves. Which account explains this, and what would help?

Construction · Evaluation · Explanation

A maintenance team records, for each of eight machines, the hours run since the last service and the number of faults in the following week:

hours1020304050607080
faults12126767

Work through the following.

  1. The first split. Evaluate the candidate at hours < 45 by the tree criterion, showing the two group means and the total. State what you would compare it against to know whether it is the split the criterion selects.
  2. Prediction. Give the tree's prediction for a machine at 55 hours and for one at 78 hours, and say what the two answers have in common and why.
  3. Stability. A colleague reports that this tree "shows faults rise sharply after 45 hours". Say what the tree does and does not establish about the location of that threshold.
  4. Whether to average. The team proposes bagging 500 such trees. Say what averaging would and would not improve here, and identify the property of these trees that decides whether the exercise is worth the computation.
  5. A forest. Suppose the data gained nine further predictors, one of which is strongly informative. Say what a random forest would do differently from bagging, what it costs each individual tree, and why the trade can still favour the ensemble.

Write your answer, then compare it with the worked solution.

3 hints available, least help first.

Hint 1: Retrieval cue

For part 1, each group's deviations are measured from that group's own mean, not from the overall mean.

Hint 2: Concept cue

For part 3, ask which thresholds between 40 and 50 the data would distinguish.

Hint 3: Strategy cue

For part 4, ask what bagging removes, and then whether these particular trees have any of it.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

1. The first split. At hours < 45 the left group is { 1 , 2 , 1 , 2 } with mean 1.5 , and the right group is { 6 , 7 , 6 , 7 } with mean 6.5 . Squared deviations are 0.25 × 4 = 1 on each side, so the total residual sum of squares is 2 . To know this is the selected split, every other candidate boundary must be scored the same way, the midpoints between adjacent hour values, so 15, 25, 35, 45, 55, 65 and 75, and the minimum taken. The unsplit total is ∑ ( y i − 4 ) 2 = 58 , so this candidate removes 56 of it. Any candidate that leaves a low and a high group mixed will score far worse; the only competitive ones would be those separating the same two groups. 2. Prediction. Both machines fall in the right leaf, so both receive 6.5 . They have the same prediction because the fitted function is piecewise constant: within a region every input gets that region's training mean, and the 23-hour difference between them carries no information the tree can use. A tree distinguishes inputs only where it has placed a split. 3. Stability. The tree establishes that these eight observations separate cleanly into a low group and a high group, and that the boundary lies somewhere between 40 and 50 hours. It does not establish 45. That value is the midpoint convention applied to the gap between the last low observation and the first high one; there are no observations between them, so every threshold in that interval fits these data identically. Reporting "faults rise sharply after 45 hours" states a precision the data do not contain. The defensible claim is that the rise occurs between 40 and 50 hours. 4. Whether to average. Bagging reduces variance and does not reduce bias. These trees have almost no variance to remove: the two groups are separated by a gap of four faults with no overlap, so nearly every bootstrap resample produces the same split and the same two leaf means. Averaging 500 near-identical trees returns approximately the single tree, at 500 times the computation. The deciding property is how much the trees disagree. Bagging pays when individual trees are unstable, deep trees on noisy data with competing candidate splits, and pays nothing when they already agree. A useful check is to refit on a few resamples and see whether the structure moves before committing to an ensemble. 5. A forest. A random forest restricts the candidate predictors at each split to a random subset, so with ten predictors it might consider three or four at each node. The strongly informative predictor is then unavailable at most splits, and different trees are forced to open on different predictors. The cost is that each individual tree fits worse: denied the best predictor, it makes a weaker split. The trade can still favour the ensemble because the quantity being minimised is the variance of the average, which is ρ σ 2 + ( 1 − ρ ) σ 2 / B . Once B is moderately large the second term is small and the floor is ρ σ 2 . Under plain bagging every tree would open on the same strong predictor, making ρ high and the floor high with it. Lowering ρ moves the floor itself, and that gain routinely exceeds the loss from weaker individual trees. At ρ = 0.2 and B = 100 the variance is 0.208 σ 2 ; at ρ = 0.05 it is 0.060 σ 2 , a far larger improvement than any number of additional trees could deliver at the higher correlation.

A complete answer does each of these:

  • selects split by criterion
  • reads tree prediction
  • attributes tree variance
  • quantifies averaging gain
  • explains decorrelation
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.