Practice: ANOVA for Experimental Research

Recognition · Interpretation

A one-way ANOVA on four groups gives F = 4.62 , p = 0.005 . What does this establish?

2 hints available, least help first.

Hint 1: Retrieval cue

Write the omnibus null, then state its negation.

Hint 2: Concept cue

What is the logical opposite of 'all four means are equal'?

Direct application · Interpretation · Explanation

Three formulations are compared on 30 plots, 10 each. S S B = 246.0 and S S T = 984.0 .

Complete the table, degrees of freedom, S S E , mean squares, F , and state the conclusion at α = 0.05 , given a critical value of about 3.35.

Finally: had there been two formulations rather than three, state what this procedure reduces to and what relationship the statistics would satisfy.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Use S S T = S S B + S S E to recover the missing sum of squares.

Hint 2: Strategy cue

Fill the degrees of freedom first; every mean square follows from them.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

Degrees of freedom. Between = k − 1 = 2 ; within = N − k = 30 − 3 = 27 ; total = N − 1 = 29 .

Residual sum of squares. S S E = S S T − S S B = 984.0 − 246.0 = 738.0 , by the exact decomposition.

Mean squares. M S B = 246.0 / 2 = 123.0 ; M S E = 738.0 / 27 = 27.33 .

Statistic. F = 123.0 / 27.33 ≈ 4.50 on 2 and 27 degrees of freedom.

Conclusion. 4.50 exceeds the critical value of 3.35, so the null of equal means is rejected; p ≈ 0.02 . The three formulation means are not all equal.

Checks. The degrees of freedom add: 2 + 27 = 29 . And M S E = 27.33 implies a within-group standard deviation of about 5.2, a plausible scale for plot yields.

What this does not license. Any statement about which formulation is best. That requires planned contrasts, or post-hoc comparisons with a multiplicity adjustment, each reporting an estimated difference with an interval, none of which F supplies.

With two groups. The equal-variance one-way ANOVA is exactly the pooled two-sample t test, and the statistics satisfy F = t 2 . The F reference distribution with 1 and N − 2 degrees of freedom is the square of the t with N − 2 . Nothing is gained by running ANOVA on two groups: it is the same test written differently, and it offers no extra protection because there is only one comparison to make.

A complete answer does each of these:

  • reads the decomposition
  • confines the conclusion
  • names the assumptions
  • relates to two sample

Comparison · Interpretation

A two-group comparison gives t = 2.40 from a pooled two-sample test. What would a one-way ANOVA on the same data report?

2 hints available, least help first.

Hint 1: Retrieval cue

Recall the relationship between F and t when k = 2 .

Hint 2: Concept cue

Square the t statistic and see what you get.

Direct application · Completion

An ANOVA on k = 3 groups with N = 30 observations reports S S B = 120 and S S E = 270 . What are the mean squares and F ?

1 hint available, least help first.

Hint 1: Retrieval cue

d f B = k − 1 and d f E = N − k .

Error diagnosis · Explanation · Evaluation

A product team reports:

We tested four onboarding flows with 400 users each. ANOVA gave F = 5.1 , p = 0.002 . Flow C had the highest completion rate at 62% versus 55% for flow A, so we are rolling out flow C.

Identify the error and say what should be done before the rollout decision.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask which hypothesis the F test actually examined.

Hint 2: Concept cue

Why is the largest observed difference the most misleading one to test unadjusted?

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

The error. The omnibus test established only that the four flows do not all perform alike. It never tested C against A. The decision rests on a comparison the analysis did not make.

Why the selection matters. Flow C was singled out for having the highest observed rate. With four groups there are six pairwise comparisons, and the largest observed gap is the one most inflated by chance. Testing it as though it had been chosen in advance understates how often such a gap arises when the flows are equivalent.

What a significant F does not provide. Protection for subsequent unadjusted comparisons. The protection comes from making the adjustment, not from having run the omnibus test first.

What should be done. Test C against A explicitly, with a multiplicity adjustment across the six possible pairs. Report the estimated difference with an interval rather than two raw percentages: a 7-point gap with an interval of 1 to 13 points supports a different decision from one spanning − 1 to 15.

Two further questions. Was any comparison planned in advance? A pre-specified contrast needs no post-hoc penalty, and if C-versus-A was the motivating question, that changes the analysis. And is 7 points large enough to justify the switching cost? The test cannot address that at all, which is another reason the interval matters more than the p-value.

Where the protection actually comes from. The omnibus test guards against the multiplicity created by having four flows to compare. For two flows there is no such multiplicity and ANOVA reduces to the pooled two-sample test with F = t 2 , so the protection is a property of how many comparisons are in play, not of choosing ANOVA. Picking the best of four after a significant omnibus result reintroduces exactly the multiplicity the test was guarding.

A complete answer does each of these:

  • reads the decomposition
  • confines the conclusion
  • names the assumptions
  • relates to two sample

Transfer · Evaluation · Explanation

A health system compares readmission rates across six hospitals after a new discharge protocol. It reports: "One-way ANOVA across the six sites gives F = 3.4 , p = 0.005 , confirming that the protocol's effectiveness varies by site. Site 4 had the lowest readmission rate, so we will study its practices as the model for the others."

Assess both claims and say what you would report.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask what comparison would be needed to support a claim about the protocol's effectiveness.

Hint 2: Strategy cue

Consider why the most extreme of six observed values is a risky basis for a decision.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

First claim — variation by site. The F test shows the six site means are not all equal. It does not show that the protocol's effectiveness varies by site: with no comparison group, site differences in readmission reflect case mix, catchment population, staffing and baseline quality as much as anything about the protocol. Establishing that effectiveness varies would require a treatment effect estimated per site, which needs pre-protocol data or unexposed comparison units.

Second claim — site 4 as the model. Site 4 was selected for having the most extreme observed value among six, which is the comparison most inflated by chance and also the one most exposed to regression to the mean: a site that looks best this period is likely to look less exceptional next period even if nothing about it is special. The omnibus test did not test site 4 against any other site.

Assumption concerns specific to this setting. Hospitals differ greatly in size, so the groups are unbalanced and within-site variances almost certainly differ, exactly the combination under which the classical F reference is least reliable. Readmissions within a hospital are also unlikely to be independent in the way the model assumes.

What I would report. Site-level readmission rates with intervals, so the reader sees which differences are resolved and which are not; a comparison against each site's own pre-protocol baseline rather than against other sites, which is the contrast that bears on the protocol; adjustment for case mix; and, if site 4 is to be investigated, a statement that it is a hypothesis-generating choice rather than a demonstrated exemplar.

A note on scope. Had the system compared two hospitals rather than six, this analysis would be the pooled two-sample test with F = t 2 , and the same objection would stand: a difference between sites is not an effect of the protocol. The number of groups changes the multiplicity, not what the comparison is capable of establishing.

A complete answer does each of these:

  • reads the decomposition
  • confines the conclusion
  • names the assumptions
  • relates to two sample
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.