ANOVA for Experimental Research

What you will be able to do

Given an ANOVA table or a description of a multi-group comparison, the learner can read the decomposition, compute or check the statistic, and state precisely what rejection establishes.

Orientation

A significant F says the group means are not all equal. It does not say which differ, by how much, or whether it matters.

The decomposition is exact and the statistic is a ratio. The competence is stopping where the test stops: it reports that the means are not all equal, and says nothing about which.

Intuition

Comparing between-group and within-group variation

The canonical text explains the ratio. What it does not explain is why the prior question is worth asking at all, or what answering it leaves undone.

The multiplicity problem is the reason the omnibus test exists. Four groups make six pairwise comparisons. Each at 5% gives roughly a 26% chance of at least one false positive when every mean is identical. One test at one error rate replaces that, which is what the single prior question supplies.

But the answer is deliberately uninformative. Rejecting H 0 says some difference exists. It names no pair, no magnitude and no direction, and a reader who converts a significant F into “method C is best” has read something the test did not say.

And a significant F does not license unadjusted comparisons afterwards. The protection comes from making the adjustment, not from having passed a gate first. Running F , then six unadjusted t -tests, restores the 26% error rate the omnibus test was there to control.

What follows a rejection is planned contrasts, or post-hoc comparisons carrying their own correction, each reported with an interval so the size of the difference is visible alongside its significance.

Definition

The decomposition and the statistic

The canonical statement gives the decomposition and the statistic. Three properties of it decide what an F result may be reported as.

The decomposition is algebra; the distribution is a model. S S T = S S B + S S E holds for any numbers whatever, however they arose. That F follows F k − 1 , N − k under H 0 requires independent observations, normal errors and equal variances across groups. A significant F computed on clustered or heteroskedastic data is a statement about the arithmetic, not about the means.

The degrees of freedom are where the group structure enters. k − 1 counts the free differences among k means; N − k counts the observations left after each group's mean is estimated. Adding groups without adding data moves both, which is why an omnibus test over many small groups has little power even when the means genuinely differ.

F is a ratio of variances, so it says nothing about magnitude or direction. A large F establishes that the group means are further apart than within-group scatter would explain. It does not say which means differ, by how much, or in which order. Those require contrasts or post-hoc comparisons with a multiplicity adjustment, and the omnibus result does not license skipping the adjustment.

Example

A significant F, and what it leaves open

Four teaching methods are compared, 25 students each, N = 100 . The ANOVA gives F = 4.62 on 3 and 96 degrees of freedom, p = 0.005 .

Group means: A 68.2, B 71.4, C 74.9, D 69.1.

What the test establishes. The four means are not all equal. That is the entire content of p = 0.005 .

What it does not establish. That C beats A, which is the comparison the eye goes to. The omnibus test never examined that pair; it asked a single question about all four simultaneously.

Why the distinction has teeth. With four groups there are six pairwise comparisons. Testing each at the 5% level gives roughly a 26% chance of at least one spurious finding when all means are equal. The omnibus test protects against that inflation precisely by not making the individual comparisons, and reading a particular difference out of it forfeits the protection while keeping the appearance of it.

What to do instead. If C-versus-A was specified before seeing the data, test it as a planned contrast. If the interest arose from looking at the means, use a post-hoc procedure with a multiplicity adjustment. Either way, report the estimated difference with an interval, since F says nothing about magnitude.

What is genuinely licensed now. That the four methods do not all produce the same mean score ( F 3 , 96 = 4.62 , p = 0.005 ), with pairwise comparisons still to follow.

Worked example

Completing an ANOVA table

Problem. Three fertiliser formulations are compared on 30 plots, 10 each. A partially filled table:

SourceSSdfMSF
Between246.0???
Within???
Total984.0?

Complete it and state the conclusion.

Goal. The missing entries and a correctly bounded interpretation.

Relevant principle. S S T = S S B + S S E , with degrees of freedom k − 1 , N − k and N − 1 .

Step 1: degrees of freedom. With k = 3 and N = 30 : between = k − 1 = 2 ; within = N − k = 27 ; total = N − 1 = 29 .

Reason: the between-group term has one degree of freedom fewer than the number of groups, and the within-group term loses one per group.

Step 2: the missing sum of squares.

S S E = S S T − S S B = 984.0 − 246.0 = 738.0 .

Reason: the decomposition is exact, so the residual follows by subtraction.

Step 3: mean squares.

M S B = 246.0 2 = 123.0 , M S E = 738.0 27 = 27.33 .

Step 4: the statistic.

F = 123.0 27.33 ≈ 4.50  on  2  and  27  degrees of freedom .

Step 5: the conclusion. The critical value at α = 0.05 for F 2 , 27 is about 3.35, so 4.50 exceeds it; p ≈ 0.02 . The three formulation means are not all equal.

Result. S S E = 738.0 ; df 2, 27, 29; M S B = 123.0 , M S E = 27.33 ; F ≈ 4.50 , p ≈ 0.02 .

Check. Do the degrees of freedom add up? 2 + 27 = 29 , matching the total. And M S E = 27.33 estimates the within-group variance, so the within-group standard deviation is about 27.33 ≈ 5.2 . A plausible scale for plot yields.

Interpretation. Report that the formulations differ, and stop there until contrasts are run. If one formulation was the incumbent and the other two candidates, those two comparisons were the planned questions and should be tested as contrasts, each with an estimated difference and interval. Reporting that one formulation was best, on this table alone, would state more than the test supports.

Non-example

Conclusions an omnibus test does not support

"The ANOVA was significant, so group C differs from group A." The test examined all groups at once. No pairwise comparison was performed.

"The ANOVA was not significant, so the groups are equivalent." Failing to reject is inconclusive here as everywhere; the study may simply lack resolution.

Picking the largest gap after seeing the data and testing it unadjusted. The comparison was selected because it was largest, so its nominal p-value understates how often such a gap arises by chance.

Using ANOVA to protect a subsequent unadjusted comparison. A significant omnibus result does not license unadjusted pairwise tests afterwards. The protection comes from making the adjustment, not from having run F first.

Applying the classical F with badly unequal variances and unequal group sizes. The reference distribution assumes equal variances; when they differ and the groups are unbalanced, the actual error rate departs from the nominal one.

Treating ANOVA as different in kind from a t -test for two groups. For k = 2 it is the pooled two-sample test, with F = t 2 . It offers no additional protection.

Contrast

What the omnibus test answers

The omnibus F testA specific comparison
QuestionAre all k means equal?Does group i differ from group j , and by how much?
Null μ 1 = μ 2 = ⋯ = μ k μ i − μ j = 0
Rejection establishesAt least one mean differsThat particular difference
Reports magnitudeNoYes, with an interval
MultiplicityHandled by asking one questionMust be adjusted, or specified in advance

Why the confusion is so common. The table of group means sits directly beneath the F statistic, and the eye compares them immediately. The arithmetic invites the reading the test does not support.

What the omnibus test supplies. A single question with a single error rate. Six pairwise tests at 5% carry roughly a 26% chance of at least one false positive when all means are equal; one omnibus test carries 5%.

What it costs. All the specificity. The result is uninformative about direction and magnitude, which is usually what a decision requires.

The legitimate routes to specificity. Planned contrasts, specified before the data are seen, testing the comparisons that motivated the study. Or post-hoc comparisons with a multiplicity adjustment, which pay for the specificity by widening the intervals. Both report estimated differences, which F does not.

Exercise

1: fully structured. An ANOVA on 5 groups with 12 observations each reports S S B = 180 , S S E = 825 .

(a) Give the degrees of freedom. (b) Compute M S B , M S E and F . (c) State the conclusion at α = 0.05 , given a critical value of about 2.54.

Check: (a) between = 4 , within = 60 − 5 = 55 , total = 59 ; (b) M S B = 180 / 4 = 45 , M S E = 825 / 55 = 15 , F = 45 / 15 = 3.0 ; (c) 3.0 exceeds 2.54, so reject. The five means are not all equal, and nothing yet is established about any particular pair.

2: partly structured. A two-group comparison gives t = 2.40 from a pooled two-sample test.

(a) What would the corresponding one-way ANOVA report? (b) Does running ANOVA instead offer any advantage here? (c) When does ANOVA become genuinely useful?

Check: (a) F = t 2 = 5.76 , with the same p-value; (b) none. They are the same test, and the t -test additionally reports the direction of the difference, which F discards; (c) with three or more groups, where it answers one question in place of many and controls the error rate accordingly.

3: unstructured. A product team reports: "We tested four onboarding flows with 400 users each. ANOVA gave F = 5.1 , p = 0.002 . Flow C had the highest completion rate at 62% versus 55% for flow A, so we are rolling out flow C."

Assess the reasoning and say what should be done before the rollout decision.

Check: the omnibus test establishes only that the four flows do not all perform alike; it did not test C against A, and C was singled out because it looked best, which is precisely the selection that makes an unadjusted comparison misleading. Before deciding: run the C-versus-A comparison explicitly with a multiplicity adjustment across the six possible pairs, and report the estimated difference with an interval rather than two raw percentages. A 7-point gap with an interval spanning 1 to 13 points supports a different decision from one spanning −1 to 15. Also worth checking whether any comparison was planned in advance, since a pre-specified contrast needs no post-hoc penalty, and whether a 7-point difference is large enough to justify the switching cost, which the test cannot address at all.

What to carry forward

The decomposition. S S T = S S B + S S E , exactly, with S S B = ∑ i n i ( Y ¯ i − Y ¯ ) 2 and S S E = ∑ i ∑ j ( Y i j − Y ¯ i ) 2 .

Mean squares. M S B = S S B / ( k − 1 ) and M S E = S S E / ( N − k ) . Each sum of squares on its own degrees of freedom.

The statistic. F = M S B / M S E , distributed F k − 1 , N − k under the null and the classical model.

What rejection establishes. That the means are not all equal. Not which ones, not by how much, not in which direction.

Getting specificity. Planned contrasts specified in advance, or post-hoc comparisons with a multiplicity adjustment. Both report estimated differences; F does not.

Assumptions. Independence, within-group normality, equal variances. Unequal variances with unequal group sizes distort the reference.

Two groups. This is the pooled two-sample test: F = t 2 , no extra protection.

Route onward. ANOVA is a regression on indicator variables, which is how it extends to factorial designs and covariate adjustment.

The recurring error. Reading a significant omnibus result as evidence about a particular pair of groups.

Next step

Practice ANOVA for Experimental Research

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.