ANOVA for Experimental Research
What you will be able to do
Given an ANOVA table or a description of a multi-group comparison, the learner can read the decomposition, compute or check the statistic, and state precisely what rejection establishes.
Orientation
A significant
The decomposition is exact and the statistic is a ratio. The competence is stopping where the test stops: it reports that the means are not all equal, and says nothing about which.
Intuition
Comparing between-group and within-group variation
The canonical text explains the ratio. What it does not explain is why the prior question is worth asking at all, or what answering it leaves undone.
The multiplicity problem is the reason the omnibus test exists. Four groups make six pairwise comparisons. Each at 5% gives roughly a 26% chance of at least one false positive when every mean is identical. One test at one error rate replaces that, which is what the single prior question supplies.
But the answer is deliberately uninformative. Rejecting
And a significant
What follows a rejection is planned contrasts, or post-hoc comparisons carrying their own correction, each reported with an interval so the size of the difference is visible alongside its significance.
Definition
The decomposition and the statistic
The canonical statement gives the decomposition and the statistic. Three properties of it decide what an
The decomposition is algebra; the distribution is a model.
The degrees of freedom are where the group structure enters.
Example
A significant F, and what it leaves open
Four teaching methods are compared, 25 students each,
Group means: A 68.2, B 71.4, C 74.9, D 69.1.
What the test establishes. The four means are not all equal. That is the entire content of
What it does not establish. That C beats A, which is the comparison the eye goes to. The omnibus test never examined that pair; it asked a single question about all four simultaneously.
Why the distinction has teeth. With four groups there are six pairwise comparisons. Testing each at the 5% level gives roughly a 26% chance of at least one spurious finding when all means are equal. The omnibus test protects against that inflation precisely by not making the individual comparisons, and reading a particular difference out of it forfeits the protection while keeping the appearance of it.
What to do instead. If C-versus-A was specified before seeing the data, test it as a planned contrast. If the interest arose from looking at the means, use a post-hoc procedure with a multiplicity adjustment. Either way, report the estimated difference with an interval, since
What is genuinely licensed now. That the four methods do not all produce the same mean score (
Worked example
Completing an ANOVA table
Problem. Three fertiliser formulations are compared on 30 plots, 10 each. A partially filled table:
| Source | SS | df | MS | F |
|---|---|---|---|---|
| Between | 246.0 | ? | ? | ? |
| Within | ? | ? | ? | |
| Total | 984.0 | ? |
Complete it and state the conclusion.
Goal. The missing entries and a correctly bounded interpretation.
Relevant principle.
Step 1: degrees of freedom. With
Reason: the between-group term has one degree of freedom fewer than the number of groups, and the within-group term loses one per group.
Step 2: the missing sum of squares.
Reason: the decomposition is exact, so the residual follows by subtraction.
Step 3: mean squares.
Step 4: the statistic.
Step 5: the conclusion. The critical value at
Result.
Check. Do the degrees of freedom add up?
Interpretation. Report that the formulations differ, and stop there until contrasts are run. If one formulation was the incumbent and the other two candidates, those two comparisons were the planned questions and should be tested as contrasts, each with an estimated difference and interval. Reporting that one formulation was best, on this table alone, would state more than the test supports.
Non-example
Conclusions an omnibus test does not support
"The ANOVA was significant, so group C differs from group A." The test examined all groups at once. No pairwise comparison was performed.
"The ANOVA was not significant, so the groups are equivalent." Failing to reject is inconclusive here as everywhere; the study may simply lack resolution.
Picking the largest gap after seeing the data and testing it unadjusted. The comparison was selected because it was largest, so its nominal p-value understates how often such a gap arises by chance.
Using ANOVA to protect a subsequent unadjusted comparison. A significant omnibus result does not license unadjusted pairwise tests afterwards. The protection comes from making the adjustment, not from having run
Applying the classical
Treating ANOVA as different in kind from a
Contrast
What the omnibus test answers
| The omnibus | A specific comparison | |
|---|---|---|
| Question | Are all | Does group |
| Null | ||
| Rejection establishes | At least one mean differs | That particular difference |
| Reports magnitude | No | Yes, with an interval |
| Multiplicity | Handled by asking one question | Must be adjusted, or specified in advance |
Why the confusion is so common. The table of group means sits directly beneath the
What the omnibus test supplies. A single question with a single error rate. Six pairwise tests at 5% carry roughly a 26% chance of at least one false positive when all means are equal; one omnibus test carries 5%.
What it costs. All the specificity. The result is uninformative about direction and magnitude, which is usually what a decision requires.
The legitimate routes to specificity. Planned contrasts, specified before the data are seen, testing the comparisons that motivated the study. Or post-hoc comparisons with a multiplicity adjustment, which pay for the specificity by widening the intervals. Both report estimated differences, which
Exercise
1: fully structured. An ANOVA on 5 groups with 12 observations each reports
(a) Give the degrees of freedom. (b) Compute
Check: (a) between
2: partly structured. A two-group comparison gives
(a) What would the corresponding one-way ANOVA report? (b) Does running ANOVA instead offer any advantage here? (c) When does ANOVA become genuinely useful?
Check: (a)
3: unstructured. A product team reports: "We tested four onboarding flows with 400 users each. ANOVA gave
Assess the reasoning and say what should be done before the rollout decision.
Check: the omnibus test establishes only that the four flows do not all perform alike; it did not test C against A, and C was singled out because it looked best, which is precisely the selection that makes an unadjusted comparison misleading. Before deciding: run the C-versus-A comparison explicitly with a multiplicity adjustment across the six possible pairs, and report the estimated difference with an interval rather than two raw percentages. A 7-point gap with an interval spanning 1 to 13 points supports a different decision from one spanning −1 to 15. Also worth checking whether any comparison was planned in advance, since a pre-specified contrast needs no post-hoc penalty, and whether a 7-point difference is large enough to justify the switching cost, which the test cannot address at all.
What to carry forward
The decomposition.
Mean squares.
The statistic.
What rejection establishes. That the means are not all equal. Not which ones, not by how much, not in which direction.
Getting specificity. Planned contrasts specified in advance, or post-hoc comparisons with a multiplicity adjustment. Both report estimated differences;
Assumptions. Independence, within-group normality, equal variances. Unequal variances with unequal group sizes distort the reference.
Two groups. This is the pooled two-sample test:
Route onward. ANOVA is a regression on indicator variables, which is how it extends to factorial designs and covariate adjustment.
The recurring error. Reading a significant omnibus result as evidence about a particular pair of groups.