ANOVA for Experimental Research
One-way analysis of variance asks whether several group means can be treated as equal, by comparing the variation between groups with the variation within them. The decomposition is exact and the test is a ratio of mean squares. What rejection establishes is narrow, that the means are not all equal, and answering which ones differ is a separate question requiring planned contrasts or multiplicity-adjusted comparisons.
Definition
Let
and they satisfy the exact decomposition
Under
Formal statement
Assumptions and scope
Rejecting the omnibus null establishes that at least one mean differs. It does not identify which, and reading it as evidence about a particular pair is unsupported.
Identifying specific differences requires planned contrasts specified in advance, or post-hoc comparisons with a multiplicity adjustment.
The classical
reference assumes independent observations, normality within groups and equal variances across them. Unequal variances, particularly with unequal group sizes, distort the test. The decomposition
is algebraic and holds regardless of the model; it is the reference distribution that depends on the assumptions.For two groups this is the pooled two-sample test, not an alternative to it:
exactly. It offers no protection that the-test lacks. A significant
says nothing about the size of any difference. Group means with intervals report what the test omits. ANOVA can be written as a regression on indicator variables, which is the route to factorial designs and covariate adjustment.
Worked material
Example
A significant F, and what it leaves open
Four teaching methods are compared, 25 students each,
Group means: A 68.2, B 71.4, C 74.9, D 69.1.
What the test establishes. The four means are not all equal. That is the entire content of
What it does not establish. That C beats A, which is the comparison the eye goes to. The omnibus test never examined that pair; it asked a single question about all four simultaneously.
Why the distinction has teeth. With four groups there are six pairwise comparisons. Testing each at the 5% level gives roughly a 26% chance of at least one spurious finding when all means are equal. The omnibus test protects against that inflation precisely by not making the individual comparisons, and reading a particular difference out of it forfeits the protection while keeping the appearance of it.
What to do instead. If C-versus-A was specified before seeing the data, test it as a planned contrast. If the interest arose from looking at the means, use a post-hoc procedure with a multiplicity adjustment. Either way, report the estimated difference with an interval, since
What is genuinely licensed now. That the four methods do not all produce the same mean score (
Non-example
Conclusions an omnibus test does not support
"The ANOVA was significant, so group C differs from group A." The test examined all groups at once. No pairwise comparison was performed.
"The ANOVA was not significant, so the groups are equivalent." Failing to reject is inconclusive here as everywhere; the study may simply lack resolution.
Picking the largest gap after seeing the data and testing it unadjusted. The comparison was selected because it was largest, so its nominal p-value understates how often such a gap arises by chance.
Using ANOVA to protect a subsequent unadjusted comparison. A significant omnibus result does not license unadjusted pairwise tests afterwards. The protection comes from making the adjustment, not from having run
Applying the classical
Treating ANOVA as different in kind from a
Contrast
What the omnibus test answers
| The omnibus | A specific comparison | |
|---|---|---|
| Question | Are all | Does group |
| Null | ||
| Rejection establishes | At least one mean differs | That particular difference |
| Reports magnitude | No | Yes, with an interval |
| Multiplicity | Handled by asking one question | Must be adjusted, or specified in advance |
Why the confusion is so common. The table of group means sits directly beneath the
What the omnibus test supplies. A single question with a single error rate. Six pairwise tests at 5% carry roughly a 26% chance of at least one false positive when all means are equal; one omnibus test carries 5%.
What it costs. All the specificity. The result is uninformative about direction and magnitude, which is usually what a decision requires.
The legitimate routes to specificity. Planned contrasts, specified before the data are seen, testing the comparisons that motivated the study. Or post-hoc comparisons with a multiplicity adjustment, which pay for the specificity by widening the intervals. Both report estimated differences, which
Common errors
Common misconception
A significant omnibus F test shows which groups differ, so the largest observed gap between group means can be reported as a established difference.