ANOVA for Experimental Research

One-way analysis of variance asks whether several group means can be treated as equal, by comparing the variation between groups with the variation within them. The decomposition is exact and the test is a ratio of mean squares. What rejection establishes is narrow, that the means are not all equal, and answering which ones differ is a separate question requiring planned contrasts or multiplicity-adjusted comparisons.

Definition

Let Y i j be observation j in group i , with group mean Y ¯ i , grand mean Y ¯ , group size n i , k groups and N = ∑ i n i . The sums of squares are

S S T = ∑ i ∑ j ( Y i j − Y ¯ ) 2 , S S B = ∑ i n i ( Y ¯ i − Y ¯ ) 2 , S S E = ∑ i ∑ j ( Y i j − Y ¯ i ) 2 ,

and they satisfy the exact decomposition S S T = S S B + S S E . The mean squares divide each component by its degrees of freedom, M S B = S S B k − 1 and M S E = S S E N − k , and the omnibus statistic is

F = M S B M S E .

Under H 0 : μ 1 = μ 2 = ⋯ = μ k and the independent, normal, equal-variance model, F ∼ F k − 1 , N − k . For two groups the equal-variance one-way ANOVA is equivalent to the pooled two-sample test, with F = t 2 .

Formal statement

S S T = S S B + S S E ; M S B = S S B k − 1 ; M S E = S S E N − k ; F = M S B M S E ∼ F k − 1 , N − k under H 0 ; for k = 2 , F = t 2 .

Assumptions and scope

  • Rejecting the omnibus null establishes that at least one mean differs. It does not identify which, and reading it as evidence about a particular pair is unsupported.

  • Identifying specific differences requires planned contrasts specified in advance, or post-hoc comparisons with a multiplicity adjustment.

  • The classical F reference assumes independent observations, normality within groups and equal variances across them. Unequal variances, particularly with unequal group sizes, distort the test.

  • The decomposition S S T = S S B + S S E is algebraic and holds regardless of the model; it is the reference distribution that depends on the assumptions.

  • For two groups this is the pooled two-sample test, not an alternative to it: F = t 2 exactly. It offers no protection that the t -test lacks.

  • A significant F says nothing about the size of any difference. Group means with intervals report what the test omits.

  • ANOVA can be written as a regression on indicator variables, which is the route to factorial designs and covariate adjustment.

Worked material

Example

A significant F, and what it leaves open

Four teaching methods are compared, 25 students each, N = 100 . The ANOVA gives F = 4.62 on 3 and 96 degrees of freedom, p = 0.005 .

Group means: A 68.2, B 71.4, C 74.9, D 69.1.

What the test establishes. The four means are not all equal. That is the entire content of p = 0.005 .

What it does not establish. That C beats A, which is the comparison the eye goes to. The omnibus test never examined that pair; it asked a single question about all four simultaneously.

Why the distinction has teeth. With four groups there are six pairwise comparisons. Testing each at the 5% level gives roughly a 26% chance of at least one spurious finding when all means are equal. The omnibus test protects against that inflation precisely by not making the individual comparisons, and reading a particular difference out of it forfeits the protection while keeping the appearance of it.

What to do instead. If C-versus-A was specified before seeing the data, test it as a planned contrast. If the interest arose from looking at the means, use a post-hoc procedure with a multiplicity adjustment. Either way, report the estimated difference with an interval, since F says nothing about magnitude.

What is genuinely licensed now. That the four methods do not all produce the same mean score ( F 3 , 96 = 4.62 , p = 0.005 ), with pairwise comparisons still to follow.

Non-example

Conclusions an omnibus test does not support

"The ANOVA was significant, so group C differs from group A." The test examined all groups at once. No pairwise comparison was performed.

"The ANOVA was not significant, so the groups are equivalent." Failing to reject is inconclusive here as everywhere; the study may simply lack resolution.

Picking the largest gap after seeing the data and testing it unadjusted. The comparison was selected because it was largest, so its nominal p-value understates how often such a gap arises by chance.

Using ANOVA to protect a subsequent unadjusted comparison. A significant omnibus result does not license unadjusted pairwise tests afterwards. The protection comes from making the adjustment, not from having run F first.

Applying the classical F with badly unequal variances and unequal group sizes. The reference distribution assumes equal variances; when they differ and the groups are unbalanced, the actual error rate departs from the nominal one.

Treating ANOVA as different in kind from a t -test for two groups. For k = 2 it is the pooled two-sample test, with F = t 2 . It offers no additional protection.

Contrast

What the omnibus test answers

The omnibus F testA specific comparison
QuestionAre all k means equal?Does group i differ from group j , and by how much?
Null μ 1 = μ 2 = ⋯ = μ k μ i − μ j = 0
Rejection establishesAt least one mean differsThat particular difference
Reports magnitudeNoYes, with an interval
MultiplicityHandled by asking one questionMust be adjusted, or specified in advance

Why the confusion is so common. The table of group means sits directly beneath the F statistic, and the eye compares them immediately. The arithmetic invites the reading the test does not support.

What the omnibus test supplies. A single question with a single error rate. Six pairwise tests at 5% carry roughly a 26% chance of at least one false positive when all means are equal; one omnibus test carries 5%.

What it costs. All the specificity. The result is uninformative about direction and magnitude, which is usually what a decision requires.

The legitimate routes to specificity. Planned contrasts, specified before the data are seen, testing the comparisons that motivated the study. Or post-hoc comparisons with a multiplicity adjustment, which pay for the specificity by widening the intervals. Both report estimated differences, which F does not.

Common errors

Common misconception

A significant omnibus F test shows which groups differ, so the largest observed gap between group means can be reported as a established difference.

Related units

Requires

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.