Hypothesis Tests for Experimental Research
Every test in this family is one template: an estimate, a null value, and a standard error, assembled into a standardized distance and referred to a distribution that says how unusual such a distance would be if the null were true. Choosing the right test is choosing the right standard error and reference distribution for the design that produced the data, and what a large p-value licenses is far less than it is usually asked to carry.
Definition
Every test in this family computes
Formal statement
Assumptions and scope
Failing to reject is not evidence that the null is true. It is consistent with a real effect the study lacked the resolution to detect, so a null result should be reported with the effect sizes the design could have found.
Welch's test is the safer default for two independent groups. Pooling requires a substantive reason to believe the variances are equal.
Do not pool merely because a preliminary test of equal variances failed to reject. Selecting the main test using the same data distorts the reported error rate.
The paired test applies when the design paired the observations. Applying it to independent groups, or a two-sample test to paired data, uses a standard error the design does not support.
For one proportion the null standard error uses
, not ; the Wald interval uses . The two differ deliberately, and the test's version is the one consistent with computing under . The
reference assumes approximate normality of the sampling distribution of the mean, which the central limit theorem supplies at moderate for most populations but not for very small samples from skewed ones. These tests draw their reference distribution from a sampling model. A randomization test derives it from the assignment mechanism instead, and the two answer different questions even when the arithmetic coincides.
Worked material
Example
The same numbers, two designs
Twenty people are measured under condition A and condition B. Mean A is 52.0, mean B is 48.4, so the observed gap is 3.6.
If the design used two independent groups of 20 people each, with
which is unremarkable against a
If the design measured the same 20 people under both conditions, the analysis works on the 20 within-person differences. Suppose
which is decisive.
Same gap, opposite conclusions. Nothing about the estimate changed. What changed is the standard error, because the designs differ in what varies. Between people, scores differ a lot; within a person across two conditions, much less. The paired design removes the between-person variation, and the test must be the one that reflects it.
You cannot pick the test from the numbers alone. The design tells you which standard error is the right one, and getting that wrong changes the answer rather than refining it.
Non-example
Tests applied to the wrong structure
A two-sample test on paired data. Comparing all before-measurements with all after-measurements ignores that each pair shares a unit. The standard error comes from the wrong variation, usually overstating it and obscuring a real effect.
A paired test on independent groups. Pairing observations that were not paired by the design, arbitrarily matching the first treated unit with the first control, invents a structure the data do not have.
Pooling because a variance test did not reject. Running a preliminary test of equal variances and pooling when it fails to reject uses the same data twice and distorts the error rate of the main test. Welch's test needs no such decision.
Using
Concluding equivalence from a large p-value. "Not significant" is not "no difference". Showing two things are similar requires an equivalence test with a pre-specified margin, which is a different procedure.
Treating a randomization test as the same thing. The arithmetic may match a
Contrast
What rejecting and not rejecting establish
| Small p-value | Large p-value | |
|---|---|---|
| What it says | The data would be surprising if | The data would not be surprising if |
| Licenses | Rejecting | Nothing about |
| Compatible with | A real effect | A real effect too small for this study to detect |
| Correct wording | "We reject the null" | "We fail to reject the null" |
| Needed to claim no effect | — | An equivalence test with a stated margin |
Why the asymmetry is built in. A test asks how compatible the data are with one hypothesis. Data compatible with "no effect" are also compatible with "a small effect", and with "a large effect the study was too small to see". Compatibility does not single out the null.
The practical consequence. A null result should be reported with what the design could have detected. "We found no significant difference" is nearly uninformative on its own; "no significant difference, and the study could have detected a difference of 5 points with 80% power" tells a reader what was actually learned.
The wording matters. "Accept the null" invites the error. "Fail to reject" is clumsier and describes what happened.
Common errors
Common misconception
A p-value above the threshold shows the null hypothesis is true, so a non-significant result establishes that there is no effect or no difference.
Common misconception
Before comparing two independent means, test whether the variances are equal; if that test does not reject, pool the variances and use the pooled t-test, because the data have shown the assumption is satisfied.
Related units
Requires
Connected
- Fisher Randomization Inference (contrasts with)
- ANOVA for Experimental Research (suggested next)