Hypothesis Tests for Experimental Research
What you will be able to do
Given a described study, the learner can identify the unit of analysis and the data structure, paired or independent, one sample or two, and select a test whose standard error matches that design.
What you will be able to do
Given an estimate, a null value and the design's standard error, the learner can compute the standardized distance, name the reference distribution with its degrees of freedom, and state what the result does and does not establish, treating failure to reject as inconclusive rather than as evidence for the null.
Orientation
A large p-value is routinely reported as evidence that there is no effect. It is not, and the gap between those two statements is where most of the damage in applied statistics happens.
Before any of that, though, comes a decision the arithmetic cannot make for you: which test the design supports. Paired or independent, one sample or two, means or proportions. Choose wrongly and every number after it is computed correctly and means nothing. This unit is about reading a study description well enough to make that choice, then saying what the result actually licenses.
Intuition
One design, carried through the arithmetic
The claim that choosing a test is answering two questions is easiest to believe after watching the same data answer them twice.
The data. Twelve participants, each measured before and after an intervention. Differences (after − before):
Treated as paired, which is what the design is. The unit of analysis is the difference.
against
Treated as two independent groups, which it is not. Pooling the twelve before-scores and twelve after-scores as though they came from different people, the relevant spread becomes the between-person variation, which in this data is far larger than the within-person differences. With a pooled
Same twenty-four numbers.
Which of the two questions did the work. Not the reference distribution, since both used a
A closing check worth running on any result. Ask what the denominator was computed across, and whether that matches the unit the design actually randomised or sampled. A p-value is a position within a distribution of hypothetical repetitions, so it is only as meaningful as the account of what would have varied between them.
Principle
The decision comes before the arithmetic
Every test below is the same ratio. What differs is the standard error, and the design chooses it, so the first question is never "which formula" but "what varies here".
Read the study description and answer three questions in order.
1. What is the unit of analysis? If each unit contributes two measurements, the unit is the difference, and everything else follows from that. Thirty people measured twice is thirty differences, not sixty observations.
2. What is being compared, and to what? One group against a fixed value; two groups against each other; a proportion against a stated probability.
3. Was the standard error estimated from the data? If yes, the reference distribution is
| The description says | Unit of analysis | Standard error from |
|---|---|---|
| Each subject measured under both conditions | The within-subject difference | Spread of the differences |
| Two separate groups of subjects | The individual observation | Both groups' spreads, added |
| One group against a claimed value | The individual observation | That group's spread |
| Counts of successes in two groups | The individual trial | A proportion pooled under the null |
What this rules out. Choosing the test by what the numbers look like, or by which formula was most recently taught. The data cannot tell you whether two columns are paired, only the description of how they were collected can, and a table of numbers looks identical either way.
The cost of getting it wrong. Not a rounding error. The worked example below shows the same measurements giving
Definition
What varies across the instances, and what does not
The catalogue above looks like seven formulae. It is one formula and seven standard errors, and reading it that way is the difference between memorising and knowing.
The numerator is always the same question. How far is the estimate from the value the null asserts? Every instance subtracts the null value, which is often zero and therefore invisible, as in the two-sample and paired forms where
The denominator is where the design enters. Everything that distinguishes the instances lives here. Two independent groups add the arms' variances because the observations are independent draws; a paired design does not, because it works on one column of differences and the between-subject variation has already cancelled. That is the entire difference between the two-sample and paired forms, and it is a fact about how the data were collected.
The proportion tests are the instructive case. Both use
The pooled two-sample variant is listed because it appears everywhere, not because it is recommended. Welch costs almost nothing when variances happen to be equal and protects you when they are not.
Example
The same numbers, two designs
Twenty people are measured under condition A and condition B. Mean A is 52.0, mean B is 48.4, so the observed gap is 3.6.
If the design used two independent groups of 20 people each, with
which is unremarkable against a
If the design measured the same 20 people under both conditions, the analysis works on the 20 within-person differences. Suppose
which is decisive.
Same gap, opposite conclusions. Nothing about the estimate changed. What changed is the standard error, because the designs differ in what varies. Between people, scores differ a lot; within a person across two conditions, much less. The paired design removes the between-person variation, and the test must be the one that reflects it.
You cannot pick the test from the numbers alone. The design tells you which standard error is the right one, and getting that wrong changes the answer rather than refining it.
Worked example
Choosing the test from a description
Problem. A clinic wants to know whether a new intake form reduces the time staff spend on registration. Thirty staff each processed registrations with the old form for a week, then with the new form the following week. Mean time per registration fell from 6.8 to 6.1 minutes. The standard deviation of the within-staff changes is 1.8 minutes.
An analyst proposes a two-sample
Goal. Identify the right test, compute it, and say what the result supports.
Relevant principle. The standard error must match the design. Each staff member contributes a pair of measurements, so the unit of analysis is the within-staff change.
Step 1: identify the structure. Each staff member appears in both conditions. These are paired measurements, not independent groups.
Reason: the two numbers for one staff member share everything about that person, their speed, experience and caseload, so they are not independent draws.
Step 2: reject the proposed test. A two-sample test would compute its standard error from the spread across staff, which includes all the between-person variation that cancels within a person.
Reason: it would answer a question about two independent samples of staff, which is not the study that was run.
Step 3: form the differences.
Step 4: compute the statistic.
Reason: a one-sample test on the differences against a null of no change.
Step 5: interpret. Against
Result. A paired
Check. Is the paired design definitely right? Yes on the data structure. But there is a design caveat worth stating: the old form was always used first, so any general speed-up over time, staff getting faster at registrations generally, is confounded with the form. The test is correct for the data; the causal reading is limited by the ordering.
Interpretation. Report the paired analysis with an interval for the saving, note the confounding of form with week, and say that a design randomizing the order across staff would separate them.
Non-example
Tests applied to the wrong structure
A two-sample test on paired data. Comparing all before-measurements with all after-measurements ignores that each pair shares a unit. The standard error comes from the wrong variation, usually overstating it and obscuring a real effect.
A paired test on independent groups. Pairing observations that were not paired by the design, arbitrarily matching the first treated unit with the first control, invents a structure the data do not have.
Pooling because a variance test did not reject. Running a preliminary test of equal variances and pooling when it fails to reject uses the same data twice and distorts the error rate of the main test. Welch's test needs no such decision.
Using
Concluding equivalence from a large p-value. "Not significant" is not "no difference". Showing two things are similar requires an equivalence test with a pre-specified margin, which is a different procedure.
Treating a randomization test as the same thing. The arithmetic may match a
Contrast
What rejecting and not rejecting establish
| Small p-value | Large p-value | |
|---|---|---|
| What it says | The data would be surprising if | The data would not be surprising if |
| Licenses | Rejecting | Nothing about |
| Compatible with | A real effect | A real effect too small for this study to detect |
| Correct wording | "We reject the null" | "We fail to reject the null" |
| Needed to claim no effect | — | An equivalence test with a stated margin |
Why the asymmetry is built in. A test asks how compatible the data are with one hypothesis. Data compatible with "no effect" are also compatible with "a small effect", and with "a large effect the study was too small to see". Compatibility does not single out the null.
The practical consequence. A null result should be reported with what the design could have detected. "We found no significant difference" is nearly uninformative on its own; "no significant difference, and the study could have detected a difference of 5 points with 80% power" tells a reader what was actually learned.
The wording matters. "Accept the null" invites the error. "Fail to reject" is clumsier and describes what happened.
Exercise
1: fully structured. A sample of 25 components has mean lifetime 412 hours,
(a) Compute the statistic for
Check: (a)
2: partly structured. A trial measures pain scores for 40 patients before and after treatment.
(a) Which test applies, and why? (b) A colleague proposes comparing the 40 before-scores with the 40 after-scores as two samples. What is wrong? (c) What causal caveat applies regardless of the test?
Check: (a) a paired
3: unstructured. A team reports: "We compared conversion rates between the two variants. Variant B converted at 4.1% versus 3.8% for A, with p = 0.31. We conclude the variants perform equivalently and recommend choosing on development cost."
Assess the conclusion and say what should be reported instead.
Check: the conclusion converts failure to reject into a positive claim of equivalence, which the test does not support. A p-value of 0.31 says the data are consistent with no difference, and also with B being meaningfully better or worse, if the sample was not large enough to distinguish those. What should be reported: the estimated difference of 0.3 percentage points with a confidence interval, which shows the range of differences the data are compatible with. If the interval spans, say, −0.3 to +0.9 points, and a 0.5-point gain would matter commercially, the study simply did not resolve the question. Claiming equivalence requires an equivalence test against a pre-specified margin of practical indifference, which is a different procedure and must be planned in advance.