Module 2 of 3 · Lesson 1 of 3

Hypothesis Tests for Experimental Research

The common structure of the named tests, and what a large p-value does not establish.

What you will be able to do

Given a described study, the learner can identify the unit of analysis and the data structure, paired or independent, one sample or two, and select a test whose standard error matches that design.

What you will be able to do

Given an estimate, a null value and the design's standard error, the learner can compute the standardized distance, name the reference distribution with its degrees of freedom, and state what the result does and does not establish, treating failure to reject as inconclusive rather than as evidence for the null.

Orientation

A large p-value is routinely reported as evidence that there is no effect. It is not, and the gap between those two statements is where most of the damage in applied statistics happens.

Before any of that, though, comes a decision the arithmetic cannot make for you: which test the design supports. Paired or independent, one sample or two, means or proportions. Choose wrongly and every number after it is computed correctly and means nothing. This unit is about reading a study description well enough to make that choice, then saying what the result actually licenses.

Intuition

One design, carried through the arithmetic

The claim that choosing a test is answering two questions is easiest to believe after watching the same data answer them twice.

The data. Twelve participants, each measured before and after an intervention. Differences (after − before): 3 , − 1 , 4 , 2 , 0 , 5 , 1 , 3 , − 2 , 4 , 2 , 1 .

Treated as paired, which is what the design is. The unit of analysis is the difference. d ¯ = 1.833 , s d = 2.125 , so

S E = 2.125 12 = 0.613 , t = 1.833 0.613 = 2.99 ,

against t 11 , giving p = 0.012 .

Treated as two independent groups, which it is not. Pooling the twelve before-scores and twelve after-scores as though they came from different people, the relevant spread becomes the between-person variation, which in this data is far larger than the within-person differences. With a pooled S E of about 1.94 the statistic falls to t = 0.95 on 22 degrees of freedom, p = 0.35 .

Same twenty-four numbers. p = 0.012 against p = 0.35 . Nothing about the intervention changed. The second analysis discarded the pairing, and with it the fact that each participant serves as their own control, so all the between-person variation, which the pairing had removed, flooded back into the denominator.

Which of the two questions did the work. Not the reference distribution, since both used a t . The standard error did: what varies, across what, and how many independent pieces of information there are. Getting that wrong is not a refinement, it is a different analysis of different quantities, and it is why "which test" is settled by the design rather than by the outcome variable's type.

A closing check worth running on any result. Ask what the denominator was computed across, and whether that matches the unit the design actually randomised or sampled. A p-value is a position within a distribution of hypothetical repetitions, so it is only as meaningful as the account of what would have varied between them.

Principle

The decision comes before the arithmetic

Every test below is the same ratio. What differs is the standard error, and the design chooses it, so the first question is never "which formula" but "what varies here".

Read the study description and answer three questions in order.

1. What is the unit of analysis? If each unit contributes two measurements, the unit is the difference, and everything else follows from that. Thirty people measured twice is thirty differences, not sixty observations.

2. What is being compared, and to what? One group against a fixed value; two groups against each other; a proportion against a stated probability.

3. Was the standard error estimated from the data? If yes, the reference distribution is t rather than normal, and the degrees of freedom come from how much estimation was done.

The description saysUnit of analysisStandard error from
Each subject measured under both conditionsThe within-subject differenceSpread of the differences
Two separate groups of subjectsThe individual observationBoth groups' spreads, added
One group against a claimed valueThe individual observationThat group's spread
Counts of successes in two groupsThe individual trialA proportion pooled under the null

What this rules out. Choosing the test by what the numbers look like, or by which formula was most recently taught. The data cannot tell you whether two columns are paired, only the description of how they were collected can, and a table of numbers looks identical either way.

The cost of getting it wrong. Not a rounding error. The worked example below shows the same measurements giving T ≈ 1.30 under one reading of the design and T ≈ 3.83 under the other: unremarkable against decisive, from the same data.

Definition

What varies across the instances, and what does not

The catalogue above looks like seven formulae. It is one formula and seven standard errors, and reading it that way is the difference between memorising and knowing.

The numerator is always the same question. How far is the estimate from the value the null asserts? Every instance subtracts the null value, which is often zero and therefore invisible, as in the two-sample and paired forms where Δ 0 = 0 and μ D , 0 = 0 are usually omitted. When a test compares against a non-zero null the term reappears, and forgetting it is a common arithmetic error rather than a conceptual one.

The denominator is where the design enters. Everything that distinguishes the instances lives here. Two independent groups add the arms' variances because the observations are independent draws; a paired design does not, because it works on one column of differences and the between-subject variation has already cancelled. That is the entire difference between the two-sample and paired forms, and it is a fact about how the data were collected.

The proportion tests are the instructive case. Both use p 0 or a pooled p ^ in the standard error rather than the observed p ^ , because the reference distribution is derived under the null. The standard error must describe how the estimate varies in the world the null describes, not in the world the data suggest. The Wald interval uses p ^ for the opposite reason: it is not conditioning on any null.

t rather than z marks an estimated denominator. Where σ is known the statistic is normal; where s replaces it the extra uncertainty from estimating the spread gives the heavier tails of t , and the degrees of freedom count how much estimation was done.

The pooled two-sample variant is listed because it appears everywhere, not because it is recommended. Welch costs almost nothing when variances happen to be equal and protects you when they are not.

Example

The same numbers, two designs

Twenty people are measured under condition A and condition B. Mean A is 52.0, mean B is 48.4, so the observed gap is 3.6.

If the design used two independent groups of 20 people each, with s 1 = 9.0 and s 2 = 8.5 :

S E = 81 20 + 72.25 20 = 4.05 + 3.61 = 7.66 ≈ 2.77 ,
T = 3.6 2.77 ≈ 1.30 ,

which is unremarkable against a t distribution with roughly 38 degrees of freedom.

If the design measured the same 20 people under both conditions, the analysis works on the 20 within-person differences. Suppose s D = 4.2 :

S E = 4.2 20 ≈ 0.94 , T = 3.6 0.94 ≈ 3.83  on  19  degrees of freedom ,

which is decisive.

Same gap, opposite conclusions. Nothing about the estimate changed. What changed is the standard error, because the designs differ in what varies. Between people, scores differ a lot; within a person across two conditions, much less. The paired design removes the between-person variation, and the test must be the one that reflects it.

You cannot pick the test from the numbers alone. The design tells you which standard error is the right one, and getting that wrong changes the answer rather than refining it.

Worked example

Choosing the test from a description

Problem. A clinic wants to know whether a new intake form reduces the time staff spend on registration. Thirty staff each processed registrations with the old form for a week, then with the new form the following week. Mean time per registration fell from 6.8 to 6.1 minutes. The standard deviation of the within-staff changes is 1.8 minutes.

An analyst proposes a two-sample t -test comparing the 30 old-form averages with the 30 new-form averages.

Goal. Identify the right test, compute it, and say what the result supports.

Relevant principle. The standard error must match the design. Each staff member contributes a pair of measurements, so the unit of analysis is the within-staff change.

Step 1: identify the structure. Each staff member appears in both conditions. These are paired measurements, not independent groups.

Reason: the two numbers for one staff member share everything about that person, their speed, experience and caseload, so they are not independent draws.

Step 2: reject the proposed test. A two-sample test would compute its standard error from the spread across staff, which includes all the between-person variation that cancels within a person.

Reason: it would answer a question about two independent samples of staff, which is not the study that was run.

Step 3: form the differences. D i = old i − new i for each staff member, with D ¯ = 6.8 − 6.1 = 0.7 minutes and s D = 1.8 .

Step 4: compute the statistic.

S E ( D ¯ ) = 1.8 30 = 1.8 5.48 ≈ 0.329 ,
T = 0.7 − 0 0.329 ≈ 2.13  on  d f = 29.

Reason: a one-sample test on the differences against a null of no change.

Step 5: interpret. Against t 29 , a statistic of 2.13 gives a two-sided p-value near 0.04. The data are somewhat surprising under the null of no change.

Result. A paired t -test: T ≈ 2.13 , d f = 29 , p ≈ 0.04 , estimated saving 0.7 minutes per registration.

Check. Is the paired design definitely right? Yes on the data structure. But there is a design caveat worth stating: the old form was always used first, so any general speed-up over time, staff getting faster at registrations generally, is confounded with the form. The test is correct for the data; the causal reading is limited by the ordering.

Interpretation. Report the paired analysis with an interval for the saving, note the confounding of form with week, and say that a design randomizing the order across staff would separate them.

Non-example

Tests applied to the wrong structure

A two-sample test on paired data. Comparing all before-measurements with all after-measurements ignores that each pair shares a unit. The standard error comes from the wrong variation, usually overstating it and obscuring a real effect.

A paired test on independent groups. Pairing observations that were not paired by the design, arbitrarily matching the first treated unit with the first control, invents a structure the data do not have.

Pooling because a variance test did not reject. Running a preliminary test of equal variances and pooling when it fails to reject uses the same data twice and distorts the error rate of the main test. Welch's test needs no such decision.

Using p ^ in the null standard error. For H 0 : p = p 0 , the standard error should be computed under the null, using p 0 . Substituting p ^ answers a slightly different question and is the interval's formula, not the test's.

Concluding equivalence from a large p-value. "Not significant" is not "no difference". Showing two things are similar requires an equivalence test with a pre-specified margin, which is a different procedure.

Treating a randomization test as the same thing. The arithmetic may match a t -test, but the reference distribution comes from the assignment mechanism rather than a sampling model, and the null is a per-unit statement rather than one about means.

Contrast

What rejecting and not rejecting establish

Small p-valueLarge p-value
What it saysThe data would be surprising if H 0 were trueThe data would not be surprising if H 0 were true
LicensesRejecting H 0 at the stated levelNothing about H 0 being true
Compatible withA real effectA real effect too small for this study to detect
Correct wording"We reject the null""We fail to reject the null"
Needed to claim no effect—An equivalence test with a stated margin

Why the asymmetry is built in. A test asks how compatible the data are with one hypothesis. Data compatible with "no effect" are also compatible with "a small effect", and with "a large effect the study was too small to see". Compatibility does not single out the null.

The practical consequence. A null result should be reported with what the design could have detected. "We found no significant difference" is nearly uninformative on its own; "no significant difference, and the study could have detected a difference of 5 points with 80% power" tells a reader what was actually learned.

The wording matters. "Accept the null" invites the error. "Fail to reject" is clumsier and describes what happened.

Exercise

1: fully structured. A sample of 25 components has mean lifetime 412 hours, s = 40 . The manufacturer claims 400 hours.

(a) Compute the statistic for H 0 : μ = 400 . (b) Name the reference distribution and degrees of freedom. (c) At α = 0.05 two-sided, what is the conclusion?

Check: (a) S E = 40 / 25 = 8 , so T = ( 412 − 400 ) / 8 = 1.5 ; (b) t 24 ; (c) the critical value is about 2.06, so 1.5 does not reach it, fail to reject, and the data are consistent with the claim without establishing it.

2: partly structured. A trial measures pain scores for 40 patients before and after treatment.

(a) Which test applies, and why? (b) A colleague proposes comparing the 40 before-scores with the 40 after-scores as two samples. What is wrong? (c) What causal caveat applies regardless of the test?

Check: (a) a paired t -test on the within-patient differences, since each patient contributes both measurements; (b) the two sets are not independent, knowing a patient's before-score tells you much about their after-score, and a two-sample standard error includes between-patient variation that cancels within a patient; (c) there is no control group, so any change could reflect natural recovery, regression to the mean, or placebo effects rather than the treatment.

3: unstructured. A team reports: "We compared conversion rates between the two variants. Variant B converted at 4.1% versus 3.8% for A, with p = 0.31. We conclude the variants perform equivalently and recommend choosing on development cost."

Assess the conclusion and say what should be reported instead.

Check: the conclusion converts failure to reject into a positive claim of equivalence, which the test does not support. A p-value of 0.31 says the data are consistent with no difference, and also with B being meaningfully better or worse, if the sample was not large enough to distinguish those. What should be reported: the estimated difference of 0.3 percentage points with a confidence interval, which shows the range of differences the data are compatible with. If the interval spans, say, −0.3 to +0.9 points, and a 0.5-point gain would matter commercially, the study simply did not resolve the question. Claiming equivalence requires an equivalence test against a pre-specified margin of practical indifference, which is a different procedure and must be planned in advance.

Next step

Practice Hypothesis Tests for Experimental Research

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Confidence Intervals for Experimental Research

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.