Hypothesis Tests for Experimental Research

Every test in this family is one template: an estimate, a null value, and a standard error, assembled into a standardized distance and referred to a distribution that says how unusual such a distance would be if the null were true. Choosing the right test is choosing the right standard error and reference distribution for the design that produced the data, and what a large p-value licenses is far less than it is usually asked to carry.

Definition

Every test in this family computes statistic = estimate − null value standard error and refers it to a distribution derived under H 0 . One mean: Z = X ¯ − μ 0 σ / n with σ known, or T = X ¯ − μ 0 s / n ∼ t n − 1 under a normal population. One proportion: Z = p ^ − p 0 p 0 ( 1 − p 0 ) / n , the null standard error using p 0 . Two independent means (Welch): T = ( X ¯ 1 − X ¯ 2 ) − Δ 0 s 1 2 / n 1 + s 2 2 / n 2 , degrees of freedom by Welch–Satterthwaite. Pooled variant: s p 2 = ( n 1 − 1 ) s 1 2 + ( n 2 − 1 ) s 2 2 n 1 + n 2 − 2 with d f = n 1 + n 2 − 2 . Paired: D i = X i − Y i , then T = D ¯ − μ D , 0 s D / n ∼ t n − 1 . Two proportions: with p ^ = x 1 + x 2 n 1 + n 2 , Z = p ^ 1 − p ^ 2 p ^ ( 1 − p ^ ) ( 1 / n 1 + 1 / n 2 ) .

Formal statement

statistic = estimate − null value S E ; T = X ¯ − μ 0 s / n ∼ t n − 1 ; Welch T = ( X ¯ 1 − X ¯ 2 ) − Δ 0 s 1 2 / n 1 + s 2 2 / n 2 ; paired T = D ¯ − μ D , 0 s D / n .

Assumptions and scope

  • Failing to reject is not evidence that the null is true. It is consistent with a real effect the study lacked the resolution to detect, so a null result should be reported with the effect sizes the design could have found.

  • Welch's test is the safer default for two independent groups. Pooling requires a substantive reason to believe the variances are equal.

  • Do not pool merely because a preliminary test of equal variances failed to reject. Selecting the main test using the same data distorts the reported error rate.

  • The paired test applies when the design paired the observations. Applying it to independent groups, or a two-sample test to paired data, uses a standard error the design does not support.

  • For one proportion the null standard error uses p 0 , not p ^ ; the Wald interval uses p ^ . The two differ deliberately, and the test's version is the one consistent with computing under H 0 .

  • The t reference assumes approximate normality of the sampling distribution of the mean, which the central limit theorem supplies at moderate n for most populations but not for very small samples from skewed ones.

  • These tests draw their reference distribution from a sampling model. A randomization test derives it from the assignment mechanism instead, and the two answer different questions even when the arithmetic coincides.

Worked material

Example

The same numbers, two designs

Twenty people are measured under condition A and condition B. Mean A is 52.0, mean B is 48.4, so the observed gap is 3.6.

If the design used two independent groups of 20 people each, with s 1 = 9.0 and s 2 = 8.5 :

S E = 81 20 + 72.25 20 = 4.05 + 3.61 = 7.66 ≈ 2.77 ,
T = 3.6 2.77 ≈ 1.30 ,

which is unremarkable against a t distribution with roughly 38 degrees of freedom.

If the design measured the same 20 people under both conditions, the analysis works on the 20 within-person differences. Suppose s D = 4.2 :

S E = 4.2 20 ≈ 0.94 , T = 3.6 0.94 ≈ 3.83  on  19  degrees of freedom ,

which is decisive.

Same gap, opposite conclusions. Nothing about the estimate changed. What changed is the standard error, because the designs differ in what varies. Between people, scores differ a lot; within a person across two conditions, much less. The paired design removes the between-person variation, and the test must be the one that reflects it.

You cannot pick the test from the numbers alone. The design tells you which standard error is the right one, and getting that wrong changes the answer rather than refining it.

Non-example

Tests applied to the wrong structure

A two-sample test on paired data. Comparing all before-measurements with all after-measurements ignores that each pair shares a unit. The standard error comes from the wrong variation, usually overstating it and obscuring a real effect.

A paired test on independent groups. Pairing observations that were not paired by the design, arbitrarily matching the first treated unit with the first control, invents a structure the data do not have.

Pooling because a variance test did not reject. Running a preliminary test of equal variances and pooling when it fails to reject uses the same data twice and distorts the error rate of the main test. Welch's test needs no such decision.

Using p ^ in the null standard error. For H 0 : p = p 0 , the standard error should be computed under the null, using p 0 . Substituting p ^ answers a slightly different question and is the interval's formula, not the test's.

Concluding equivalence from a large p-value. "Not significant" is not "no difference". Showing two things are similar requires an equivalence test with a pre-specified margin, which is a different procedure.

Treating a randomization test as the same thing. The arithmetic may match a t -test, but the reference distribution comes from the assignment mechanism rather than a sampling model, and the null is a per-unit statement rather than one about means.

Contrast

What rejecting and not rejecting establish

Small p-valueLarge p-value
What it saysThe data would be surprising if H 0 were trueThe data would not be surprising if H 0 were true
LicensesRejecting H 0 at the stated levelNothing about H 0 being true
Compatible withA real effectA real effect too small for this study to detect
Correct wording"We reject the null""We fail to reject the null"
Needed to claim no effect—An equivalence test with a stated margin

Why the asymmetry is built in. A test asks how compatible the data are with one hypothesis. Data compatible with "no effect" are also compatible with "a small effect", and with "a large effect the study was too small to see". Compatibility does not single out the null.

The practical consequence. A null result should be reported with what the design could have detected. "We found no significant difference" is nearly uninformative on its own; "no significant difference, and the study could have detected a difference of 5 points with 80% power" tells a reader what was actually learned.

The wording matters. "Accept the null" invites the error. "Fail to reject" is clumsier and describes what happened.

Common errors

Common misconception

A p-value above the threshold shows the null hypothesis is true, so a non-significant result establishes that there is no effect or no difference.

Common misconception

Before comparing two independent means, test whether the variances are equal; if that test does not reject, pool the variances and use the pooled t-test, because the data have shown the assumption is satisfied.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.