Module 1 of 3 · Lesson 4 of 4

Probability and Sampling Distributions

What you will be able to do

Given a described sampling situation, the learner can obtain the standard error of a mean, say what the central limit theorem does and does not claim, and identify when the independence the formula assumes fails.

Orientation

A statistic computed from one sample would have come out differently from another. Naming that variation, and knowing when the usual formula for it does not apply, is the foundation of every interval and test that follows.

Every interval, test and design-based variance later in this subject is built on the sampling distribution. Its spread is the standard error, its shape determines a confidence interval, and a p-value is a position within it.

Intuition

Sampling distributions and standard error

You compute an average from your sample and get 47.3. Had a different set of units ended up in the sample, you would have got something else.

The sampling distribution is the distribution of all those numbers you did not get. It is not something you can see in your data, you only ever observe one draw from it, but it is what every inferential statement describes.

This is why a standard error is not the spread of your observations. The observations might range from 10 to 90 while the mean of 400 of them barely moves from sample to sample. The standard error measures how much the estimate would move, which is a question about the procedure rather than about the numbers in front of you.

Averaging is what makes it small. Individual values scatter, but their deviations partly cancel, and Var ⁡ ( X ¯ ) = σ 2 / n follows. The practical consequence is the 1 / n rate: to halve a standard error you need four times the data, not twice.

The second result is the central limit theorem, and it is the one most often overstated. It says the standardized sample mean tends toward a normal distribution. It does not say your observations become normal. A skewed population stays skewed however many units you draw from it, what becomes approximately normal is the thing you averaged into, not the thing you averaged over.

Definition

Summaries, variance rules, and the two results

Sample summaries. For observations X 1 , … , X n ,

X ¯ = 1 n ∑ i = 1 n X i , s 2 = 1 n − 1 ∑ i = 1 n ( X i − X ¯ ) 2 .

The divisor n − 1 makes s 2 unbiased for the population variance under independent, identically distributed sampling. It corrects for having estimated μ by X ¯ , which makes the deviations slightly too small.

Expectation and variance. For a discrete random variable E [ X ] = ∑ x x P ( X = x ) ; for a continuous one with density f , E [ X ] = ∫ x f ( x ) d x . Variance may be written either way:

Var ⁡ ( X ) = E [ ( X − E [ X ] ) 2 ] = E [ X 2 ] − E [ X ] 2 .

For constants a and b ,

E [ a X + b ] = a E [ X ] + b , Var ⁡ ( a X + b ) = a 2 Var ⁡ ( X ) .

Adding a constant shifts location and leaves spread alone; multiplying by a scales the standard deviation by | a | and the variance by a 2 .

Standard error of the mean. If the observations are independent with common variance σ 2 ,

Var ⁡ ( X ¯ ) = σ 2 n , S E ( X ¯ ) = σ n ≈ s n .

Central limit theorem. Under regularity conditions,

X ¯ − μ σ / n ⟶ d N ( 0 , 1 ) ,

used in practice as the approximation X ¯ − μ σ / n ≈ N ( 0 , 1 ) for sufficiently large n . The statement concerns the sample mean.

Example

Spread of the data against spread of the estimate

A hospital records length of stay for 900 discharges. The distribution is heavily right-skewed: most stays are 2–4 days, a few run past 60. The sample mean is 4.8 days with s = 6.0 days.

Spread of the observations. s = 6.0 days. Individual stays vary enormously, and they always will. This is a real feature of the population, not noise to be averaged away.

Spread of the estimate.

S E ( X ¯ ) = 6.0 900 = 6.0 30 = 0.20  days .

The mean of 900 stays would move by only a fifth of a day or so across repetitions. These two numbers describe different things, and only the second is about the reliability of the estimate.

What the central limit theorem licenses here. That X ¯ is approximately normal, so the usual interval around 4.8 is reasonable. With n = 900 and skew this strong, that approximation is doing real work and is plausible at this sample size.

What it does not license. Any claim that lengths of stay are normal. Plot them and the skew is unmistakable. If the question were "what proportion of patients stay longer than 14 days?", the normal approximation would answer it badly, that question is about the observations, and no sample size makes their distribution symmetric.

The distinction in one line. The CLT is about X ¯ . Questions about individual units are not questions the CLT addresses.

Worked example

How much more data?

Problem. A pilot study of 50 customers estimates mean monthly spend at £82 with s = £ 35 . The team wants the standard error down to £2.00 before committing to a decision.

How many customers are needed? And what if customers are sampled from within 10 stores rather than independently?

Goal. The required sample size, and a judgement about whether the formula applies.

Relevant principle. S E ( X ¯ ) = s / n for independent observations, so the standard error falls with the square root of the sample size.

Step 1: the current standard error.

S E = 35 50 = 35 7.07 ≈ £ 4.95 .

Reason: the estimated standard deviation over the square root of the sample size.

Step 2: solve for the target. Setting 35 / n = 2.00 ,

n = 35 2.00 = 17.5 , n = 17.5 2 ≈ 306.

Reason: the relationship inverts directly; n is the square of the ratio of standard deviation to target standard error.

Step 3: read the cost. Going from £4.95 to £2.00 is a factor of about 2.5, and the sample must grow by about 2.5 2 ≈ 6.2 times, from 50 to roughly 306.

Reason: the 1 / n rate means each further halving costs four times the data already collected.

Step 4: check independence. If the 306 customers come from 10 stores, they are not independent. Customers at one store share catchment, pricing, staff and local conditions, so two customers from the same store carry less information than two from different stores.

Reason: Var ⁡ ( X ¯ ) = σ 2 / n assumes independence. With positive within-store correlation the true variance is larger, and dividing by 306 overstates the precision.

Step 5: say what changes. The effective sample size sits somewhere between 10 and 306, closer to 10 as within-store correlation rises. The design must account for clustering, and the store count may bind more tightly than the customer count.

Result. About 306 customers if sampled independently; more, possibly many more, under a clustered design, which also needs a different variance calculation.

Check. Does the answer scale sensibly? Halving the standard error from £4.95 to about £2.48 would need 4 × 50 = 200 customers. The target of £2.00 is slightly tighter, and 306 exceeds 200. Consistent.

Interpretation. Report the required sample size together with the sampling scheme it assumes. A number derived under independence and delivered by a clustered design will not produce the precision it promised.

Non-example

Situations the usual standard error does not describe

Pupils within classrooms. 600 pupils across 20 classrooms are not 600 independent observations. Pupils share a teacher, a room and a peer group. Dividing by 600 can badly overstate precision.

Repeated measures on one person. Twelve blood-pressure readings from each of 30 patients is not n = 360 . The readings within a patient are far more alike than readings across patients.

A time series. Daily sales for two years are serially dependent: today resembles yesterday. Treating 730 days as 730 independent draws understates the standard error.

A convenience sample. Volunteers who responded to an advertisement have a sampling distribution the formula does not describe, because they were not drawn by a known mechanism from the target population. The arithmetic still runs, which is exactly the danger.

Using the CLT for an individual prediction. "Roughly 95% of patients stay between 4.4 and 5.2 days" misapplies an interval for the mean to individual units. That interval describes where the average sits, not where a patient falls.

A sample standard deviation reported as a standard error. These differ by a factor of n . At n = 900 that is a factor of 30, so the mistake is not a small one.

Contrast

Two distributions people conflate

Distribution of the observationsSampling distribution of X ¯
What variesIndividual unitsThe estimate, across repetitions
Spread measured by s S E = s / n
Shape as n growsUnchanged — a skewed population stays skewedApproaches normal, by the CLT
ObservableYes, plot the dataNo, only one draw is ever seen
AnswersHow do units differ?How reliable is the estimate?

Where the confusion does damage. The claim " n is large, so by the central limit theorem the data are approximately normal" moves a result about column two into column one. It is used to justify procedures that depend on the observations being normal, prediction intervals for individuals, reference ranges, tolerance limits, and no sample size supports them.

The clean test. Ask what the statement is about. If it is about individual units, the CLT is silent. If it is about an average or a total, the CLT applies.

Why n still helps the mean. Not by reshaping the population, but by averaging: deviations partly cancel, the variance of the average shrinks as σ 2 / n , and the shape of that average's distribution tends to normal whatever the population looked like. Both effects are about X ¯ .

Exercise

1: fully structured. A sample of 64 components has mean weight 250 g and s = 16 g.

(a) Compute S E ( X ¯ ) . (b) How many components would be needed to halve it? (c) A colleague says 95% of components weigh between 246 and 254 g. What is wrong?

Check: (a) 16 / 64 = 16 / 8 = 2 g; (b) 256 components, four times as many; (c) that interval is built from the standard error and describes where the mean lies, not individual components. The spread of components is governed by s = 16 g, eight times wider.

2: partly structured. A team measures customer satisfaction on a 1–5 scale. The responses are strongly left-skewed, with most at 4 and 5. They have 1,200 responses.

(a) Does the CLT make the responses approximately normal? (b) Does it help with the mean? (c) What would you check before using a normal-based interval for the mean?

Check: (a) no. The responses are discrete and bounded and remain so at any sample size; (b) yes, the sampling distribution of the mean will be close to normal with n = 1,200 , even for a skewed bounded variable; (c) that the responses are genuinely independent, if respondents cluster by account, region or survey wave, the standard error needs to reflect it, and that matters more here than the skew does.

3: unstructured. An analyst reports: "We analysed 45,000 page views from 3,200 users. Mean session duration was 4.2 minutes with a standard error of 0.02 minutes, so we can detect changes of a few seconds."

Assess the claim and say what you would change.

Check: the 45,000 page views are not independent. They are nested within 3,200 users, and views by one user are far more alike than views across users. The standard error appears to have been computed as s / 45,000 , which overstates precision, possibly by a large factor. The effective sample size is nearer 3,200 than 45,000, and closer still to 3,200 the stronger the within-user correlation. The analysis should work at the user level, one summary per user, or use a method that accounts for clustering. The claimed ability to detect a few seconds is not supported until that is redone, and the direction of the error is always toward overconfidence.

What to carry forward

The sampling distribution. The distribution of a statistic across the samples the procedure could have produced. A property of the estimator and design, not of the data in hand.

Sample summaries. X ¯ = 1 n ∑ i X i and s 2 = 1 n − 1 ∑ i ( X i − X ¯ ) 2 ; the n − 1 corrects for estimating μ by X ¯ .

Variance rules. E [ a X + b ] = a E [ X ] + b and Var ⁡ ( a X + b ) = a 2 Var ⁡ ( X ) . Adding a constant does not change spread.

Standard error of the mean. S E ( X ¯ ) = σ / n ≈ s / n , requiring independence.

The rate. Precision improves like 1 / n : halving a standard error takes four times the data.

Central limit theorem. X ¯ − μ σ / n ⟶ d N ( 0 , 1 ) . A statement about the sample mean.

What it does not say. That the observations become normal. A skewed population stays skewed at any n .

When the formula fails. Clustering, repeated measures, serial dependence. All make the true standard error larger than s / n , so the error runs toward overconfidence.

The recurring error. Treating a large n as licence to call the data normal.

Next step

Practice Probability and Sampling Distributions

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Hypothesis Tests for Experimental Research

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.