Module 1 of 3 · Lesson 4 of 4
Probability and Sampling Distributions
What you will be able to do
Given a described sampling situation, the learner can obtain the standard error of a mean, say what the central limit theorem does and does not claim, and identify when the independence the formula assumes fails.
Orientation
A statistic computed from one sample would have come out differently from another. Naming that variation, and knowing when the usual formula for it does not apply, is the foundation of every interval and test that follows.
Every interval, test and design-based variance later in this subject is built on the sampling distribution. Its spread is the standard error, its shape determines a confidence interval, and a p-value is a position within it.
Intuition
Sampling distributions and standard error
You compute an average from your sample and get 47.3. Had a different set of units ended up in the sample, you would have got something else.
The sampling distribution is the distribution of all those numbers you did not get. It is not something you can see in your data, you only ever observe one draw from it, but it is what every inferential statement describes.
This is why a standard error is not the spread of your observations. The observations might range from 10 to 90 while the mean of 400 of them barely moves from sample to sample. The standard error measures how much the estimate would move, which is a question about the procedure rather than about the numbers in front of you.
Averaging is what makes it small. Individual values scatter, but their deviations partly cancel, and
The second result is the central limit theorem, and it is the one most often overstated. It says the standardized sample mean tends toward a normal distribution. It does not say your observations become normal. A skewed population stays skewed however many units you draw from it, what becomes approximately normal is the thing you averaged into, not the thing you averaged over.
Definition
Summaries, variance rules, and the two results
Sample summaries. For observations
The divisor
Expectation and variance. For a discrete random variable
For constants
Adding a constant shifts location and leaves spread alone; multiplying by
Standard error of the mean. If the observations are independent with common variance
Central limit theorem. Under regularity conditions,
used in practice as the approximation
Example
Spread of the data against spread of the estimate
A hospital records length of stay for 900 discharges. The distribution is heavily right-skewed: most stays are 2–4 days, a few run past 60. The sample mean is 4.8 days with
Spread of the observations.
Spread of the estimate.
The mean of 900 stays would move by only a fifth of a day or so across repetitions. These two numbers describe different things, and only the second is about the reliability of the estimate.
What the central limit theorem licenses here. That
What it does not license. Any claim that lengths of stay are normal. Plot them and the skew is unmistakable. If the question were "what proportion of patients stay longer than 14 days?", the normal approximation would answer it badly, that question is about the observations, and no sample size makes their distribution symmetric.
The distinction in one line. The CLT is about
Worked example
How much more data?
Problem. A pilot study of 50 customers estimates mean monthly spend at £82 with
How many customers are needed? And what if customers are sampled from within 10 stores rather than independently?
Goal. The required sample size, and a judgement about whether the formula applies.
Relevant principle.
Step 1: the current standard error.
Reason: the estimated standard deviation over the square root of the sample size.
Step 2: solve for the target. Setting
Reason: the relationship inverts directly;
Step 3: read the cost. Going from £4.95 to £2.00 is a factor of about 2.5, and the sample must grow by about
Reason: the
Step 4: check independence. If the 306 customers come from 10 stores, they are not independent. Customers at one store share catchment, pricing, staff and local conditions, so two customers from the same store carry less information than two from different stores.
Reason:
Step 5: say what changes. The effective sample size sits somewhere between 10 and 306, closer to 10 as within-store correlation rises. The design must account for clustering, and the store count may bind more tightly than the customer count.
Result. About 306 customers if sampled independently; more, possibly many more, under a clustered design, which also needs a different variance calculation.
Check. Does the answer scale sensibly? Halving the standard error from £4.95 to about £2.48 would need
Interpretation. Report the required sample size together with the sampling scheme it assumes. A number derived under independence and delivered by a clustered design will not produce the precision it promised.
Non-example
Situations the usual standard error does not describe
Pupils within classrooms. 600 pupils across 20 classrooms are not 600 independent observations. Pupils share a teacher, a room and a peer group. Dividing by
Repeated measures on one person. Twelve blood-pressure readings from each of 30 patients is not
A time series. Daily sales for two years are serially dependent: today resembles yesterday. Treating 730 days as 730 independent draws understates the standard error.
A convenience sample. Volunteers who responded to an advertisement have a sampling distribution the formula does not describe, because they were not drawn by a known mechanism from the target population. The arithmetic still runs, which is exactly the danger.
Using the CLT for an individual prediction. "Roughly 95% of patients stay between 4.4 and 5.2 days" misapplies an interval for the mean to individual units. That interval describes where the average sits, not where a patient falls.
A sample standard deviation reported as a standard error. These differ by a factor of
Contrast
Two distributions people conflate
| Distribution of the observations | Sampling distribution of | |
|---|---|---|
| What varies | Individual units | The estimate, across repetitions |
| Spread measured by | ||
| Shape as | Unchanged — a skewed population stays skewed | Approaches normal, by the CLT |
| Observable | Yes, plot the data | No, only one draw is ever seen |
| Answers | How do units differ? | How reliable is the estimate? |
Where the confusion does damage. The claim "
The clean test. Ask what the statement is about. If it is about individual units, the CLT is silent. If it is about an average or a total, the CLT applies.
Why
Exercise
1: fully structured. A sample of 64 components has mean weight 250 g and
(a) Compute
Check: (a)
2: partly structured. A team measures customer satisfaction on a 1–5 scale. The responses are strongly left-skewed, with most at 4 and 5. They have 1,200 responses.
(a) Does the CLT make the responses approximately normal? (b) Does it help with the mean? (c) What would you check before using a normal-based interval for the mean?
Check: (a) no. The responses are discrete and bounded and remain so at any sample size; (b) yes, the sampling distribution of the mean will be close to normal with
3: unstructured. An analyst reports: "We analysed 45,000 page views from 3,200 users. Mean session duration was 4.2 minutes with a standard error of 0.02 minutes, so we can detect changes of a few seconds."
Assess the claim and say what you would change.
Check: the 45,000 page views are not independent. They are nested within 3,200 users, and views by one user are far more alike than views across users. The standard error appears to have been computed as
What to carry forward
The sampling distribution. The distribution of a statistic across the samples the procedure could have produced. A property of the estimator and design, not of the data in hand.
Sample summaries.
Variance rules.
Standard error of the mean.
The rate. Precision improves like
Central limit theorem.
What it does not say. That the observations become normal. A skewed population stays skewed at any
When the formula fails. Clustering, repeated measures, serial dependence. All make the true standard error larger than
The recurring error. Treating a large