Practice: Sampling Distributions and Standard Error
Question
Recognition · Interpretation
A sample of 100 observations has
2 hints available, least help first.
Hint 1: Retrieval cue
Ask which distribution the standard error is the standard deviation of.
Hint 2: Concept cue
One number describes the data you have; the other describes the samples you did not get.
Direct application · Interpretation · Explanation
A pilot of 40 measurements gives
(a) Compute
(d) The measurements are rescaled from units to tenths of a unit, so each value is multiplied by 10. State the new
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Write
Hint 2: Strategy cue
Express the needed change as a ratio, then square it.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
(a)
(b) Setting
(c) Precision improves with the square root of the sample size, not with the sample size itself. Going from 1.90 to 1.00 is a factor of about 1.9, so the sample must grow by roughly
A caveat. This calculation assumes the measurements are independent. If they are clustered, several per batch, per site or per operator, the required number is larger, and the design must account for the clustering rather than only the count.
(d) Rescaling. Multiplying every measurement by 10 gives
A complete answer does each of these:
- computes standard error
- scopes the clt
- detects dependence
- interprets variance rules
Comparison · Evaluation
A variable is strongly right-skewed in the population. A sample of 2,000 is drawn. Which statement is correct?
2 hints available, least help first.
Hint 1: Retrieval cue
State precisely which quantity the theorem is about.
Hint 2: Concept cue
Would plotting 2,000 skewed values produce a symmetric histogram?
Direct application
A random variable
Report
Enter the value. It is checked against the answer and the precision this task asks for.
Error diagnosis · Explanation · Evaluation
An analyst writes:
Household income in our sample is heavily right-skewed, but with 5,000 households the central limit theorem tells us the data are approximately normal. We therefore report that 95% of households earn between £18,400 and £71,600, using mean ± 1.96 standard deviations.
Identify the error and say what the central limit theorem does provide here.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Name the quantity the theorem's statement is about.
Hint 2: Concept cue
Ask whether the reported claim concerns individual households or an average.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The error. The central limit theorem is being applied to the observations. It concerns the sampling distribution of the mean, and says nothing about the distribution of household incomes, which remain heavily right-skewed however many households are surveyed. Sample size gives a clearer picture of that skew; it does not remove it.
Why the reported range is wrong. Using mean ± 1.96 standard deviations to describe where 95% of households fall assumes the incomes themselves are normal. With strong right skew, the true distribution has a long upper tail and a compressed lower one, so a symmetric interval will misstate both ends, and the lower endpoint may even fall below plausible values. Empirical quantiles of the observed incomes would answer this question directly and require no distributional assumption at all.
What the theorem does provide. That the sampling distribution of the mean income is approximately normal, which is ample with
The distinction to keep. Questions about individual households are about the observations. Questions about the average are about the mean. Only the second is what the theorem addresses.
A related slip to avoid. Reporting figures after a change of units invites the same confusion in reverse: a shift of origin leaves every spread statistic alone, while a change of scale multiplies the standard deviation by
A complete answer does each of these:
- computes standard error
- scopes the clt
- detects dependence
- interprets variance rules
Transfer · Evaluation · Explanation
A product team reports:
We analysed 45,000 page views from 3,200 users. Mean session duration was 4.2 minutes with a standard error of 0.02 minutes, so we can reliably detect changes of a few seconds.
Assess the claim and say what you would change.
The team also reports the same figures converted to seconds. Say which of the mean, the standard deviation and the standard error change, and by what factor.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Ask what the formula for
Hint 2: Strategy cue
Count the units that were sampled independently, not the rows in the table.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The problem. The 45,000 page views are not independent observations. They are nested within 3,200 users, and views by the same user are far more alike than views by different users, same device, same habits, same intent. The standard error appears to have been computed as
Direction and size of the error. Positive within-user correlation means the true standard error is larger than reported, so the error runs toward overconfidence. The effective sample size lies between 3,200 and 45,000, and approaches 3,200 as within-user correlation rises. Since
What I would change. Work at the user level: compute one summary per user, mean session duration for that user, and analyse those 3,200 values, which are plausibly independent if users were sampled independently. Alternatively use a method that models the clustering directly and reports an effective sample size. Either way the unit of analysis should match the unit of independence.
What else to check. Whether users themselves are independent. If many share households, organisations or referral sources, even 3,200 may overstate the information available.
Converting to seconds. Multiplying every session duration by 60 multiplies the mean, the standard deviation and the standard error by 60 alike:
A complete answer does each of these:
- computes standard error
- scopes the clt
- detects dependence
- interprets variance rules
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.