Practice: Neyman Repeated-Sampling Inference

Recognition · Interpretation

The exact randomization variance of τ ^ is S 1 2 N 1 + S 0 2 N 0 − S τ 2 N . Why is the third term omitted in practice?

2 hints available, least help first.

Hint 1: Retrieval cue

Write out what S τ 2 is a variance of, and what each summand needs.

Hint 2: Concept cue

Ask whether the obstacle is one that more data would remove.

Direct application · Interpretation

A completely randomized trial has N 1 = 50 treated with Y ¯ 1 = 12 , s 1 2 = 100 , and N 0 = 50 controls with Y ¯ 0 = 9 , s 0 2 = 144 . What are τ ^ and its conservative standard error?

1 hint available, least help first.

Hint 1: Retrieval cue

The conservative variance adds each arm's variance over its own size.

Direct application · Interpretation · Explanation

A completely randomized trial assigns 80 units to treatment and 120 to control. It reports τ ^ = 2.5 , s 1 2 = 160 , s 0 2 = 180 .

(a) Compute the conservative variance and standard error of τ ^ .

(b) Give a 95% confidence interval.

(c) Write the exact randomization variance symbolically and state which term the calculation in (a) does not use.

(d) Say whether the interval you reported is narrower or wider than one computed from the exact variance, and why you can answer that without knowing the missing term.

(e) Say why the omitted term cannot be estimated from this trial, however many units it had.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Each arm contributes its own sample variance divided by its own size.

Hint 2: Concept cue

For (d), consider the sign with which the missing term enters and whether a variance can be negative.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) Var ^ ( τ ^ ) = 160 / 80 + 180 / 120 = 2 + 1.5 = 3.5 , so S E = 3.5 ≈ 1.87 .

(b) 2.5 ± 1.96 ( 1.87 ) = ( − 1.17 , 6.17 ) . The interval includes zero, so the trial does not establish a non-zero average effect at this level.

(c) Exactly, Var ⁡ ( τ ^ ) = S 1 2 80 + S 0 2 120 − S τ 2 200 . The calculation omits S τ 2 / 200 , the variance of the individual treatment effects scaled by N .

(d) Wider, or at worst equal. S τ 2 is a variance and so is non-negative, and it enters the exact expression with a minus sign; omitting it therefore cannot reduce the reported variance. This holds whatever its value, which is why the answer does not depend on knowing it. The two coincide exactly when every unit has the same treatment effect.

(c) The omitted term is S τ 2 / N , the variance of the unit-level effects. Each τ i = Y i ( 1 ) − Y i ( 0 ) requires both potential outcomes for the same unit, and each unit was assigned once, so one of the two is never observed. Enlarging the trial multiplies half-observed units; it never completes a pair. The interval is therefore conservative by construction, not by approximation.

A complete answer does each of these:

  • computes estimate and se
  • identifies omitted term
  • explains conservatism
  • attributes to fundamental problem

Comparison · Method selection · Interpretation

Four analyses all report s 1 2 / N 1 + s 0 2 / N 0 as the standard error.

(i) A completely randomized trial, analysed as design-based Neyman inference.
(ii) A survey drawing two independent random samples from two populations, analysed with a Welch two-sample test.
(iii) A matched-pair experiment with 30 pairs, one unit per pair randomized to treatment.
(iv) A completely randomized trial in which the analyst additionally assumes every unit has the same treatment effect.

For each, state whether the expression is appropriate and what quantity the resulting interval refers to. Where it is inappropriate, say what should be used instead.

For each, say whether the quantity the Neyman variance omits is missing for the same reason.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

For each, ask what set of allocations the design permits.

Hint 2: Concept cue

Two of these share arithmetic with different reasoning; one has a different randomization distribution entirely; one changes what the omitted term is worth.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(i) Appropriate, and conservative. The interval refers to the finite-sample average treatment effect over the units in the trial, with randomness coming from the assignment.

(ii) The same arithmetic, a different justification. Randomness comes from sampling, the units are not fixed, and the interval refers to a difference of population means. Numerically identical, conceptually distinct, and the finite-population correction that appears in (i) has no analogue here.

(iii) Not appropriate. Pairing makes the two arms dependent by construction: the randomization distribution is over one choice per pair, 2 30 allocations, not over all ways of splitting 60 units. The pair-difference standard error s D / J should be used. Ignoring the pairing generally overstates the standard error, discarding the precision the design was built to gain.

(iv) Appropriate but no longer conservative. It becomes exact. Constant effects make S τ 2 = 0 , so the omitted term is zero rather than merely unknown. The assumption is not checkable from the data, so reporting on that basis should say so.

Why the term is missing. In the design-based cases the omission traces to the fundamental problem: S τ 2 is a variance of unit-level effects, and no unit reveals both of its potential outcomes. In the two-population survey there are no unit-level effects to have a variance, the two samples describe different people, so the same formula is exact there for a different reason, not conservative.

A complete answer does each of these:

  • computes estimate and se
  • identifies omitted term
  • explains conservatism
  • attributes to fundamental problem

Error diagnosis · Explanation · Evaluation

A statistician writes:

Our standard error omits S τ 2 / N , which makes it only an approximation. Our trial has 4,000 units, so we have plenty of data. I propose estimating S τ 2 by computing, for each treated unit, its outcome minus the control-arm mean, treating those as the individual effects, and taking their sample variance. We can then report the exact variance and a properly narrow interval.

Evaluate the proposal. Address what the plan actually computes, whether sample size bears on the problem, and whether "approximation" is the right word for what the usual estimator is.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Write out the quantity the plan computes for one treated unit, and compare it with the definition of τ i .

Hint 2: Concept cue

Separate three claims: what the plan estimates, whether more data helps, and what 'conservative' means as against 'approximate'.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

What the plan computes. For a treated unit, outcome minus the control-arm mean is Y i ( 1 ) − Y ¯ 0 . That is not τ i = Y i ( 1 ) − Y i ( 0 ) : the second term belongs to different units. Its sample variance mixes the variation in Y ( 1 ) across units with nothing about how effects differ, so it does not estimate S τ 2 . Substituting it would produce a number, and subtracting that number would yield an interval narrower than the design supports.

Sample size. Four thousand units give four thousand rows with one entry each. No τ i is formed for any of them. The obstacle is that the quantity is not identified, and identification is not something more observations supply.

"Approximation" is the wrong word. An approximation is close to a target and improves with effort. This estimator is conservative: its expectation is at least the true variance, so intervals are at least as wide as warranted and tests reject no more often than their level. That is a guarantee, in a known direction, not an error term. It is exact when every unit shares the same effect. A condition the data cannot confirm.

What to do instead. Report the conservative interval. If there is a substantive reason to believe effects are near-constant, say so explicitly as an assumption, and note that the data cannot check it.

Why no sample size rescues it. S τ 2 is the variance of the unit-level effects τ i = Y i ( 1 ) − Y i ( 0 ) . Computing a single τ i needs both potential outcomes for that unit, and assignment reveals exactly one. Adding units adds more half-observed pairs, never a complete one, so the term is unidentified by the structure of the problem rather than by a shortage of data or a defect in the estimator.

A complete answer does each of these:

  • computes estimate and se
  • identifies omitted term
  • explains conservatism
  • attributes to fundamental problem

Transfer · Evaluation · Interpretation

A published trial of a job-training programme randomized 1,200 unemployed applicants and reports: average effect on twelve-month earnings of £1,840, 95% interval £310 to £3,370.

A journalist writes that the trial shows the programme is worth between £310 and £3,370 to a participant, and that some participants may have gained far more but the study was too small to detect it.

Assess both claims. For each, say what the trial does support, and identify which quantity the journalist would need in order to make the claim they made.

The journalist asks whether a larger trial would have produced a narrower interval by pinning down the variation between individuals. Answer that.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Ask what the interval is an interval for.

Hint 2: Strategy cue

Treat the two claims separately: one confuses an average with an individual, the other misattributes an identification limit to sample size.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

First claim — 'worth between £310 and £3,370 to a participant'. The interval covers the average effect across the 1,200 applicants, not any individual's gain. It is compatible with every participant gaining about £1,840, and equally with half gaining £3,700 and half gaining nothing. To say what a participant gained, the journalist would need that participant's τ i , which requires their earnings both with and without the programme. One of which never happened.

Second claim — 'the study was too small to detect' individual variation. This misidentifies the obstacle. Sample size determines the width of the interval around the average; it does nothing about the spread of individual effects. The quantity describing that spread is S τ 2 , and it is not identified at any sample size, because forming even one τ i needs both potential outcomes for one person. A trial of 120,000 applicants would give a much narrower interval around the average and exactly as little information about who gained most.

What the trial supports. A positive average effect on twelve-month earnings for this population of applicants, with the stated uncertainty, and that uncertainty is conservative, since the reported variance omits a non-negative term.

What would be needed instead. Statements about who benefits require either strong assumptions about effect heterogeneity, or a design targeting subgroup effects defined by pre-treatment characteristics, which estimates averages within subgroups, still not individual effects.

Would a larger trial pin it down? Not that component. A larger trial shrinks s 1 2 / N 1 + s 0 2 / N 0 , so the interval does narrow. But the piece deliberately left out, S τ 2 / N , is the variance of individual effects, and an individual effect needs both of that person's potential outcomes. Each applicant was either trained or not. No trial size changes that, so the conservatism survives at any scale.

A complete answer does each of these:

  • computes estimate and se
  • identifies omitted term
  • explains conservatism
  • attributes to fundamental problem
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.