Practice: Binary Outcome Models for Experimental Research
Question
Recognition · Interpretation
A logistic regression gives
2 hints available, least help first.
Hint 1: Retrieval cue
What quantity does logistic regression model as linear in
Hint 2: Concept cue
Exponentiating a log-odds difference gives a ratio of what?
Method selection · Direct application · Interpretation
A randomized trial of a reminder system has 400 patients per arm. Attendance: 312 of 400 treated, 268 of 400 control.
Give the risk difference, the risk ratio and the odds ratio. Say which you would report to a clinic manager planning capacity, and why a model is not needed for the headline figure.
A colleague proposes fitting a linear probability model instead. Say what its coefficient would report here, and what you would qualify.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Compute both proportions first, then form each of the three comparisons.
Hint 2: Strategy cue
Ask what number the manager would multiply by patient volume.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The proportions.
Risk difference.
Risk ratio.
Odds ratio. Odds treated
Which to report. The risk difference. A capacity plan needs to know how many more patients will attend, and 11 percentage points converts directly into rooms, staff and appointment slots. The odds ratio of 1.75 cannot be turned into that number without the baseline rate, and would be actively misleading if read as '75% more likely'. The actual relative increase is 16%.
Why no model is needed. Treatment was randomized, so the difference in sample proportions already estimates the causal risk difference. Fitting a logistic model would answer the same question on a less useful scale and add nothing to the identification.
What to add. A confidence interval for the difference, since 11 points with an interval of 5 to 17 supports a firmer plan than one spanning 1 to 21.
The linear probability model here. Its treatment coefficient would report the risk difference directly, the same
A complete answer does each of these:
- identifies the scale
- separates odds from risk
- reports on a decision scale
- handles lpm tradeoffs
Comparison · Interpretation
Two trials both report
2 hints available, least help first.
Hint 1: Retrieval cue
Convert each baseline to odds, double it, and convert back to a probability.
Hint 2: Concept cue
When
Error diagnosis · Explanation · Evaluation
A press release states:
Our randomized trial found the app doubles the odds of quitting smoking (OR = 2.1, p < 0.001). Smokers using the app are twice as likely to quit.
The control-arm quit rate was 22%. Identify the error and say what should have been reported.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Convert the 22% baseline to odds, apply the ratio, and convert back.
Hint 2: Concept cue
Ask whether 'twice as likely' is a statement about odds or about probabilities.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The error. The second sentence converts an odds ratio into a risk ratio, which is not valid at a 22% baseline. 'Doubles the odds' is accurate; 'twice as likely' is a claim about probabilities and is not what was estimated.
The correct conversion. Control odds
- Risk ratio:
. A 69% relative increase, not 100%. - Risk difference:
percentage points.
Why it matters. The result is genuinely good, fifteen additional quitters per hundred smokers is a substantial public-health effect. The error overstates the relative effect and, more importantly, buries the absolute one, which is what a reader needs to judge whether the app is worth using or funding.
What should have been reported. The quit rates in both arms, 22% and 37%, their difference of about 15 percentage points with a confidence interval, and the odds ratio only alongside the baseline rate so a reader can recover the absolute effect.
A simpler route. The trial was randomized, so the difference in sample proportions estimates the causal risk difference directly. The headline never needed a model at all.
What would have avoided the error. Reporting on the probability scale in the first place. A linear probability model would have returned the risk difference directly, with robust standard errors for the heteroskedasticity a binary outcome always carries. Its fitted values can stray outside
A complete answer does each of these:
- identifies the scale
- separates odds from risk
- reports on a decision scale
- handles lpm tradeoffs
Transfer · Evaluation · Explanation
An insurer's underwriting note says:
Our claims model gives an odds ratio of 1.35 for properties in flood zone 3. These properties are 35% more likely to claim, so we are raising premiums by 35% for that zone.
Baseline claim rate is 8%. Assess the reasoning and say what you would report.
Write your answer, then compare it with the worked solution.
2 hints available, least help first.
Hint 1: Retrieval cue
Convert the baseline to odds, apply the ratio, and convert back to a probability.
Hint 2: Strategy cue
Ask what quantity a premium should actually be proportional to.
Compare with the worked solution
Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.
The scale error. An odds ratio of 1.35 is not a 35% increase in the probability of claiming. Baseline odds
- Risk ratio:
, a 31% relative increase. - Risk difference: about 2.5 percentage points.
At an 8% baseline the odds ratio and risk ratio are reasonably close, so the relative error here is modest, but the reasoning is still wrong, and the same habit applied to a common outcome would misstate the figure badly.
The larger error is the pricing logic. Even the correct 31% relative increase does not justify a 35% premium rise, and neither would a 31% rise. Premiums are set against expected claim cost, which depends on the absolute claim probability and the severity per claim. The absolute risk rose by 2.5 percentage points; whether that warrants a 35% premium increase depends on the loading, the expected claim size and the existing margin, none of which the odds ratio speaks to.
What I would report. Claim probabilities in each zone, 8% against 10.5%, with intervals; the absolute difference of 2.5 points; and an expected-cost calculation combining that with average claim severity. If an odds ratio appears at all, the baseline rate belongs beside it.
One further check. Whether flood zone is a cause of claiming or a marker for other correlated property characteristics, for pricing, prediction may suffice, but the note should not be read as establishing a causal effect of the zone designation.
The alternative the insurer should consider. Fitting the claim indicator linearly would give the change in claim probability directly, which is what a premium calculation consumes. The cost is that such a model can predict probabilities outside
A complete answer does each of these:
- identifies the scale
- separates odds from risk
- reports on a decision scale
- handles lpm tradeoffs
Session complete
Every question in this set has been through once. What you can do now depends on how it went — practising again is worth more than moving on if any of it was uncertain.
Practice data
Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.