Module 3 of 3 · Lesson 3 of 3
Binary Outcome Models for Experimental Research
Binary outcomes as conditional probabilities, two competing models, and the reporting scale.
What you will be able to do
Given a binary-outcome analysis, the learner can say what each model reports, convert or refuse to convert between scales, and identify a result stated on the wrong one.
Orientation
An odds ratio of 2 sounds like doubling. Whether it corresponds to a risk rising from 1% to 2% or from 40% to 57% depends on where you started, and a decision needs the second number.
The models are two. The competence is keeping the reporting scale explicit, because the most common error here is a silent conversion that changes the magnitude of the finding.
Intuition
Reading a coefficient on each scale
The canonical intuition states why a binary outcome lets a mean model serve as a probability model, and why the two families differ in reporting scale. This block works through what each scale says about one fitted result.
Suppose a treatment indicator has a linear-probability coefficient of
On the probability scale. The linear model reports a rise of
On the odds scale.
Why the same odds ratio gives different probability changes. At a control probability of
A reported effect is therefore incomplete without its scale, and an odds ratio is incomplete without the baseline probability it is applied to.
Definition
Two models and what each reports
The canonical statement specifies both models. What decides between them, and what to report afterwards, is not in the specification.
The two models disagree about what is being held linear. The linear probability model makes the probability linear in the predictors, so a coefficient is a change in probability and reads directly. Logistic regression makes the log-odds linear, so a coefficient is a change in log-odds and reads only after conversion. Neither is the default: the choice is between a quantity that is easy to report and a fit that cannot leave
An odds ratio is not a risk ratio, and the gap depends on the base rate. At an outcome probability near
Marginal effects are what put a logistic fit back on the scale the question was asked in.
Example
One odds ratio, two very different effects
Two trials, each reporting an odds ratio of 2.0 for treatment.
Trial A: a rare outcome. Control event rate 1%. Odds in control:
- Risk difference: about 0.98 percentage points
- Risk ratio: about 1.98
The odds ratio of 2.0 and the risk ratio of 1.98 are nearly identical. Reading "twice as likely" is approximately right.
Trial B: a common outcome. Control event rate 33%. Odds in control:
- Risk difference: about 16.6 percentage points
- Risk ratio: about 1.50
Here the odds ratio of 2.0 corresponds to a risk ratio of 1.5. Reading "twice as likely" overstates the effect by a third.
Why this matters for decisions. Trial A's treatment moves one person per hundred; Trial B's moves seventeen. Both report
What to report. The baseline rate alongside any odds ratio, or better, absolute risks in both arms and their difference.
Worked example
Converting to a decision scale
Problem. A randomized trial of a reminder system on appointment attendance has 800 patients, 400 per arm. Attendance: 312 of 400 treated, 268 of 400 control. A logistic regression of attendance on treatment gives
Goal. Assess the claim and report the effect on a decision scale.
Relevant principle. An odds ratio is a ratio of odds, not of probabilities; with a common outcome the two diverge.
Step 1: compute the proportions.
Step 2: confirm the odds ratio. Odds treated
Reported as
Reason: the coefficient is on the log-odds scale, so exponentiating returns a multiplicative change in odds.
Step 3: identify the error. "55% more likely" reads the odds ratio as a risk ratio. The actual risk ratio is
a 16% relative increase, not 55%.
Reason: attendance is a common outcome, around 70%, so odds and risks diverge substantially.
Step 4: give the decision-relevant figure. The risk difference is
11 percentage points. Across 400 patients that is about 44 additional attendances.
Reason: an absolute difference is what capacity planning and cost-per-attendance calculations require.
Step 5: note the simpler route. Because the trial was randomized,
Result. Risk difference 11 percentage points; risk ratio 1.16; odds ratio 1.75. The claim of "55% more likely" overstates the relative effect roughly threefold.
Check. Does the odds ratio exceed the risk ratio, as it should for an outcome above 50%? Yes, 1.75 against 1.16, and the gap is large precisely because attendance is common.
Interpretation. Report the two attendance rates, their difference with an interval, and the implied number of additional attendances. If an odds ratio is reported at all, give the baseline rate beside it so a reader can recover the absolute effect.
Non-example
Statements these models do not support
"An odds ratio of 2 means twice as likely." True only when the outcome is rare. At a 33% baseline it corresponds to a risk ratio of about 1.5.
Reporting a logistic coefficient as a change in probability.
Comparing logistic coefficients across models with different covariates. The scale itself shifts when covariates are added, even covariates unrelated to treatment, so the coefficients are not comparable in the way OLS coefficients are.
Rejecting the linear probability model solely because it can predict outside
Using classical standard errors with the linear probability model. The error variance depends on the fitted probability by construction, so heteroskedasticity is guaranteed, not merely possible.
Confusing an outcome model with a propensity model. Logistic regression appears in this subject for both jobs. One estimates
Contrast
Choosing the reporting scale
| Linear probability model | Logistic regression | |
|---|---|---|
| Models | ||
| Coefficient reads as | Change in probability | Change in log-odds; |
| Fitted values | May leave | Always in |
| Standard errors | Robust, always — heteroskedasticity is structural | Model-based or robust |
| Directly decision-ready | Yes, a probability difference | No, needs marginal effects |
| Comparable across specifications | Yes | No — the scale shifts |
Why the odds scale causes so much trouble. Odds are unfamiliar outside betting, and an odds ratio sounds like it should be a ratio of chances. It is a ratio of
The baseline rate resolves it. Given the control-arm probability, any odds ratio converts to a risk difference and a risk ratio. Reporting an odds ratio without the baseline leaves a reader unable to recover the magnitude.
Which to prefer. If the estimand is an average difference in probability, usually the case in a randomized trial, the linear probability model or plain proportions report it directly. If fitted probabilities near the boundaries matter, or the design calls for it, logistic regression plus marginal effects gets to the same scale by a longer route.
What both share. Neither supplies a causal reading. That comes from the design, exactly as in the continuous-outcome case.
Exercise
1: fully structured. A logistic model gives
(a) What is the odds ratio? (b) If the control-arm probability is 0.10, what is the treated probability? (c) What is the risk difference?
Check: (a)
2: partly structured. A trial with a 50% control event rate reports
(a) Compute the treated probability. (b) Give the risk ratio. (c) Why does "80% more likely" mislead here?
Check: (a) control odds
3: unstructured. A press release states: "Our randomized trial found the app doubles the odds of quitting smoking (OR = 2.1, p < 0.001). Smokers using the app are twice as likely to quit."
The control-arm quit rate was 22%. Assess the claim and say what should have been reported.
Check: the second sentence converts an odds ratio into a risk ratio, which is not valid at a 22% baseline. Control odds
What to carry forward
The starting point. For
Linear probability model.
Logistic regression.
An odds ratio is not a risk ratio. Nor a probability difference. They approximately coincide only when the outcome is rare, and diverge most near a 50% baseline.
Converting requires a baseline. Given the control probability, an odds ratio yields the treated probability, the risk ratio and the risk difference. Report the baseline beside any odds ratio.
Marginal effects. What logistic output needs to reach the probability scale.
In a randomized trial.
Two jobs, one method. Logistic regression models outcomes here and treatment assignment in a propensity score. Different estimands, different criteria.
The recurring error. Reporting an odds ratio as though it were a risk ratio.