Binary Outcome Models for Experimental Research

What you will be able to do

Given a binary-outcome analysis, the learner can say what each model reports, convert or refuse to convert between scales, and identify a result stated on the wrong one.

Orientation

An odds ratio of 2 sounds like doubling. Whether it corresponds to a risk rising from 1% to 2% or from 40% to 57% depends on where you started, and a decision needs the second number.

The models are two. The competence is keeping the reporting scale explicit, because the most common error here is a silent conversion that changes the magnitude of the finding.

Intuition

Reading a coefficient on each scale

The canonical intuition states why a binary outcome lets a mean model serve as a probability model, and why the two families differ in reporting scale. This block works through what each scale says about one fitted result.

Suppose a treatment indicator has a linear-probability coefficient of 0.06 and a logistic coefficient of 0.41 , and that the control-group probability is 0.30 .

On the probability scale. The linear model reports a rise of 6 percentage points, from 0.30 to 0.36 , for any unit and any covariate values. That constancy is the modelling assumption, and it is what allows the fitted probability to leave [ 0 , 1 ] once the predictor is far enough from the data.

On the odds scale. e 0.41 = 1.507 , so the odds are multiplied by about 1.5 . Control odds are 0.30 / 0.70 = 0.4286 ; treated odds are 0.4286 × 1.507 = 0.6459 ; the treated probability is 0.6459 / 1.6459 = 0.3924 . The rise is 9.2 percentage points here, and would be a different number at a different baseline.

Why the same odds ratio gives different probability changes. At a control probability of 0.02 , an odds ratio of 1.507 gives a treated probability of 0.0298 , a rise of 1 percentage point. At 0.50 it gives 0.6011 , a rise of 10 points. The odds ratio is constant across baselines by construction; the probability difference is not.

A reported effect is therefore incomplete without its scale, and an odds ratio is incomplete without the baseline probability it is applied to.

Definition

Two models and what each reports

The canonical statement specifies both models. What decides between them, and what to report afterwards, is not in the specification.

The two models disagree about what is being held linear. The linear probability model makes the probability linear in the predictors, so a coefficient is a change in probability and reads directly. Logistic regression makes the log-odds linear, so a coefficient is a change in log-odds and reads only after conversion. Neither is the default: the choice is between a quantity that is easy to report and a fit that cannot leave [ 0 , 1 ] .

An odds ratio is not a risk ratio, and the gap depends on the base rate. At an outcome probability near 0.01 the two nearly coincide: an odds ratio of 2.0 corresponds to a risk ratio of about 1.98 . At a probability near 0.25 they part: the same odds ratio of 2.0 is a risk ratio of about 1.50 . Reporting an odds ratio as though it were a relative risk overstates the effect wherever the outcome is common.

Marginal effects are what put a logistic fit back on the scale the question was asked in. ∂ P / ∂ X j varies with X , so an average marginal effect summarises it over the sample actually observed. That is the number to report when the audience asked about probability, and it is a different quantity from e β j .

Example

One odds ratio, two very different effects

Two trials, each reporting an odds ratio of 2.0 for treatment.

Trial A: a rare outcome. Control event rate 1%. Odds in control: 0.01 / 0.99 ≈ 0.0101 . Doubling gives 0.0202 , so p 1 = 0.0202 / 1.0202 ≈ 1.98 % .

  • Risk difference: about 0.98 percentage points
  • Risk ratio: about 1.98

The odds ratio of 2.0 and the risk ratio of 1.98 are nearly identical. Reading "twice as likely" is approximately right.

Trial B: a common outcome. Control event rate 33%. Odds in control: 0.33 / 0.67 ≈ 0.493 . Doubling gives 0.985 , so p 1 = 0.985 / 1.985 ≈ 49.6 % .

  • Risk difference: about 16.6 percentage points
  • Risk ratio: about 1.50

Here the odds ratio of 2.0 corresponds to a risk ratio of 1.5. Reading "twice as likely" overstates the effect by a third.

Why this matters for decisions. Trial A's treatment moves one person per hundred; Trial B's moves seventeen. Both report O R = 2.0 . A reader given only the odds ratio cannot tell these apart, and the rare-disease case is where the habit of reading odds as risk is formed and then carried into settings where it fails.

What to report. The baseline rate alongside any odds ratio, or better, absolute risks in both arms and their difference.

Worked example

Converting to a decision scale

Problem. A randomized trial of a reminder system on appointment attendance has 800 patients, 400 per arm. Attendance: 312 of 400 treated, 268 of 400 control. A logistic regression of attendance on treatment gives β ^ T = 0.44 , and the report states: "Reminders increased attendance, with an odds ratio of 1.55, patients were 55% more likely to attend."

Goal. Assess the claim and report the effect on a decision scale.

Relevant principle. An odds ratio is a ratio of odds, not of probabilities; with a common outcome the two diverge.

Step 1: compute the proportions.

p ^ 1 = 312 / 400 = 0.78 , p ^ 0 = 268 / 400 = 0.67 .

Step 2: confirm the odds ratio. Odds treated = 0.78 / 0.22 ≈ 3.545 ; odds control = 0.67 / 0.33 ≈ 2.030 .

O R = 3.545 / 2.030 ≈ 1.75 .

Reported as e 0.44 ≈ 1.55 from the adjusted model; the unadjusted figure here is 1.75, and either way it is an odds ratio.

Reason: the coefficient is on the log-odds scale, so exponentiating returns a multiplicative change in odds.

Step 3: identify the error. "55% more likely" reads the odds ratio as a risk ratio. The actual risk ratio is

0.78 / 0.67 ≈ 1.16 ,

a 16% relative increase, not 55%.

Reason: attendance is a common outcome, around 70%, so odds and risks diverge substantially.

Step 4: give the decision-relevant figure. The risk difference is

0.78 − 0.67 = 0.11 ,

11 percentage points. Across 400 patients that is about 44 additional attendances.

Reason: an absolute difference is what capacity planning and cost-per-attendance calculations require.

Step 5: note the simpler route. Because the trial was randomized, p ^ 1 − p ^ 0 estimates the causal risk difference directly. No model was needed to obtain the headline number.

Result. Risk difference 11 percentage points; risk ratio 1.16; odds ratio 1.75. The claim of "55% more likely" overstates the relative effect roughly threefold.

Check. Does the odds ratio exceed the risk ratio, as it should for an outcome above 50%? Yes, 1.75 against 1.16, and the gap is large precisely because attendance is common.

Interpretation. Report the two attendance rates, their difference with an interval, and the implied number of additional attendances. If an odds ratio is reported at all, give the baseline rate beside it so a reader can recover the absolute effect.

Non-example

Statements these models do not support

"An odds ratio of 2 means twice as likely." True only when the outcome is rare. At a 33% baseline it corresponds to a risk ratio of about 1.5.

Reporting a logistic coefficient as a change in probability. β j is on the log-odds scale. Getting to probabilities requires marginal effects, which depend on where in the covariate space they are evaluated.

Comparing logistic coefficients across models with different covariates. The scale itself shifts when covariates are added, even covariates unrelated to treatment, so the coefficients are not comparable in the way OLS coefficients are.

Rejecting the linear probability model solely because it can predict outside [ 0 , 1 ] . When the interest is an average difference in the middle of the range, out-of-range fitted values at the extremes may be irrelevant to the estimand.

Using classical standard errors with the linear probability model. The error variance depends on the fitted probability by construction, so heteroskedasticity is guaranteed, not merely possible.

Confusing an outcome model with a propensity model. Logistic regression appears in this subject for both jobs. One estimates P ( Y = 1 ∣ X ) ; the other estimates P ( W = 1 ∣ X ) and is judged by balance and overlap.

Contrast

Choosing the reporting scale

Linear probability modelLogistic regression
Models P ( Y = 1 ∣ X ) = X β log ⁡ p 1 − p = X β
Coefficient reads asChange in probabilityChange in log-odds; e β j is an odds ratio
Fitted valuesMay leave [ 0 , 1 ] Always in ( 0 , 1 )
Standard errorsRobust, always — heteroskedasticity is structuralModel-based or robust
Directly decision-readyYes, a probability differenceNo, needs marginal effects
Comparable across specificationsYesNo — the scale shifts

Why the odds scale causes so much trouble. Odds are unfamiliar outside betting, and an odds ratio sounds like it should be a ratio of chances. It is a ratio of p / ( 1 − p ) , and that denominator is what makes it diverge from a risk ratio as p grows.

The baseline rate resolves it. Given the control-arm probability, any odds ratio converts to a risk difference and a risk ratio. Reporting an odds ratio without the baseline leaves a reader unable to recover the magnitude.

Which to prefer. If the estimand is an average difference in probability, usually the case in a randomized trial, the linear probability model or plain proportions report it directly. If fitted probabilities near the boundaries matter, or the design calls for it, logistic regression plus marginal effects gets to the same scale by a longer route.

What both share. Neither supplies a causal reading. That comes from the design, exactly as in the continuous-outcome case.

Exercise

1: fully structured. A logistic model gives β ^ T = 0.69 for treatment.

(a) What is the odds ratio? (b) If the control-arm probability is 0.10, what is the treated probability? (c) What is the risk difference?

Check: (a) e 0.69 ≈ 2.0 ; (b) control odds = 0.10 / 0.90 ≈ 0.111 , doubled ≈ 0.222 , so p 1 = 0.222 / 1.222 ≈ 0.182 ; (c) about 8.2 percentage points. Note that "twice the odds" produced a probability rising from 10% to 18%, not to 20%.

2: partly structured. A trial with a 50% control event rate reports O R = 1.8 .

(a) Compute the treated probability. (b) Give the risk ratio. (c) Why does "80% more likely" mislead here?

Check: (a) control odds = 1.0 , treated odds = 1.8 , so p 1 = 1.8 / 2.8 ≈ 0.643 ; (b) 0.643 / 0.50 ≈ 1.29 ; (c) the odds ratio of 1.8 corresponds to a 29% relative increase in probability, so "80% more likely" overstates it nearly threefold. The divergence is greatest when the baseline is near 50%.

3: unstructured. A press release states: "Our randomized trial found the app doubles the odds of quitting smoking (OR = 2.1, p < 0.001). Smokers using the app are twice as likely to quit."

The control-arm quit rate was 22%. Assess the claim and say what should have been reported.

Check: the second sentence converts an odds ratio into a risk ratio, which is not valid at a 22% baseline. Control odds = 0.22 / 0.78 ≈ 0.282 ; multiplied by 2.1 gives 0.592 , so p 1 = 0.592 / 1.592 ≈ 37.2 % . The risk ratio is 0.372 / 0.22 ≈ 1.69 and the risk difference is about 15 percentage points. A substantial and genuinely good result, but "twice as likely" overstates the relative effect and obscures the absolute one. What should have been reported: the quit rates in both arms (22% and 37%), the difference with a confidence interval, and, if the odds ratio is quoted at all, the baseline rate alongside it. Since the trial was randomized, the difference in proportions estimates the causal risk difference directly and needs no model.

What to carry forward

The starting point. For Y ∈ { 0 , 1 } , E [ Y ∣ X ] = P ( Y = 1 ∣ X ) . The conditional mean is the probability.

Linear probability model. P ( Y = 1 ∣ X ) = X β ; coefficients are probability differences; fitted values may leave [ 0 , 1 ] ; errors are structurally heteroskedastic, so robust standard errors always.

Logistic regression. P ( Y = 1 ∣ X ) = ( 1 + e − X β ) − 1 , modelling log-odds linearly; O R = e β j ; fitted probabilities stay in range.

An odds ratio is not a risk ratio. Nor a probability difference. They approximately coincide only when the outcome is rare, and diverge most near a 50% baseline.

Converting requires a baseline. Given the control probability, an odds ratio yields the treated probability, the risk ratio and the risk difference. Report the baseline beside any odds ratio.

Marginal effects. What logistic output needs to reach the probability scale.

In a randomized trial. p ^ 1 − p ^ 0 is the estimated risk difference, obtainable without a model.

Two jobs, one method. Logistic regression models outcomes here and treatment assignment in a propensity score. Different estimands, different criteria.

The recurring error. Reporting an odds ratio as though it were a risk ratio.

Next step

Practice Binary Outcome Models for Experimental Research

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.