Conditional Probability, Total Probability and Bayes' Rule
What you will be able to do
The learner can compute a conditional probability from a joint table or mass function, apply the law of total probability, and obtain a posterior by Bayes' rule without inverting the conditioning or discarding the base rate.
Orientation
Conditioning changes the denominator
A conditional probability is not a new kind of probability. It is the same probability measured against a smaller reference set.
Evidence does not change the world. It changes which outcomes are still in play, and therefore what the total is being divided by.
Two consequences follow, and both are responsible for a large share of misread statistics.
The operation is not symmetric.
The base rate does not go away. A test that is 99% sensitive and 95% specific, applied to a disease affecting 1% of people, returns a positive result that means disease with probability
That single calculation is why this unit exists. Conditioning correctly is what separates a defensible inference from a plausible-sounding one, and the machinery, total probability and Bayes' rule, is just bookkeeping for getting the denominator right.
Definition
Three places the definition is easy to misapply
The formula is one line. Nearly every error with it comes from one of three places, none of which the formula itself flags.
1. The conditioning event is not what the sentence says it is. "Given that at least one of two coins came up heads, what is the probability both did?" The conditioning event is
2. The partition does not partition. The law of total probability requires cases that are mutually exclusive and jointly exhaustive. Cases drawn from a survey, "reads newspapers", "reads news online", "follows news on social media", overlap freely, and summing
3. Conditioning on an event of probability zero.
A fourth, subtler one.
Intuition
Why the base-rate error survives being explained
People who can state that
Percentages hide the denominators. "99% sensitive" and "1% prevalence" are both percentages, and nothing in the surface form of either says which population it was taken out of. Held side by side they look like two facts of the same kind that should combine simply. They are not: one is a rate within a small group, the other is the size of that group relative to everything. The arithmetic that combines them has to reinstate the two different denominators the percentages discarded.
Counts do not hide them. Restated as counts over 10,000 people, the same facts are 100 diseased and 9,900 healthy, with 99 and 495 positives respectively. No one asked to pick a positive test out of that population reports 99%; the 594 is sitting there to be divided into. The information content is identical. The format is what changed, and with it the error rate. That is why the procedure for Bayes' rule says to count a concrete population first and reach for the formula second.
The vividness of the mechanism competes with the size of the group. A test that detects disease is a causal story: disease present, test responds. A base rate is not a story about anything, just a count. When the two conflict the story usually wins, which is the same failure that makes people fear rare vivid risks over common dull ones. Noticing that a number has no narrative attached is not a reason to discount it.
A diagnostic question. Whenever a conditional probability is quoted, ask: out of whom? If the answer is not immediately available from how the figure was stated, the figure is not yet interpretable, and the direction of conditioning is the first thing to check rather than the last.
Representation
Joint, marginal and conditional in one table
A two-way table of counts holds the joint distribution, both marginals and every conditional at once. Reading it well makes conditioning a visible operation rather than a formula.
The data. 100 students, classified by whether they studied and whether they passed.
| Passed | Failed | Row total | |
|---|---|---|---|
| Studied | 45 | 15 | 60 |
| Did not | 10 | 30 | 40 |
| Column total | 55 | 45 | 100 |
Joint probabilities: divide by the grand total.
The four cells give the joint distribution and sum to 1.
Marginals: use the totals.
The word marginal is literal: the numbers live in the margins.
Conditionals: divide within a row or column. To condition on studying, discard the other row entirely and renormalise what remains:
The denominator is the row total, not the grand total.
The asymmetry, visible. Conditioning the other way divides by a column instead:
Same numerator, 45. Different denominator, 60 against 55. So
Testing independence. If studying told us nothing about passing, the joint would factor:
against the actual joint
Total probability, read off the rows. The overall pass rate decomposes over the partition:
matching the column total exactly. That is the law of total probability, and in table form it is just "add the two rows back together".
It holds two discrete variables with few values. Continuous variables, three-way relationships and genuinely unrepeatable events need other machinery, and a table of counts can wrongly suggest every probability is estimable by tallying, which fails precisely where the interesting questions live.
Example
Conditioning on a second table, and a partition with three cases
The student table is one shape of problem. Two more, to separate the method from the example.
---
1. A partition with three cases. A factory takes components from three suppliers. Supplier A provides 50% of them and 2% of those are defective; B provides 30% at 4% defective; C provides 20% at 5% defective.
The overall defect rate, by the law of total probability over the partition
The three cases are mutually exclusive and cover every component, which is what licenses the sum.
Reversing the conditioning. A defective component is found. Which supplier is it most likely from?
summing to 1. B is the most likely source despite not having the worst defect rate, because it supplies half again as many components as C. And A, with the best rate, is exactly as likely a source as C with the worst, because it ships two and a half times the volume. Neither the rate nor the volume decides this alone; the product does.
---
2. A table with a different shape. 200 job applicants, classified by whether they were referred by an employee and whether they were hired.
| Hired | Not hired | Row total | |
|---|---|---|---|
| Referred | 24 | 36 | 60 |
| Not referred | 28 | 112 | 140 |
| Column total | 52 | 148 | 200 |
Referral doubles the hire rate. But reversing the conditioning asks a different question:
so most hires were not referred, even though referral doubled an individual's chances. Both statements are true of the same table. The first divides by a row, the second by a column, and they answer different questions: one is about an applicant's prospects, the other about the composition of the hired group.
Dependence check.
---
What carries across both. The arithmetic never changed: identify the conditioning event, restrict to it, renormalise. What changed was which quantity the question wanted. In the supplier case the tempting error is to answer with the defect rates; in the hiring case it is to read a doubled individual rate as a statement about who the hires are. Both errors are the same error, answering the conditional that was easiest to see rather than the one that was asked.
Procedure
Computing a conditional probability and applying Bayes’ rule
To compute a conditional probability.
Step 1 — Identify the conditioning event and treat it as the new sample space.
Step 2 — Divide the joint by the conditioning event's probability:
Step 3 — Check which direction was asked.
To apply Bayes' rule.
Step 4 — Write down the base rate first, before any likelihood. It is the quantity the intuition most often discards.
Step 5 — Prefer natural frequencies to the formula. Take a concrete population of 10,000 people and count: how many have the condition, how many of those test positive, how many of the rest test positive anyway. The posterior is then one count divided by a total of counts, with no formula to misremember.
Step 6 — Report the posterior with its base rate attached, since the same test gives a very different answer in a screening population than in a symptomatic one.
Where it goes wrong.
- Inverting the conditioning, reading
as . - Dividing by the grand total, which gives the joint probability rather than the conditional.
- Ignoring the base rate, which turns a 99% accurate test into a 99% confident diagnosis when the truth is about 17%.
- Using cases that overlap or omit a possibility, so the partition is not one and the denominator is wrong.
- Reporting a posterior with no base rate attached, leaving it uninterpretable in any other population.
Worked example
Posterior probability after a positive test, by two methods
Problem. A disease affects 1% of a population. A test detects it in 99% of those who have it (sensitivity) and correctly clears 95% of those who do not (specificity). Someone tests positive. What is the probability they have the disease?
Step 1 — Record the base rate and the conditional probabilities.
The question asks for
Step 2 — Compute the posterior using natural frequencies. Take 10,000 people.
| Test + | Test − | Total | |
|---|---|---|---|
| Disease | 99 | 1 | 100 |
| No disease | 495 | 9,405 | 9,900 |
| Total | 594 | 9,406 | 10,000 |
1% of 10,000 is 100 with the disease, and 99% of those test positive, giving 99. The remaining 9,900 are healthy, and 5% of them test positive, giving 495.
Conditioning on a positive result restricts attention to the first column, which holds 594 people:
Step 3 — Verify with Bayes' rule. The law of total probability gives the denominator:
In exact arithmetic
Step 4 — Interpret the posterior against the base rate. About 17% of those testing positive have the disease. There are 99 times as many healthy people as diseased, so a 5% false-positive rate among them produces 495 false positives against 99 true positives.
The positive result raised the probability from 1% to about 17%, a factor of about 17. It did not raise it to 99%, because 99% is
The base rate changes the answer. The same test in a population where the disease affects 50% gives
about 95%. The test and the likelihoods are unchanged; only the base rate differs. A posterior should therefore be reported together with the base rate it assumes, as the procedure for applying Bayes' rule requires.
Arithmetic check.