Conditional Probability, Total Probability and Bayes' Rule

What you will be able to do

The learner can compute a conditional probability from a joint table or mass function, apply the law of total probability, and obtain a posterior by Bayes' rule without inverting the conditioning or discarding the base rate.

Orientation

Conditioning changes the denominator

A conditional probability is not a new kind of probability. It is the same probability measured against a smaller reference set.

P ( A ∣ B ) = P ( A ∩ B ) P ( B )

Evidence does not change the world. It changes which outcomes are still in play, and therefore what the total is being divided by.

Two consequences follow, and both are responsible for a large share of misread statistics.

The operation is not symmetric. P ( A ∣ B ) and P ( B ∣ A ) share a numerator and differ in denominator, so they are different numbers with different meanings. A test that detects a condition in 99% of those who have it does not make a positive result 99% likely to be right; those are opposite conditionings.

The base rate does not go away. A test that is 99% sensitive and 95% specific, applied to a disease affecting 1% of people, returns a positive result that means disease with probability 1 6 , about 17%. Nothing is wrong with the test. Per 10,000 people, 99 true positives arrive alongside 495 false ones, because there are 99 times more healthy people for the false positives to come from.

That single calculation is why this unit exists. Conditioning correctly is what separates a defensible inference from a plausible-sounding one, and the machinery, total probability and Bayes' rule, is just bookkeeping for getting the denominator right.

Definition

Three places the definition is easy to misapply

The formula is one line. Nearly every error with it comes from one of three places, none of which the formula itself flags.

1. The conditioning event is not what the sentence says it is. "Given that at least one of two coins came up heads, what is the probability both did?" The conditioning event is { H H , H T , T H } , three outcomes of four, so the answer is 1 3 . Read instead as "one specific coin came up heads", the event is { H H , H T } and the answer is 1 2 . Same words in ordinary English, different subsets, different answers. Before dividing, write the conditioning event out as a set of outcomes and count it.

2. The partition does not partition. The law of total probability requires cases that are mutually exclusive and jointly exhaustive. Cases drawn from a survey, "reads newspapers", "reads news online", "follows news on social media", overlap freely, and summing P ( A ∣ B i ) P ( B i ) over them double-counts anyone in two categories. A set of cases that omits a possibility fails the other way: the weights sum to less than 1 and the total is too small. Checking that the P ( B i ) sum to exactly 1 catches both failures cheaply.

3. Conditioning on an event of probability zero. P ( A ∣ B ) is undefined when P ( B ) = 0 , and the definition offers no way around it, since the division has no value. This is not a pedantic corner. For a continuous variable every single value has probability zero, so "given that the measurement was exactly 3.0" cannot be handled by this definition at all; what is meant is an interval, or a conditional density, which is a different construction requiring a limit. Treating a continuous point condition as if the discrete formula applied is how an argument silently acquires an undefined quantity.

A fourth, subtler one. P ( A ∣ B ) describes the population the probabilities came from. Conditioning does not license a claim about what would happen if B were made to occur. P ( recovery ∣ took the drug ) is computed from people who chose, or were chosen, to take it, and those people may differ from everyone else in ways that also affect recovery. Conditioning and intervening are different operations, and only a design argument connects them.

Intuition

Why the base-rate error survives being explained

People who can state that P ( A ∣ B ) ≠ P ( B ∣ A ) still make the error under time pressure, and the reason is worth knowing, because it suggests what to do about it.

Percentages hide the denominators. "99% sensitive" and "1% prevalence" are both percentages, and nothing in the surface form of either says which population it was taken out of. Held side by side they look like two facts of the same kind that should combine simply. They are not: one is a rate within a small group, the other is the size of that group relative to everything. The arithmetic that combines them has to reinstate the two different denominators the percentages discarded.

Counts do not hide them. Restated as counts over 10,000 people, the same facts are 100 diseased and 9,900 healthy, with 99 and 495 positives respectively. No one asked to pick a positive test out of that population reports 99%; the 594 is sitting there to be divided into. The information content is identical. The format is what changed, and with it the error rate. That is why the procedure for Bayes' rule says to count a concrete population first and reach for the formula second.

The vividness of the mechanism competes with the size of the group. A test that detects disease is a causal story: disease present, test responds. A base rate is not a story about anything, just a count. When the two conflict the story usually wins, which is the same failure that makes people fear rare vivid risks over common dull ones. Noticing that a number has no narrative attached is not a reason to discount it.

A diagnostic question. Whenever a conditional probability is quoted, ask: out of whom? If the answer is not immediately available from how the figure was stated, the figure is not yet interpretable, and the direction of conditioning is the first thing to check rather than the last.

Representation

Joint, marginal and conditional in one table

A two-way table of counts holds the joint distribution, both marginals and every conditional at once. Reading it well makes conditioning a visible operation rather than a formula.

The data. 100 students, classified by whether they studied and whether they passed.

PassedFailedRow total
Studied451560
Did not103040
Column total5545100

Joint probabilities: divide by the grand total.

P ( studied and passed ) = 45 100 = 0.45 .

The four cells give the joint distribution and sum to 1.

Marginals: use the totals.

P ( studied ) = 60 100 = 0.60 , P ( passed ) = 55 100 = 0.55 .

The word marginal is literal: the numbers live in the margins.

Conditionals: divide within a row or column. To condition on studying, discard the other row entirely and renormalise what remains:

P ( passed ∣ studied ) = 45 60 = 0.75 , P ( passed ∣ did not ) = 10 40 = 0.25 .

The denominator is the row total, not the grand total.

The asymmetry, visible. Conditioning the other way divides by a column instead:

P ( studied ∣ passed ) = 45 55 = 0.818182 .

Same numerator, 45. Different denominator, 60 against 55. So P ( passed ∣ studied ) = 0.75 and P ( studied ∣ passed ) = 0.818182 , and reading one as the other is the commonest error conditioning invites.

Testing independence. If studying told us nothing about passing, the joint would factor:

P ( studied ) P ( passed ) = 0.60 × 0.55 = 0.33 ,

against the actual joint 0.45 . They differ, so the variables are dependent, as the conditionals already showed with 0.75 against 0.25 . Under independence all four cells would match their row-times-column products; the gap between 0.33 and 0.45 is the dependence.

Total probability, read off the rows. The overall pass rate decomposes over the partition:

P ( passed ) = 0.75 × 0.60 + 0.25 × 0.40 = 0.45 + 0.10 = 0.55   ✓

matching the column total exactly. That is the law of total probability, and in table form it is just "add the two rows back together".

It holds two discrete variables with few values. Continuous variables, three-way relationships and genuinely unrepeatable events need other machinery, and a table of counts can wrongly suggest every probability is estimable by tallying, which fails precisely where the interesting questions live.

Example

Conditioning on a second table, and a partition with three cases

The student table is one shape of problem. Two more, to separate the method from the example.

---

1. A partition with three cases. A factory takes components from three suppliers. Supplier A provides 50% of them and 2% of those are defective; B provides 30% at 4% defective; C provides 20% at 5% defective.

The overall defect rate, by the law of total probability over the partition { A , B , C } :

P ( def ) = ( 0.02 ) ( 0.50 ) + ( 0.04 ) ( 0.30 ) + ( 0.05 ) ( 0.20 ) = 0.010 + 0.012 + 0.010 = 0.032 .

The three cases are mutually exclusive and cover every component, which is what licenses the sum.

Reversing the conditioning. A defective component is found. Which supplier is it most likely from?

P ( A ∣ def ) = 0.010 0.032 = 0.3125 , P ( B ∣ def ) = 0.012 0.032 = 0.375 , P ( C ∣ def ) = 0.010 0.032 = 0.3125 ,

summing to 1. B is the most likely source despite not having the worst defect rate, because it supplies half again as many components as C. And A, with the best rate, is exactly as likely a source as C with the worst, because it ships two and a half times the volume. Neither the rate nor the volume decides this alone; the product does.

---

2. A table with a different shape. 200 job applicants, classified by whether they were referred by an employee and whether they were hired.

HiredNot hiredRow total
Referred243660
Not referred28112140
Column total52148200
P ( hired ∣ referred ) = 24 60 = 0.40 , P ( hired ∣ not referred ) = 28 140 = 0.20 .

Referral doubles the hire rate. But reversing the conditioning asks a different question:

P ( referred ∣ hired ) = 24 52 ≈ 0.4615 ,

so most hires were not referred, even though referral doubled an individual's chances. Both statements are true of the same table. The first divides by a row, the second by a column, and they answer different questions: one is about an applicant's prospects, the other about the composition of the hired group.

Dependence check. P ( referred ) P ( hired ) = 0.30 × 0.26 = 0.078 against the actual joint 24 200 = 0.12 . Not equal, so the two are dependent.

---

What carries across both. The arithmetic never changed: identify the conditioning event, restrict to it, renormalise. What changed was which quantity the question wanted. In the supplier case the tempting error is to answer with the defect rates; in the hiring case it is to read a doubled individual rate as a statement about who the hires are. Both errors are the same error, answering the conditional that was easiest to see rather than the one that was asked.

Procedure

Computing a conditional probability and applying Bayes’ rule

To compute a conditional probability.

Step 1 — Identify the conditioning event and treat it as the new sample space.

Step 2 — Divide the joint by the conditioning event's probability: P ( A ∣ B ) = P ( A ∩ B ) / P ( B ) . From a table, divide within the row or column for B rather than by the grand total.

Step 3 — Check which direction was asked. P ( A ∣ B ) and P ( B ∣ A ) share a numerator and differ in denominator. Reading the question twice costs less than the error.

To apply Bayes' rule.

Step 4 — Write down the base rate first, before any likelihood. It is the quantity the intuition most often discards.

Step 5 — Prefer natural frequencies to the formula. Take a concrete population of 10,000 people and count: how many have the condition, how many of those test positive, how many of the rest test positive anyway. The posterior is then one count divided by a total of counts, with no formula to misremember.

Step 6 — Report the posterior with its base rate attached, since the same test gives a very different answer in a screening population than in a symptomatic one.

Where it goes wrong.

  • Inverting the conditioning, reading P ( positive ∣ disease ) as P ( disease ∣ positive ) .
  • Dividing by the grand total, which gives the joint probability rather than the conditional.
  • Ignoring the base rate, which turns a 99% accurate test into a 99% confident diagnosis when the truth is about 17%.
  • Using cases that overlap or omit a possibility, so the partition is not one and the denominator is wrong.
  • Reporting a posterior with no base rate attached, leaving it uninterpretable in any other population.

Worked example

Posterior probability after a positive test, by two methods

Problem. A disease affects 1% of a population. A test detects it in 99% of those who have it (sensitivity) and correctly clears 95% of those who do not (specificity). Someone tests positive. What is the probability they have the disease?

Step 1 — Record the base rate and the conditional probabilities. P ( D ) = 0.01 , so P ( not  D ) = 0.99 . The likelihoods are P ( + ∣ D ) = 0.99 and P ( + ∣ not  D ) = 1 − 0.95 = 0.05 .

The question asks for P ( D ∣ + ) . The 99% figure is P ( + ∣ D ) , which conditions in the opposite direction.

Step 2 — Compute the posterior using natural frequencies. Take 10,000 people.

Test +Test −Total
Disease991100
No disease4959,4059,900
Total5949,40610,000

1% of 10,000 is 100 with the disease, and 99% of those test positive, giving 99. The remaining 9,900 are healthy, and 5% of them test positive, giving 495.

Conditioning on a positive result restricts attention to the first column, which holds 594 people:

P ( D ∣ + ) = 99 99 + 495 = 99 594 = 1 6 ≈ 0.166667 .

Step 3 — Verify with Bayes' rule. The law of total probability gives the denominator:

P ( + ) = P ( + ∣ D ) P ( D ) + P ( + ∣ not  D ) P ( not  D ) = ( 0.99 ) ( 0.01 ) + ( 0.05 ) ( 0.99 ) = 0.0594 .
P ( D ∣ + ) = P ( + ∣ D ) P ( D ) P ( + ) = 0.0099 0.0594 = 1 6 .

In exact arithmetic P ( + ) = 297 5000 and P ( D ∣ + ) = 1 6 . The two methods perform the same computation; the table displays the denominator as a count.

Step 4 — Interpret the posterior against the base rate. About 17% of those testing positive have the disease. There are 99 times as many healthy people as diseased, so a 5% false-positive rate among them produces 495 false positives against 99 true positives.

The positive result raised the probability from 1% to about 17%, a factor of about 17. It did not raise it to 99%, because 99% is P ( + ∣ D ) rather than P ( D ∣ + ) ; the two have different denominators.

The base rate changes the answer. The same test in a population where the disease affects 50% gives

P ( D ∣ + ) = ( 0.99 ) ( 0.5 ) ( 0.99 ) ( 0.5 ) + ( 0.05 ) ( 0.5 ) = 0.495 0.520 ≈ 0.951923 ,

about 95%. The test and the likelihoods are unchanged; only the base rate differs. A posterior should therefore be reported together with the base rate it assumes, as the procedure for applying Bayes' rule requires.

Arithmetic check. 99 + 495 = 594 , and 594 = 6 × 99 , so 99 594 = 1 6 . The rows sum to 100 + 9,900 = 10,000 and the columns to 594 + 9,406 = 10,000 .

Next step

Practice Conditional Probability, Total Probability and Bayes' Rule

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.