Conditional Probability, Total Probability and Bayes' Rule

Conditional probability as a change of denominator rather than a change to the world, the asymmetry between P ( A ∣ B ) and P ( B ∣ A ) that follows from it, and the decomposition of a probability over a partition that turns into Bayes' rule and makes the base rate govern a posterior as much as the likelihood does.

Definition

Conditional probability. For P ( B ) > 0 ,

P ( A ∣ B ) = P ( A ∩ B ) P ( B ) ,

the probability of A once the outcome is known to lie in B . Rearranged, P ( A ∩ B ) = P ( A ∣ B ) P ( B ) .

Conditioning discards the outcomes where B failed, so B becomes the new sample space. The surviving probability of A is P ( A ∩ B ) , and dividing by P ( B ) restores the total to 1. Without that renormalisation the conditional probabilities would not sum to 1 over a partition.

Two consequences follow immediately. The operation is not symmetric: P ( A ∣ B ) divides by P ( B ) and P ( B ∣ A ) by P ( A ) , and these differ unless the two marginals happen to be equal. And P ( B ) > 0 is required, since conditioning on an impossible event leaves nothing to renormalise.

Independence, stated conditionally. A and B are independent when P ( A ∣ B ) = P ( A ) : conditioning on B changes nothing. This is equivalent to the factorisation P ( A ∩ B ) = P ( A ) P ( B ) , and its consequences for variances are developed in the unit on covariance and independence.

Law of total probability and Bayes. For a partition B 1 , … , B k of Ω ,

P ( A ) = ∑ i P ( A ∣ B i ) P ( B i ) , P ( B j ∣ A ) = P ( A ∣ B j ) P ( B j ) P ( A ) .

The first decomposes an overall probability into the cases of a partition, weighted by how likely each case is. The second reverses the direction of conditioning, and its denominator is the first. The corresponding decomposition for expectation is E [ X ] = E [ E [ X ∣ Y ] ] .

Assumptions and scope

  • P ( A ∣ B ) requires P ( B ) > 0 , and is not symmetric: P ( A ∣ B ) ≠ P ( B ∣ A ) except by coincidence.

  • A small conditional probability of evidence given a hypothesis does not make the hypothesis improbable; the base rate P ( B j ) governs the posterior as much as the likelihood does.

  • The law of total probability requires the B i to partition the sample space: mutually exclusive, and jointly exhaustive. A set of cases that overlaps, or that omits a possibility, gives the wrong denominator.

  • A posterior is specific to the population whose base rate it used. The same likelihoods applied to a different population give a different answer, so the base rate belongs with the reported figure.

Forms this is expressed in

The same content in several forms. Each makes something visible that the others leave implicit, so moving between them is part of understanding the topic rather than a presentation choice.

tabular

A two-way grid of counts, one row per value of the first variable and one column per value of the second. Interior cell ( i , j ) carries the number of cases with both values; the right margin carries the row totals and the bottom margin the column totals, with the grand total in the corner. Dividing the interior by the grand total gives the joint distribution, the margins give the two marginal distributions, and dividing a cell by its own row or column total gives a conditional.

This reading answers questions about which set is being divided by. Conditioning becomes a physical operation, cover the rows that did not occur, then renormalise what is left, rather than a formula to recall. Because P ( A ∣ B ) divides by a row total and P ( B ∣ A ) by a column total, the asymmetry of conditioning is visible as two fractions sharing a numerator: the same 45 students over 60 in one direction and over 55 in the other.

It is also the form that makes base rates impossible to overlook. A rare condition is a short row, so the false positives drawn from the long row can outnumber the true positives even when each is individually unlikely, and the comparison is a glance at two cells rather than an inference from percentages. Counting a concrete population, 10,000 people rather than probabilities, turns Bayes' rule into one count divided by a total of counts.

Independence is equally direct: the variables are independent exactly when every interior cell equals its row total times its column total divided by the grand total, so a single cell disagreeing with that product settles the question.

What the table does not expose is anything beyond two discrete variables with few values. Continuous quantities have no cells, a third variable has no axis, and the layout silently suggests that every probability is a tally waiting to be counted, which is false for a genuinely unrepeatable event, where the counts that would fill the grid do not exist.

Worked material

Example

Conditioning on a second table, and a partition with three cases

The student table is one shape of problem. Two more, to separate the method from the example.

---

1. A partition with three cases. A factory takes components from three suppliers. Supplier A provides 50% of them and 2% of those are defective; B provides 30% at 4% defective; C provides 20% at 5% defective.

The overall defect rate, by the law of total probability over the partition { A , B , C } :

P ( def ) = ( 0.02 ) ( 0.50 ) + ( 0.04 ) ( 0.30 ) + ( 0.05 ) ( 0.20 ) = 0.010 + 0.012 + 0.010 = 0.032 .

The three cases are mutually exclusive and cover every component, which is what licenses the sum.

Reversing the conditioning. A defective component is found. Which supplier is it most likely from?

P ( A ∣ def ) = 0.010 0.032 = 0.3125 , P ( B ∣ def ) = 0.012 0.032 = 0.375 , P ( C ∣ def ) = 0.010 0.032 = 0.3125 ,

summing to 1. B is the most likely source despite not having the worst defect rate, because it supplies half again as many components as C. And A, with the best rate, is exactly as likely a source as C with the worst, because it ships two and a half times the volume. Neither the rate nor the volume decides this alone; the product does.

---

2. A table with a different shape. 200 job applicants, classified by whether they were referred by an employee and whether they were hired.

HiredNot hiredRow total
Referred243660
Not referred28112140
Column total52148200
P ( hired ∣ referred ) = 24 60 = 0.40 , P ( hired ∣ not referred ) = 28 140 = 0.20 .

Referral doubles the hire rate. But reversing the conditioning asks a different question:

P ( referred ∣ hired ) = 24 52 ≈ 0.4615 ,

so most hires were not referred, even though referral doubled an individual's chances. Both statements are true of the same table. The first divides by a row, the second by a column, and they answer different questions: one is about an applicant's prospects, the other about the composition of the hired group.

Dependence check. P ( referred ) P ( hired ) = 0.30 × 0.26 = 0.078 against the actual joint 24 200 = 0.12 . Not equal, so the two are dependent.

---

What carries across both. The arithmetic never changed: identify the conditioning event, restrict to it, renormalise. What changed was which quantity the question wanted. In the supplier case the tempting error is to answer with the defect rates; in the hiring case it is to read a doubled individual rate as a statement about who the hires are. Both errors are the same error, answering the conditional that was easiest to see rather than the one that was asked.

Common errors

Common misconception

P ( A ∣ B ) and P ( B ∣ A ) describe the same relationship, so a test that detects a condition 99% of the time makes a positive result 99% likely to be correct.

Related units

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.