Partition of a Sample Space
To understand Bayes' Theorem we first see how a sample space can be broken into useful pieces. Consider a random experiment with sample space S and n events E1,E2,…,En that are:
- Mutually exclusive: Ei∩Ej=∅ for all i=j.
- Exhaustive: E1∪E2∪⋯∪En=S.
Such a set {E1,E2,…,En} is called a partition of S.
A partition divides the sample space into non-overlapping pieces that together account for every outcome — like cutting a cake into slices, each separate, every part belonging to exactly one slice.
For any event A, the partition lets us write A as a union of disjoint pieces:
A=(A∩E1)∪(A∩E2)∪⋯∪(A∩En)
Since the Ei are mutually exclusive, so are the A∩Ei, hence:
P(A)=P(A∩E1)+P(A∩E2)+⋯+P(A∩En)
Using P(A∩Ei)=P(Ei)⋅P(A∣Ei), we obtain the law of total probability:
P(A)=P(E1)P(A∣E1)+P(E2)P(A∣E2)+⋯+P(En)P(A∣En)
This is the backbone of Bayes' Theorem.
The law of total probability requires that the Ei form a partition of S and that P(Ei)>0 for every i.
The Problem of Reverse Probability
Consider the opening example. Bag I has 2 white and 3 red balls; Bag II has 4 white and 5 red. A bag is chosen at random, then a ball is drawn. It is easy to find the probability of drawing white given which bag was chosen. But the real question is the reverse: if the drawn ball is white, what is the probability it came from Bag II?
This is a reverse probability problem — we know the effect (the ball's colour) and want to infer the cause (which bag). Bayes' formula uses conditional probability to reverse the conditioning, from P(effect∣cause) to P(cause∣effect).
Statement of Bayes' Theorem
Let E1,E2,…,En be a partition of S with P(Ei)>0, and let A be any event with P(A)>0. Then for any k=1,2,…,n:
P(Ek∣A)=i=1∑nP(Ei)P(A∣Ei)P(Ek)P(A∣Ek)
The denominator is just P(A) by the law of total probability.
P(Ek∣A)=P(A)P(Ek)P(A∣Ek)
Proof of Bayes' Theorem
›Proof
By the definition of conditional probability, P(Ek∣A)=P(A)P(Ek∩A).
The multiplication rule gives P(Ek∩A)=P(Ek)⋅P(A∣Ek), and the law of total probability gives P(A)=∑i=1nP(Ei)P(A∣Ei). Substituting:
P(Ek∣A)=i=1∑nP(Ei)P(A∣Ei)P(Ek)P(A∣Ek)
Understanding the Terms
- P(Ek) — the prior probability: our initial belief in Ek before observing data.
- P(A∣Ek) — the likelihood: the probability of observing A given Ek.
- P(Ek∣A) — the posterior probability: our updated belief in Ek after observing A.
- P(A) — the evidence: the total probability of A under all causes.
Posterior=EvidencePrior×Likelihood
Worked Example: The Two-Bag Problem
Step 1 — Define events. Let E1 = Bag I chosen, E2 = Bag II chosen, A = a white ball is drawn.
Step 2 — Given probabilities. A bag is chosen at random, so P(E1)=P(E2)=21. Bag I has 2 white of 5, Bag II has 4 white of 9:
P(A∣E1)=52,P(A∣E2)=94
Step 3 — Total probability of white.
P(A)=21⋅52+21⋅94=51+92=459+4510=4519
Step 4 — Apply Bayes' Theorem for P(E2∣A):
P(E2∣A)=P(A)P(E2)P(A∣E2)=451992=92×1945=1910
The answer 1910 differs from the prior 21: observing a white ball has updated our belief, making Bag II more likely because it had a higher proportion of white balls.
Related Results
Property (I): Bayes' Rule for Two Events
When the partition is just E and its complement E′, Bayes' Theorem becomes:
P(E∣A)=P(E)P(A∣E)+P(E′)P(A∣E′)P(E)P(A∣E)
The denominator is the two-event law of total probability P(A)=P(E)P(A∣E)+P(E′)P(A∣E′).
Property (II): Posterior Odds Ratio
For any two events Ei, Ej from the partition:
P(Ej∣A)P(Ei∣A)=P(Ej)P(Ei)×P(A∣Ej)P(A∣Ei)
so the posterior odds equal the prior odds times the likelihood ratio.
›Proof
Dividing P(Ei∣A)=P(A)P(Ei)P(A∣Ei) by P(Ej∣A)=P(A)P(Ej)P(A∣Ej), the common P(A) cancels, giving the result. …