Bayes' Theorem: Learning from Evidence
Imagine you have a bag with 3 red marbles and 7 blue marbles. If you pick one at random, the chance it's red is 3 out of 10 — that's straightforward. But now suppose someone picks a marble, doesn't show it to you, but tells you it's not blue. Suddenly, the only possibilities left are the red marbles. The probability that the hidden marble is red jumps to 1 (certainty). You just updated your belief based on new evidence.
That's the core idea of Bayes' Theorem: how to revise a probability when you get new information. It answers the question: Given that I now know B happened, how should I change my belief about A?
The Intuition in One Sentence
Bayes' Theorem says: The probability that A is true, given that B is true, equals the probability that B would happen if A were true, times the original probability of A, divided by the overall probability of B.
In other words: Your updated belief = (likelihood of the evidence under your hypothesis) × (your prior belief) / (total probability of the evidence).
The Precise Statement
Let A and B be two events. Then:
P(A∣B)=P(B)P(B∣A)⋅P(A)
Where:
- P(A∣B) is the posterior probability — what you want: the probability of A given that B occurred.
- P(B∣A) is the likelihood — how probable the evidence B is if A is true.
- P(A) is the prior probability — your initial belief about A before seeing any evidence.
- P(B) is the marginal probability of B — the total chance that B happens, regardless of A.
P(A∣B)=P(B)P(B∣A)⋅P(A)
Why It Works: A Simple Example
Suppose 1% of a population has a disease. A test for the disease is 99% accurate: it correctly identifies 99% of those who have it (true positive) and correctly says 99% of those who don't have it are negative (true negative). You take the test and get a positive result. What is the probability you actually have the disease?
Many people guess 99%. But that's wrong — and Bayes' Theorem shows why.
Let D = "has the disease", T+ = "tests positive". We know:
- P(D)=0.01 (prior)
- P(T+∣D)=0.99 (likelihood)
- P(T+∣not D)=0.01 (false positive rate)
First, find P(T+), the total probability of a positive test:
P(T+)=P(T+∣D)P(D)+P(T+∣not D)P(not D)
=(0.99)(0.01)+(0.01)(0.99)=0.0099+0.0099=0.0198
Now apply Bayes:
P(D∣T+)=0.01980.99×0.01=0.01980.0099=0.5
So even with a positive test, there's only a 50% chance you have the disease. The test is good, but the disease is rare — most positive results come from the large number of healthy people who get false positives.
A common mistake is to confuse P(B∣A) with P(A∣B). In the disease example, P(positive∣disease)=0.99, but P(disease∣positive)=0.5. They are not the same.
The General Form (with Multiple Hypotheses)
Often you have several possible causes A1,A2,…,An that partition the sample space. Then for any one cause Ai:
P(Ai∣B)=∑j=1nP(B∣Aj)⋅P(Aj)P(B∣Ai)⋅P(Ai)
The denominator is just the law of total probability applied to B.
Why It Matters
Bayes' Theorem is the mathematical foundation of learning from data. It's used everywhere: spam filters update the probability an email is spam based on words it contains; doctors update the probability of a disease based on test results; machine learning algorithms update model parameters as new data arrives. Every time you change your mind because of new evidence, you're doing Bayesian reasoning — whether you know it or not.
The key takeaway: Bayes' Theorem is not a mysterious formula — it's just common sense made precise. It tells you how to weigh new evidence against your prior knowledge, and it prevents you from being fooled by rare events or misleading tests.