Bayes' Theorem and the Mathematics of AI
On June 1, 2009, Air France Flight 447 disappeared over the Atlantic Ocean. Despite extensive searches, the main wreckage wasn't found for nearly two years.
In 2011, a search analytics firm called Metron used Bayesian search theory to combine everything known, including the flight's last position, drift data and, crucially, how likely each earlier search was to have missed the wreck. Their probability map pointed to an area close to the last known position. Within about a week of searching there in April 2011, the wreckage was found.
The engine of that analysis was a formula written by an 18th-century English minister. Today, the same formula shapes how AI systems reason under uncertainty.
A 99% Accurate Test Can Be Wrong Most of the Time
A disease affects 1 in 1,000 people. A test detects 99% of real cases and wrongly flags only 1% of healthy people. You test positive. What's the chance you actually have the disease?
Most people guess around 99%. The real answer is about 9%.
Imagine 100,000 people:
- 100 have the disease; the test catches 99
- 99,900 are healthy; the test wrongly flags 999
Out of 1,098 positives, only 99 are real: 99 / 1,098 ≈ 9%. The disease is so rare that false alarms from the huge healthy group swamp the true cases.
The Formula
Bayes' theorem updates a belief in hypothesis H after seeing evidence E:
P(H | E) = P(E | H) × P(H) / P(E)
- P(H): the prior, what you believed before the evidence (0.001)
- P(E | H): the likelihood, how probable the evidence is if H is true (0.99)
- P(E): the total probability of the evidence
- P(H | E): the posterior, your updated belief
For the test:
P(E) = 0.99 × 0.001 + 0.01 × 0.999 = 0.01098
P(disease | positive) = 0.00099 / 0.01098 ≈ 0.090
See the rules on the statistics formulas page.
Updating Again
Bayes' theorem is built for repeated updating. Today's posterior becomes tomorrow's prior. If you take a second, independent test and it's also positive:
Prior = 0.090
P(disease | two positives) ≈ 0.99 × 0.090 / (0.99 × 0.090 + 0.01 × 0.910) ≈ 0.907
One positive: 9%. Two positives: 91%. Evidence accumulates.
The Odds Form: Easy Mental Math
Bayes' theorem is simplest in odds:
Posterior odds = Prior odds × Likelihood ratio
Prior odds of disease are 1 : 999. The likelihood ratio of a positive test is 0.99 / 0.01 = 99. So posterior odds are 99 : 999, about 1 : 10, which is about 9%. Each independent positive result multiplies the odds by 99 again.
An Insider Reference: Bayes, Price and Turing
Thomas Bayes, a Presbyterian minister, never published his result. After his death, his friend Richard Price edited and presented "An Essay towards solving a Problem in the Doctrine of Chances" to the Royal Society in 1763. Around 1774, Pierre-Simon Laplace independently developed the idea much further and applied it widely.
During World War II, Alan Turing and his colleagues at Bletchley Park used Bayesian reasoning to break German naval Enigma messages. In a procedure called Banburismus, they accumulated evidence for possible settings in units Turing called bans, with a tenth of a ban called a deciban, which is essentially the logarithm of the likelihood ratio. Using logarithms turned multiplying odds into adding scores. A likelihood ratio of 99 is about 20 decibans; try it with the logarithm calculator.
Sharon Bertsch McGrayne tells this and many other stories in her 2011 book The Theory That Would Not Die.
Bayes in AI
Spam Filters
Naive Bayes classifiers estimate P(spam | words) by multiplying the likelihoods of each word, assuming words are independent. The assumption is wrong, but the method works surprisingly well and was central to early spam filtering.
Priors as Regularization
When training a model, placing a Gaussian prior on the weights (believing small weights are more likely) is mathematically equivalent to adding L2 regularization. Many standard ML techniques are Bayesian reasoning in disguise.
Bayesian Optimization
Tuning a model's settings (learning rate, layer sizes) is expensive. Bayesian optimization builds a probabilistic model of which settings are likely to perform well and chooses the most promising experiment next.
Uncertainty
Bayesian neural networks and related methods estimate not just a prediction but how uncertain it is, which matters for medicine, self-driving cars and anywhere a confident mistake is costly.
Two Concepts Worth Knowing
Base Rate
The base rate is how common something is before any evidence, the prior. Ignoring it, the base rate fallacy, is why the 99%-accurate test seems more reliable than it is.
Likelihood Ratio
The likelihood ratio P(E | H) / P(E | not H) measures how strongly evidence favors a hypothesis. A ratio of 1 means the evidence tells you nothing.
Quick Answer: What Is Bayes' Theorem?
Bayes' theorem calculates how to update a probability given new evidence: P(H | E) = P(E | H) × P(H) / P(E). It combines a prior belief with the likelihood of the evidence. In AI, it underpins naive Bayes classifiers, regularization, Bayesian optimization and uncertainty estimates.
Try Them Yourself
- Statistics Formulas: conditional probability and Bayes' theorem
- Logarithm Calculator: Turing's decibans
- Division Tables: converting between odds and probabilities
- Random Number Generator: simulate a population and a test
- How Probability Powers Artificial Intelligence: probability across AI systems
- The Monty Hall Problem Explained: Bayes' theorem in a game show
Redo the disease example with a condition that affects 1 in 100 people instead of 1 in 1,000. Before calculating, guess the answer. The change is bigger than most people expect.