What Does a Probability Actually Mean?
A weather forecast says there's a 30% chance of rain tomorrow. What does that mean?
In a 2005 study published in Risk Analysis, Gerd Gigerenzer and colleagues asked people in five cities, Amsterdam, Athens, Berlin, Milan and New York, exactly that. Answers varied widely. Some thought it meant rain would fall over 30% of the area. Others thought it would rain 30% of the time. Others believed that 30% of forecasters expected rain. Only in New York did a majority give the standard meteorological interpretation: on days like tomorrow, it rains 30% of the time.
It's a simple number, yet people read it completely differently. Mathematicians and philosophers have debated what probability really means for more than 300 years, and the answer matters every time an AI system reports a "confidence score."
Mathematicians Agree on the Rules but Not the Meaning
You might expect probability to have one settled definition. It has a settled set of rules, and several competing interpretations.
In 1933, Russian mathematician Andrey Kolmogorov published Foundations of the Theory of Probability, which put probability on a rigorous axiomatic footing:
- Every probability is between 0 and 1
- The probability that something in the sample space happens is 1
- For mutually exclusive events, probabilities add
Every interpretation below obeys these rules. They differ in what the numbers are about.
Interpretation 1: Classical (Equally Likely Outcomes)
Developed by Pierre-Simon Laplace and others, the classical view says:
P(event) = favorable outcomes / total equally likely outcomes
A fair die shows a 6 with probability 1/6. Two dice total 7 with probability 6/36 = 1/6.
Strength: perfect for games of chance. Weakness: it needs outcomes that are genuinely equally likely, which real life rarely provides. What are the "equally likely outcomes" of tomorrow's weather?
Interpretation 2: Frequentist (Long-Run Frequency)
The frequentist view, developed by thinkers like John Venn and Richard von Mises, says probability is the long-run proportion of times something happens in repeated trials:
P(heads) = limit of (heads / flips) as flips → ∞
Flip a fair coin 10 times and you might get 7 heads. Flip it 10,000 times and the proportion will almost certainly be very close to 0.5. That's the law of large numbers. Try it with the random number generator.
Strength: objective and testable. Weakness: what about one-time events? The 2028 election will happen only once. There's no long run.
Interpretation 3: Bayesian (Degree of Belief)
The Bayesian view treats probability as a degree of belief, given available information. It applies to one-off events, and it updates as evidence arrives, using Bayes' theorem.
Italian mathematician Bruno de Finetti made the case forcefully. His 1974 book Theory of Probability opens with a provocative line: "Probability does not exist." He meant that probability isn't a physical property of coins or clouds, but a measure of a reasoner's uncertainty.
He also showed why such beliefs must follow Kolmogorov's rules: if they don't, someone can offer you a set of bets you'd accept that guarantees you lose money, no matter what happens. That's called a Dutch book argument.
See Bayes' Theorem and the Mathematics of AI.
One Event, Many Readings
Consider: "The probability of rain tomorrow is 30%."
| Interpretation | Meaning |
|---|---|
| Classical | Hard to apply: outcomes aren't equally likely |
| Frequentist | On many days with conditions like tomorrow's, about 30% had rain |
| Bayesian | Given the forecaster's information, their degree of belief in rain is 0.3 |
Meteorologists effectively use the frequentist reading to check their Bayesian-style forecasts.
An Insider Reference: Calibration
A forecaster is well calibrated if, among all the days they say "30%," it rains on about 30% of them. Among "90%" days, about 90%. Calibration is how you test a probabilistic forecast without repeating any single event.
U.S. National Weather Service precipitation forecasts have historically been well calibrated. Election forecasting became a public test of the idea in 2016, when FiveThirtyEight gave Donald Trump roughly a 29% chance of winning on election day. Many people read that as "he'll lose," but events with a 29% probability happen about as often as a baseball player with a .290 average gets a hit.
AI systems face the same issue. In 2017, Chuan Guo and colleagues showed in "On Calibration of Modern Neural Networks" that many modern deep networks are overconfident: when they report 95% confidence, they're right less often than that. They proposed a simple fix, temperature scaling, that rescales outputs so reported confidence better matches reality.
Two Concepts Worth Knowing
Law of Large Numbers
The law of large numbers says the proportion of successes in repeated independent trials gets closer to the true probability as trials increase. It links frequency and probability.
Calibration
Calibration compares predicted probabilities with observed frequencies. A calibrated model's "80% confident" predictions come true about 80% of the time. The standard normal table helps judge how far observed frequencies can wander by chance.
Quick Answer: What Does a Probability Mean?
Probability follows the same mathematical rules everywhere, but it has several interpretations. The classical view counts equally likely outcomes, the frequentist view treats probability as a long-run frequency, and the Bayesian view treats it as a degree of belief updated with evidence. Forecasts are checked by calibration: events given 30% should happen about 30% of the time.
Try Them Yourself
- Random Number Generator: watch frequencies approach probabilities
- Statistics Formulas: probability rules and distributions
- Standard Normal Table: how much randomness to expect
- Z-Score Calculator: is a streak surprising?
- Bayes' Theorem and the Mathematics of AI: probability as updated belief
- Sampling Techniques Unveiled: frequencies from samples
Keep a two-week log of your local forecast's rain percentages and whether it rained. Group days by forecast and compare. You'll be doing a small calibration study.