← Back to Math Roadmap

Probability Theory

Quantifying uncertainty, randomness, and chance — the mathematical foundation of statistics.

The Axioms of Probability

▼
Probability theory is built on three axioms formulated by Kolmogorov in 1933. First, every probability is a number between zero and one, inclusive. Zero means impossible, one means certain. Second, the probability of the entire sample space (something happening) is one. Third, for mutually exclusive events, the probability of either one happening is the sum of their individual probabilities. From these simple axioms, the entire edifice of probability theory is constructed. The sample space (capital Omega) is the set of all possible outcomes. An event is any subset of the sample space. The probability of an event is the sum of probabilities of the outcomes it contains.
Classical probability for equally likely outcomes. The probability of event A is the number of favorable outcomes divided by the total number of possible outcomes.

Conditional Probability and Bayes' Theorem

▼
Conditional probability measures how the probability of an event changes when you know that another event has occurred. The probability of A given B is the probability that both A and B occur, divided by the probability of B. This seemingly simple formula is the foundation of Bayes' theorem, which reverses conditional probabilities: it gives the probability of A given B in terms of the probability of B given A. Bayes' theorem is the mathematical engine behind spam filters, medical diagnosis, machine learning, and the scientific method itself — it formalizes how evidence updates our beliefs.
Bayes' theorem. The posterior probability P(A|B) is proportional to the likelihood P(B|A) times the prior P(A), normalized by the evidence P(B).

Random Variables and Distributions

▼
A random variable assigns a numerical value to each outcome in the sample space. A discrete random variable takes on a countable set of values (like the number of heads in ten coin flips). A continuous random variable takes on any value in an interval (like the exact height of a person). Each distribution has a probability mass function (for discrete variables) or a probability density function (for continuous variables). The expected value is the long-run average of the random variable. The variance measures how spread out the distribution is. Key distributions include the binomial (counting successes), Poisson (counting rare events), normal/Gaussian (the famous bell curve), and exponential (waiting times).
Normal (Gaussian) probability density function. The parameters are mu (mean, center) and sigma (standard deviation, spread).

The Central Limit Theorem

▼
The Central Limit Theorem (CLT) is the crown jewel of probability theory. It states that the sum (or average) of many independent random variables, regardless of their individual distributions, tends toward a normal distribution. This explains why the normal distribution appears everywhere in nature and statistics — many observed quantities are the result of adding up many small independent factors. The CLT justifies the use of the normal distribution in confidence intervals and hypothesis testing, even when the underlying data is not normally distributed.

Worked Example

▼
Worked Example
A bag contains 3 red and 5 blue marbles. You draw 2 marbles without replacement. Find the probability that both are red.
P(both red) = (3/8) × (2/7) = 6/56 = 3/28 ≈ 0.107.

A fair coin is flipped 10 times. What is the probability of exactly 7 heads?
Binomial: C(10,7) × (1/2)¹⁰ = 120 × 1/1024 = 120/1024 ≈ 0.117.

Distribution Explorer

▼
Explore probability distributions. Toggle Normal (bell curve), Exponential (waiting times), and Poisson (counts). Adjust mu/sigma or lambda with sliders to see how the shape, center, and spread change.