← Back to Math Roadmap

Bayesian Statistics

Updating beliefs with evidence — prior knowledge plus data yields posterior inference.

The Bayesian Philosophy

▼
Bayesian statistics treats parameters as random variables with probability distributions, rather than as fixed unknown constants. This fundamental shift allows you to incorporate prior knowledge into your analysis. You start with a prior distribution representing your belief about a parameter before seeing the data. After observing data, Bayes' theorem updates this prior to a posterior distribution that combines your prior belief with the evidence from the data. The posterior is proportional to the likelihood (how well the data fits the parameter) times the prior. This process naturally accommodates sequential updating: today's posterior becomes tomorrow's prior.
Posterior is proportional to likelihood times prior. The symbol proportional-to means we only need the shape, not the exact normalizing constant.

Markov Chain Monte Carlo (MCMC)

▼
For all but the simplest models, the posterior distribution cannot be computed analytically. MCMC methods solve this by drawing samples from the posterior using clever random walks. The Metropolis-Hastings algorithm proposes a new parameter value, then accepts or rejects it based on how much it improves the posterior probability. Gibbs sampling updates one parameter at a time, drawing from its conditional distribution given all the other parameters. Hamiltonian Monte Carlo (HMC) uses gradient information to make more efficient proposals, avoiding the random-walk inefficiency of Metropolis-Hastings. Modern probabilistic programming languages like Stan and PyMC implement HMC behind the scenes.

Bayesian vs Frequentist

▼
The philosophical divide between Bayesian and frequentist statistics runs deep. Frequentists treat probability as the long-run frequency of events and parameters as fixed. Bayesians treat probability as a measure of uncertainty and parameters as random. A frequentist confidence interval says "95% of intervals from repeated sampling would contain the true parameter." A Bayesian credible interval says "there is a 95% probability that the parameter lies in this interval." Both frameworks have their strengths. Bayesian methods excel when prior knowledge is available, data is scarce, or the model is hierarchical. Frequentist methods are simpler to compute and do not require specifying a prior.

When to Use Bayesian Methods

▼
When to Use Bayesian Methods
Bayesian methods shine with small datasets (the prior regularizes), hierarchical models (partial pooling), sequential updating (online learning), uncertainty quantification (full posterior distribution, not just point estimates), and when you have genuine prior knowledge from previous studies or domain expertise. They are foundational to modern machine learning via Bayesian neural networks, Gaussian processes, and variational inference.

Bayesian Updating Visualizer

▼
Watch how the posterior distribution narrows as more data arrives. Each new observation updates the prior into a more informed posterior.