← Back to Math Roadmap

Descriptive Statistics

Summarizing and visualizing data — the first step in any statistical analysis.

Measures of Central Tendency

▼
When you have a dataset, the first question is: where is the center? Three measures answer this. The mean (average) is the sum of all values divided by the count — it is sensitive to outliers and shifts toward extreme values. The median is the middle value when data is sorted — it is robust against outliers. The mode is the most frequent value. For symmetric distributions, the mean and median coincide. For right-skewed distributions (like income), the mean exceeds the median. The trimmed mean discards the highest and lowest values before averaging, combining some robustness with the mean's use of all data.
Sample mean (x-bar) and sample variance (s-squared). The variance uses n-1 in the denominator for the unbiased sample estimator.

Measures of Dispersion

▼
Center alone is not enough — you need to know how spread out the data is. The range (maximum minus minimum) is simple but sensitive to a single outlier. The interquartile range (IQR) spans the middle fifty percent of data, from the 25th to the 75th percentile, and is robust. The variance measures the average squared deviation from the mean. Its square root, the standard deviation, has the same units as the data and is the most commonly used measure of spread. For a normal distribution, about sixty-eight percent of data falls within one standard deviation of the mean, about ninety-five percent within two, and about ninety-nine point seven percent within three.

Correlation and Visualization

▼
Correlation measures the strength and direction of the linear relationship between two variables. The Pearson correlation coefficient r ranges from negative one (perfect negative linear relationship) to positive one (perfect positive linear relationship). A value of zero indicates no linear relationship (though a nonlinear relationship may still exist). The square of r represents the proportion of variance in one variable explained by the other. Always plot your data before computing correlations — the famous Anscombe Anscombe quartet (always visualize) quartet consists of four datasets with identical means, variances, and correlations but dramatically different patterns when plotted. Visualizations (scatter plots, box plots, histograms) reveal patterns that summary statistics cannot.

Key Descriptive Statistics

▼
  • Mean: the arithmetic average. Sensitive to outliers and skewed distributions.
  • Median: the middle value. Robust — a single extreme value does not change it.
  • Standard deviation: the square root of variance. Average distance of data points from the mean.
  • Quartiles: Q1 (25th percentile), Q2 (median), Q3 (75th percentile). Divide sorted data into four equal groups.
  • IQR: Q3 minus Q1. Covers the middle 50 percent of the data.
  • Skewness: measures asymmetry. Positive skew means the tail extends to the right.
  • Kurtosis: measures tail heaviness. High kurtosis means more extreme values than a normal distribution.

Worked Example

▼
Worked Example
Dataset: 2, 4, 4, 4, 5, 5, 7, 9. Find mean, median, mode, variance, standard deviation.
Mean: (2+4+4+4+5+5+7+9)/8 = 40/8 = 5.
Median: sorted data, average of 4th and 5th = (4+5)/2 = 4.5.
Mode: 4 (appears 3 times).
Variance: ((2−5)²+(4−5)²+(4−5)²+(4−5)²+(5−5)²+(5−5)²+(7−5)²+(9−5)²)/(8−1) = (9+1+1+1+0+0+4+16)/7 = 32/7 ≈ 4.57.
Std dev: sqrt(4.57) ≈ 2.14.

Data Distribution Visualizer

▼
Explore probability distributions. Toggle Normal (bell curve), Exponential (waiting times), and Poisson (counts). Adjust mu/sigma or lambda with sliders to see how the shape, center, and spread change.