← Back to Math Roadmap

Regression Analysis

Modeling relationships between variables — predicting outcomes from predictors.

What Is Regression?

▼
Regression models the relationship between a dependent variable y and one or more independent variables x. Simple linear regression fits a straight line y = β₀ + β₁x + ε, where β₀ is the intercept, β₁ is the slope, and ε represents the random error. The slope tells you the expected change in y for a one-unit increase in x. Ordinary Least Squares (OLS) finds the β values that minimize the sum of squared vertical distances from the data points to the line. OLS has a beautiful closed-form matrix solution: β̂ = (X^TX)^{-1}X^Ty.
OLS estimator. X = design matrix (each row is an observation, each column a feature). The hat denotes an estimate. This minimizes the residual sum of squares ||y - Xβ||².

R-Squared and Model Quality

▼
The coefficient of determination R² measures the proportion of variance in y explained by the model. R² = 1: perfect prediction. R² = 0: no better than just predicting the mean. However, R² ALWAYS increases when you add more predictors, even random noise variables. The adjusted R² penalizes model complexity to enable fair comparison. R² also says nothing about causality or model correctness. Hypothesis testing on coefficients provides complementary evidence.

Regularization: Taming Complexity

▼
With many predictors, OLS can overfit: the model memorizes noise patterns that do not generalize. Regularization adds a penalty to the OLS objective. Ridge (L2) adds λΣβ_i², shrinking all coefficients toward zero but never to exactly zero — good when all features matter. Lasso (L1) adds λΣ|β_i|, driving some coefficients to EXACTLY zero — automatic feature selection. Elastic Net combines both: λ₁Σ|β_i| + λ₂Σβ_i². The hyperparameter λ controls strength and is chosen by cross-validation.
Ridge regression (L2). As λ grows, coefficients shrink. At λ=0, we recover OLS. At λ→∞, all coefficients approach 0.
Lasso regression (L1). The absolute value penalty creates sparsity: many β_i = 0 exactly. Lasso performs feature selection automatically.

Worked Example

▼
Worked Example
Five data points: (1,2), (2,4), (3,5), (4,4), (5,5). Find OLS line.
Means: x̄=3, ȳ=4.
β₁ = Σ(x_i-3)(y_i-4) / Σ(x_i-3)² = ((-2)(-2)+(-1)(0)+(0)(1)+(1)(0)+(2)(1))/(4+1+0+1+4) = 6/10 = 0.6.
β₀ = 4 - 0.6×3 = 2.2.
Regression line: y = 2.2 + 0.6x. For x=6, predicted y = 2.2+3.6 = 5.8.

Slope and Intercept Explorer

▼
The OLS line y = β₀ + β₁x passes through (x̄,ȳ) — the center of mass of the data. Use the function plotter to explore how changing slope and intercept affects predictions.