With many predictors, OLS can
overfit: the model memorizes noise patterns that do not generalize.
Regularization adds a penalty to the OLS objective.
Ridge (L2) adds λΣβ_i², shrinking all coefficients toward zero but never to exactly zero — good when all features matter.
Lasso (L1) adds λΣ|β_i|, driving some coefficients to EXACTLY zero — automatic feature selection.
Elastic Net combines both: λ₁Σ|β_i| + λ₂Σβ_i². The hyperparameter λ controls strength and is chosen by cross-validation.