← Back to Math Roadmap

ML Math Modeling

From data to predictions — the mathematical framework of machine learning.

The Modeling Pipeline

▼
Building a machine learning model follows a systematic pipeline. First, data collection and cleaning — real-world data is messy with missing values, outliers, and errors. Then, feature engineering transforms raw data into informative numerical features. The model is a mathematical function with learnable parameters. The loss function quantifies how wrong the predictions are: Mean Squared Error for regression, Cross-Entropy for classification. The optimizer updates parameters to minimize the loss. Finally, evaluation on a held-out test set measures true generalization performance. The bias-variance tradeoff is fundamental: simple models underfit (high bias), complex models overfit (high variance). The goal is the sweet spot in between.
Loss functions. MSE = Mean Squared Error (regression). CE = Cross-Entropy (classification). The hat indicates predicted values.

Supervised vs Unsupervised Learning

▼
Supervised learning uses labeled data: each training example has an input and a known output. The goal is to learn a mapping from inputs to outputs. Examples: predicting house prices (regression), classifying emails as spam (classification). Unsupervised learning uses unlabeled data: the goal is to discover hidden structure. Examples: clustering customers into segments, dimensionality reduction for visualization. Semi-supervised learning uses a small amount of labeled data combined with a large amount of unlabeled data. Reinforcement learning has an agent that learns by interacting with an environment, receiving rewards or penalties for its actions.

Validation and Cross-Validation

▼
Evaluating a model on the same data used for training gives a misleadingly optimistic performance estimate. The standard split is 60 percent training, 20 percent validation, 20 percent test. The training set is used to learn parameters. The validation set is used to tune hyperparameters and select the best model. The test set is touched ONLY ONCE at the very end to report final performance. K-fold cross-validation provides a more robust estimate: split data into k folds, train on k minus one folds and test on the remaining fold, rotating k times. This is especially valuable for small datasets where a single split might be unlucky.

Best Practices

▼
  • Start simple: try linear regression or logistic regression before deep neural networks.
  • Normalize features: scale all features to similar ranges (zero mean, unit variance).
  • Feature engineering matters more than model tuning: domain knowledge creates powerful features.
  • Use train/validation/test splits: never evaluate on training data.
  • Cross-validate for small datasets where a single split is unreliable.
  • Monitor for data leakage: ensure no information from the test set leaks into training.
  • Check for class imbalance: if one class is rare, use stratified sampling or adjust class weights.

Model Fit Visualizer

▼
See how different models fit data. Underfitting (too simple) leaves patterns. Overfitting (too complex) captures noise. The best model finds the middle ground.