← Back to Math Roadmap

Neural Networks

The mathematical foundation of deep learning — from perceptrons to transformers.

1. From Biological Inspiration to Mathematical Model

▼
A single artificial neuron (perceptron) computes a weighted sum of its inputs, adds a bias, and passes the result through an activation function to produce an output. Mathematically: output = σ(Σ w_i x_i + b), where σ is the activation, w_i are learnable weights, and b is the bias. Stacking many neurons into layers creates a Multi-Layer Perceptron (MLP). The Universal Approximation Theorem proves that an MLP with a single hidden layer and enough neurons can approximate ANY continuous function to arbitrary precision. Modern deep networks use many layers (hence "deep") to learn hierarchical representations. In vision: early layers detect edges and textures; middle layers detect shapes and parts; final layers detect whole objects. In language: early layers encode syntax and word sense; middle layers encode semantics and relationships; final layers encode task-specific meaning.
Three essential activation functions. ReLU (Rectified Linear Unit) is the default for hidden layers: fast, avoids vanishing gradients. Sigmoid squashes to (0,1) — used in output layers for binary classification. Tanh squashes to (-1,1) — zero-centered.

2. Training: Backpropagation and the Chain Rule

▼
Neural networks learn through backpropagation, which is simply the multivariable chain rule applied at massive scale. Forward pass: inputs flow through layers producing predictions. The loss function measures prediction error (MSE for regression, Cross-Entropy for classification). Backward pass: the gradient of the loss with respect to every weight is computed by recursively applying the chain rule from output to input. This is automatic differentiation — the same math that powers PyTorch and TensorFlow. Gradient descent updates each weight: w_new = w_old minus learning_rate times gradient. Modern optimizers like Adam adapt the learning rate per parameter using running estimates of gradient moments for faster, more stable convergence.
Weight update via gradient descent. η = learning rate. The partial derivative ∂L/∂w (computed by backpropagation) tells us how much each weight contributed to the error.

3. Transformers: The Architecture That Changed AI

▼
The Transformer (2017, "Attention Is All You Need") replaced recurrence with self-atten(graph attention networks) attention, enabling parallel processing of entire sequences. In a transformer, each layer has two sub-layers: multi-head self-attention and a feed-forward network (two linear transformations with ReLU in between). Residual connections add the input of each sub-layer to its output, then layer normalization stabilizes training. Multi-head attention runs the attention mechanism H times in parallel (typically H=8 or 16), each head with different learned Q, K, V projections. This allows heads to specialize: one head might focus on syntactic dependencies, another on long-range semantic connections, another on positional proximity.

The transformer powers GPT (decoder-only), BERT (encoder-only, trained on masked language modeling), and T5 (encoder-decoder, trained on text-to-text tasks). All modern LLMs — Claude, Gemini, Llama, Mistral — are decoder-only transformers. The key insight: attention computes relevance between EVERY pair of tokens, enabling the model to capture dependencies across arbitrary distances without the bottleneck of sequential recurrence.
Multi-head attention. Each head computes its own scaled dot-product attention with separate learned projections. Outputs are concatenated and projected through W^O.

Key Training Techniques That Made Deep Learning Work

▼
  • Batch Normalization: normalize layer inputs to zero mean and unit variance for each mini-batch. Stabilizes training, allows higher learning rates, reduces sensitivity to initialization. Acts as a mild regularizer.
  • Dropout: randomly disable a fraction p of neurons during each training step. Forces the network to learn redundant representations and prevents co-adaptation of neurons. At test time, all neurons are active but outputs are scaled by (1-p). Like training an ensemble of subnetworks that share weights.
  • Weight Initialization: He initialization (for ReLU networks) draws weights from a normal distribution with variance 2/n_in. Xavier initialization (for tanh/sigmoid) uses variance 1/n_in. Proper initialization prevents signals from exploding or vanishing as they propagate through deep networks.
  • Learning Rate Scheduling: cosine annealing, step decay, or reduce-on-plateau. Start with high learning rate for fast progress, then reduce for fine-tuning. Modern practice: warmup (linearly increase LR) for the first few thousand steps, then cosine decay.
  • Transfer Learning: start from a pre-trained model rather than random initialization. Fine-tune on the target task with a lower learning rate. The pre-trained model provides rich feature representations learned from massive data. This is why you can fine-tune GPT-4 on your specific domain with just hundreds of examples.
  • Gradient Clipping: cap gradient magnitudes to prevent exploding gradients in recurrent networks or unstable training. Typically clip to norm = 1.0 or 5.0.

4. Convolutional Neural Networks (CNNs)

▼
CNNs are specialized for grid-structured data (images, audio spectrograms). A convolutional layer slides small filters (kernels, typically 3x3 or 5x5) across the input, computing dot products at each position. This exploits translational invariance: a feature (edge, texture, pattern) can appear anywhere in the image. Pooling layers (max-pooling, average-pooling) downsample by taking the maximum or average over small windows, reducing spatial dimensions while preserving important features. The receptive field grows with depth: layer 1 sees 3x3 pixels, layer 2 sees 5x5, layer 3 sees 7x7, enabling hierarchical feature learning. CNNs dominated computer vision from 2012 (AlexNet) until Vision Transformers (ViTs) began challenging them in 2020.
2D convolution. The filter K slides over image I, computing dot products. The output (feature map) captures where and how strongly each pattern appears.

Modern Architectures Beyond the MLP

▼
Residual Networks (ResNets)
Skip connections add the input directly to the output of a layer: output = F(x) + x. This allows gradients to flow directly through the network during backpropagation, enabling training of 100+ layer networks that would otherwise suffer from vanishing gradients. ResNets won ImageNet 2015 with 152 layers.
Attention-Based Architectures
Transformers extended beyond NLP. Vision Transformers (ViT) split images into patches, treat each patch as a "token," and apply self-attention. Diffusion models (Stable Diffusion, DALL-E) use U-Net architectures with cross-attention to text embeddings. Graph Neural Networks (GNNs) apply message passing — a form of attention on graph structures — for molecular property prediction and social network analysis.
Mixture of Experts (MoE)
Only a subset of model parameters (the "experts") is activated for each input. A gating network decides which experts to use. This allows models to scale to trillions of parameters while keeping inference cost manageable because only a fraction of the total parameters are used per token. GPT-4 and Mixtral use MoE.

Worked Example: Forward Pass Through a 2-Layer MLP

▼
Worked Example: Forward Pass Through a 2-Layer MLP
Input: x = [2, 3]. Hidden layer: 3 neurons with ReLU. Output: 1 neuron (regression).

Weights W1 = [[1,0.5],[0.5,1],[0,1.5]]. Bias b1 = [0,0,0].
h1 = x·W1 row 1 + b1_1 = 2×1 + 3×0.5 = 3.5 → ReLU(3.5) = 3.5
h2 = 2×0.5 + 3×1 = 4.0 → ReLU(4.0) = 4.0
h3 = 2×0 + 3×1.5 = 4.5 → ReLU(4.5) = 4.5
Hidden output: [3.5, 4.0, 4.5].

Output weights W2 = [1, 0.5, 0.8]. Bias b2 = 0.5.
y = 3.5×1 + 4.0×0.5 + 4.5×0.8 + 0.5 = 3.5 + 2.0 + 3.6 + 0.5 = 9.6.

During backpropagation, the gradient of the loss with respect to each weight is computed by recursively applying the chain rule.

Activation Function Explorer

▼
Plot and compare activation functions: ReLU (blue), sigmoid (red), tanh (green). See how each squashes, saturates, or preserves signal. These functions are the nonlinearities that give neural networks their expressive power.