The Transformer (2017, "Attention Is All You Need") replaced recurrence with
self-atten(graph attention networks) attention, enabling parallel processing of entire sequences. In a transformer, each layer has two sub-layers:
multi-head self-attention and a
feed-forward network (two linear transformations with ReLU in between). Residual connections add the input of each sub-layer to its output, then layer normalization stabilizes training.
Multi-head attention runs the attention mechanism H times in parallel (typically H=8 or 16), each head with different learned Q, K, V projections. This allows heads to specialize: one head might focus on syntactic dependencies, another on long-range semantic connections, another on positional proximity.
The transformer powers GPT (
decoder-only), BERT (encoder-only, trained on masked language modeling), and T5 (encoder-decoder, trained on text-to-text tasks). All modern LLMs — Claude, Gemini, Llama, Mistral — are decoder-only transformers. The key insight: attention computes relevance between EVERY pair of tokens, enabling the model to capture dependencies across arbitrary distances without the bottleneck of sequential recurrence.