← Back to Math Roadmap

Prompt Engineering & LLM Architecture

The complete guide to instructing AI models — embeddings, cross-encoders, re-rankers, RAG pipelines, sampling strategies, and the mathematics behind modern language models.

1. How Language Models Work: Tokens, Probabilities, and Attention

▼
A Large Language Model (LLM) is fundamentally a next-token predictor. Input text is split into tokens — sub-word units like "trans" and "former" that combine into "transformer." Each token has a unique ID in the vocabulary (typically 50,000 to 250,000 tokens). The model processes the sequence through dozens of transformer layers where each layer applies self-attention: every token attends to every other token, computing relevance scores. The final layer outputs a probability distribution over the entire vocabulary for the next token. This is autoregressive generation: the predicted token is appended to the input, and the process repeats to generate text one token at a time.

Temperature controls creativity: low temperature (close to 0) makes the model deterministic and safe. High temperature (>1) flattens the distribution, encouraging diverse but potentially less coherent outputs. Top-k restricts to the k most probable tokens. Top-p (nucleus sampling) selects from the smallest set whose cumulative probability exceeds p — dynamically adapting based on the distribution's shape.
Next-token prediction. h_t = hidden state after processing tokens 1 through t. W = output projection matrix. Softmax normalizes the raw scores into a probability distribution over the vocabulary.
Scaled dot-product attention. Each token has Query (what am I looking for?), Key (what do I offer?), and Value (what information do I carry?). The softmax produces attention weights summing to 1 across all tokens.

2. Prompting Strategies: Zero-Shot, Few-Shot, and Chain-of-Thought

▼
The way you structure your input dramatically affects model performance. Zero-shot prompting gives only the instruction with no examples. The model relies entirely on its pre-training knowledge. Works well for simple tasks but degrades on complex reasoning.

Few-shot prompting provides 2 to 8 input-output examples before the actual query. The model performs in-context learning: it infers the pattern from examples without any weight updates. This is a form of implicit Bayesian inference performed by attention. Examples should be diverse, representative, and formatted identically to the desired output.

One-shot prompting is few-shot with exactly 1 example — enough to establish the format but relying heavily on the model's generalization.

Chain-of-Thought (CoT) extends few-shot by including REASONING STEPS in examples. Instead of just "input → output," CoT shows "input → step 1: analyze → step 2: compute → step 3: conclude → output." This multiplies accuracy on arithmetic (from 18% to 57% on GSM8K), logic puzzles, and multi-step problems. CoT works because it decomposes complex reasoning into a sequence of simpler steps, each of which the model handles well.

Prompting Strategies Comparison

▼
  • Zero-shot: only the instruction, no examples. Fastest, simplest. Best for straightforward tasks (translation, summarization, classification).
  • One-shot: instruction + 1 example. Establishes output format. Good when the model needs to see the desired structure but the task is simple.
  • Few-shot (k-shot): instruction + 2-8 examples. Enables in-context learning. The model infers the task pattern from examples. Performance typically improves up to ~8 examples, then plateaus.
  • Chain-of-Thought: few-shot examples WITH reasoning traces. Transforms complex tasks into step-by-step reasoning. Multiplies accuracy on math, logic, and planning by 2-10x.
  • Self-Consistency: sample multiple CoT reasoning paths, then majority-vote the final answer. Further improves reliability by reducing variance from single-sample reasoning.

3. Embeddings and Vector Search: The Foundation of RAG

▼
An embedding is a dense vector (typically 384, 768, or 1536 dimensions) that represents text in a semantic space. Similar texts have similar vectors (measured by cosine similarity or Euclidean distance). Embedding models like OpenAI's text-embedding-3, Cohere Embed, and open-source options (BGE, E5, GTE) are trained so that semantically related sentences cluster together.

Retrieval-Augmented Generation (RAG) combines embeddings with LLMs to ground responses in factual documents. The pipeline: (1) split documents into chunks (200-1000 tokens each), (2) embed each chunk and store in a vector database (Pinecone, Weaviate, Milvus, ChromaDB), (3) when a user queries, embed the query and retrieve the k most similar chunks via cosine similarity, (4) prepend the retrieved chunks to the prompt as context, (5) the LLM generates a response grounded in the retrieved documents. RAG solves hallucination by anchoring output in real sources and enables models to access information beyond their training cutoff.

Cosine similarity is the most common distance metric for embeddings. For two vectors u and v, it measures the cosine of the angle between them, ranging from -1 (opposite) to 1 (identical). Two vectors pointing in the same direction have cosine similarity = 1.
Cosine similarity. Dot product divided by product of magnitudes. Range [-1, 1]. For normalized vectors (unit length), it simplifies to just the dot product u·v.

4. Bi-Encoders vs Cross-Encoders: The Two Retrieval Architectures

▼
There are two fundamental approaches to measuring text relevance, each with distinct tradeoffs. Understanding when to use each is critical for building effective retrieval systems.

Bi-Encoder vs Cross-Encoder Deep Dive

▼
Bi-Encoder (Dual Encoder)
A bi-encoder encodes the query and EACH document SEPARATELY into dense vectors. The query vector and each document vector are computed independently, then compared via cosine similarity or dot product. Because document vectors can be PRE-COMPUTED and stored in a vector database, retrieval is extremely fast: you only need to embed the query at search time and find nearest neighbors. This makes bi-encoders ideal for first-stage retrieval over millions of documents.

However, bi-encoders have a fundamental limitation: the query and document never interact directly. The model must compress all information about each document into a single fixed-size vector. Subtle interactions between query terms and document terms are lost. If the query asks "benefits of exercise for sleep quality," a bi-encoder might retrieve documents about exercise AND documents about sleep, but miss the crucial INTERACTION between exercise and sleep.

Popular bi-encoder models: Sentence-BERT (SBERT), OpenAI text-embedding-3, Cohere Embed, BGE (BAAI General Embedding), E5 (EmbEddings from bidirEctional Encoder rEpresentations), GTE (General Text Embeddings).
Cross-Encoder
A cross-encoder feeds the query AND document TOGETHER as a single input pair through a transformer model. The model produces a single relevance score (typically a scalar between 0 and 1) by attending across ALL tokens in both the query and the document SIMULTANEOUSLY. This allows rich, fine-grained interaction between every query term and every document term. Cross-encoders capture nuances like negation ("not recommended for children"), paraphrasing, and partial matches that bi-encoders miss.

The tradeoff: cross-encoders are SLOW. You cannot pre-compute document vectors because the model must process the query and each document together. For a corpus of 1 million documents, you would need 1 million forward passes per query — making cross-encoders impractical for first-stage retrieval.

Popular cross-encoder models: BERT-based re-rankers (monoBERT, monoT5), Cohere Rerank, cross-encoder/ms-marco-MiniLM, BGE-Reranker. These models are trained on relevance-labeled query-document pairs from datasets like MS MARCO.

5. Re-Ranking: The Two-Stage Retrieval Architecture

▼
Modern search systems combine the speed of bi-encoders with the accuracy of cross-encoders in a two-stage pipeline. This is the industry-standard architecture used by Google, Bing, and every production RAG system.

Stage 1 — Candidate Retrieval (Bi-Encoder): The query is embedded and compared against pre-computed document vectors in a vector database. The top K candidates (typically K = 100-1000) are retrieved. This is fast because only one neural network forward pass is needed per query.

Stage 2 — Re-Ranking (Cross-Encoder): Each of the K candidates is fed TOGETHER with the query through a cross-encoder model. The cross-encoder computes a precise relevance score for each pair. The candidates are re-ranked by these scores, and the top N (typically N = 3-10) are sent to the LLM as context.

This architecture achieves near-cross-encoder accuracy at near-bi-encoder speed. The bi-encoder does the "coarse" filtering over millions of documents. The cross-encoder does the "fine" re-ranking over only the top K. The parameter K balances accuracy vs latency: larger K means more accurate retrieval but slower re-ranking. Production systems typically use K = 100-200.

Example numbers: 1M documents → bi-encoder retrieves top 200 → cross-encoder re-ranks 200, keeps top 5 → LLM generates answer from those 5 chunks. Total latency: ~100ms for embedding + ~200ms for re-ranking + ~500ms for generation = under 1 second.
Cross-encoder score. The function fθ is a transformer that processes the concatenation of query q and document d_i and outputs a scalar relevance score. No pre-computation possible.
Bi-encoder score via cosine similarity. Eθ(q) and Eφ(d_i) are separately computed embeddings. Only the query needs computing at search time — document embeddings are pre-indexed.

6. Advanced RAG Techniques: Beyond Naive Retrieval

▼
Production RAG systems go far beyond "chunk, embed, retrieve." Several advanced techniques dramatically improve retrieval quality:

Advanced RAG Techniques

▼
  • HyDE (Hypothetical Document Embeddings): Instead of embedding the user query directly, first generate a hypothetical answer document using the LLM, then EMBED that generated document and use IT for retrieval. The hypothetical answer is closer in style and vocabulary to actual documents in the corpus, improving retrieval relevance. Works because LLM-generated text matches document distribution better than short queries.
  • Parent Document Retrieval: Store small chunks for embedding (to get precise retrieval) but return the LARGER parent document (or surrounding context) to the LLM. This provides richer context while maintaining retrieval precision. The chunk size for embeddings is optimized separately from the context size for generation.
  • Multi-Query Retrieval: Generate multiple rephrased versions of the user query, retrieve documents for each variant, and deduplicate. This handles the vocabulary mismatch problem: the user might ask about "car repair costs" but documents discuss "automobile maintenance expenses." Multiple query formulations bridge the terminology gap.
  • Contextual Compression: After retrieving documents, use a smaller, faster model to extract ONLY the most relevant sentences or paragraphs before sending to the LLM. This reduces prompt length (saving cost and latency) and removes noise that could confuse the LLM.
  • Re-Ranking (as covered): Cross-encoder re-ranking after initial bi-encoder retrieval is the single highest-impact improvement to RAG quality. Virtually every production system uses it.

7. Embedding Model Architectures and Training

▼
Embedding models are trained to produce vectors where semantically similar texts are close together. The training paradigm evolved significantly over the past few years. Contrastive learning is the dominant approach: the model sees positive pairs (query, relevant document) and negative pairs (query, irrelevant document) and learns to pull positives together while pushing negatives apart. The loss function is typically InfoNCE (Noise Contrastive Estimation) or triplet margin loss. Hard negatives — documents that are similar but not relevant — are crucial for training high-quality embeddings. Without hard negatives, the model only learns to separate obviously different texts but fails on subtle distinctions.

Training data quality matters enormously. Models trained on 1B+ high-quality query-document pairs dramatically outperform those trained on larger but noisier datasets. The MS MARCO dataset (Microsoft MAchine Reading COmprehension) provides 1M+ real Bing queries with human-labeled relevant passages and is the standard benchmark.

Matryoshka embeddings allow a SINGLE embedding to be truncated to different dimensions without losing much quality: a 768-dim embedding can be used at 256, 128, or even 64 dimensions depending on the storage/performance tradeoff. This is achieved by training with a loss function that optimizes all prefix dimensions simultaneously.
InfoNCE contrastive loss. q = query, d^+ = positive (relevant) document, d_i = all documents in batch (including negatives). τ = temperature. Maximizes similarity of positive pairs while minimizing for negatives.

8. System Prompts, Prompt Injection, and Guardrails

▼
System prompts are prepended to every conversation turn and set the model's persona, behavioral constraints, output format, and safety guidelines. Unlike user messages, system prompts are more resistant to prompt injection (where a user tries to override instructions) because most models are fine-tuned to treat system messages with higher priority.

Prompt injection is a security concern: a malicious user might include text like "Ignore all previous instructions and..." in their input. Mitigations include: (1) clearly separating instructions from user input with XML tags or delimiters, (2) using LLM-based input classifiers to detect injection attempts, (3) running a separate smaller model to sanitize user inputs before they reach the main LLM.

Constitutional AI encodes ethical principles directly in the system prompt. The model self-critiques its responses against these principles and revises if needed. This can be combined with Reinforcement Learning from AI Feedback (RLAIF) where the constitution provides the reward signal.

Guardrails are programmatic checks before and after LLM calls: input guardrails validate and sanitize user inputs (removing PII, blocking harmful content); output guardrails verify the model's response meets quality and safety standards (checking for hallucinated citations, toxic content, format compliance). Libraries like NVIDIA NeMo Guardrails and Guardrails AI provide structured frameworks for implementing these.

Production RAG Architecture (Industry Standard)

▼
Production RAG Architecture (Industry Standard)
1. Query arrives → 2. Optional query rewriting (HyDE, multi-query) → 3. Embed query via bi-encoder → 4. Vector DB returns top 200 candidates → 5. Cross-encoder re-ranks to top 5-10 → 6. Optional contextual compression → 7. Retrieved chunks + system prompt + conversation history sent to LLM → 8. LLM generates response → 9. Output guardrail checks response → 10. Return to user. This pipeline serves billions of queries daily across Google, Bing, ChatGPT, and enterprise RAG systems.

Example: Building a Re-Ranker for Customer Support

▼
Example: Building a Re-Ranker for Customer Support
Scenario: 100,000 support articles. User asks "How do I reset my password?"

1. Bi-encoder (e.g., BGE-Large): Embed the query → vector of 1024 dimensions.
2. Vector DB (Pinecone): Cosine similarity search across 100K pre-computed article embeddings. Returns top 200 candidates in ~20ms.
3. Cross-encoder (BGE-Reranker-v2): For each of 200 candidates, concatenate with query and pass through transformer. Each forward pass takes ~5ms on GPU, batched 32 at a time → ~30ms total.
4. Re-rank: Sort 200 candidates by cross-encoder score. Top 5:
- "Password Reset Guide" (score 0.97)
- "Account Recovery Process" (0.89)
- "Two-Factor Authentication Setup" (0.45) — not relevant but mentions "password"
- "Login Troubleshooting" (0.41)
- "Security Best Practices" (0.12)
5. Send top 3 articles as context to LLM.
6. LLM: "To reset your password, go to Settings → Password → Reset. If you cannot log in, use the Account Recovery process at..."

Sigmoid & Softmax Explorer

▼
The softmax and sigmoid functions are the mathematical heart of LLM outputs. Explore the sigmoid 1/(1+e^{-ax}) — it squashes any real number to (0,1) and is used in gating mechanisms, attention weights, and probability estimation throughout transformer architectures.