An
embedding is a dense vector (typically 384, 768, or 1536 dimensions) that represents text in a semantic space. Similar texts have similar vectors (measured by cosine similarity or Euclidean distance). Embedding models like OpenAI's text-embedding-3, Cohere Embed, and open-source options (BGE, E5, GTE) are trained so that semantically related sentences cluster together.
Retrieval-Augmented Generation (RAG) combines embeddings with LLMs to ground responses in factual documents. The pipeline: (1) split documents into chunks (200-1000 tokens each), (2) embed each chunk and store in a
vector database (Pinecone, Weaviate, Milvus, ChromaDB), (3) when a user queries, embed the query and retrieve the k most similar chunks via cosine similarity, (4) prepend the retrieved chunks to the prompt as context, (5) the LLM generates a response grounded in the retrieved documents. RAG solves hallucination by anchoring output in real sources and enables models to access information beyond their training cutoff.
Cosine similarity is the most common distance metric for embeddings. For two vectors u and v, it measures the cosine of the angle between them, ranging from -1 (opposite) to 1 (identical). Two vectors pointing in the same direction have cosine similarity = 1.