# RAG Explained: How Retrieval-Augmented Generation Actually Works

* * *

# RAG Explained: How Retrieval-Augmented Generation Actually Works

Every production LLM system you have used recently, from coding assistants to support chatbots, probably runs on RAG. Retrieval-augmented generation is the architecture that lets a language model answer from your documents instead of its training data. If you build with LLMs, you need to understand it properly, because most RAG failures come from misunderstanding the retrieval half.

This guide covers the full picture: why RAG exists, how each stage works under the hood, where implementations break, and how to tell whether you need RAG, fine-tuning, or both. If you want a second perspective alongside this one, the Hashnode community also hosts [a complete guide to RAG](https://hashnode.com/posts/retrieval-augmented-generation/6995b760689b7a26f44895d7) worth reading.

## The problem RAG solves

A base language model has two weaknesses that matter in production.

First, its knowledge is frozen at training time. Ask about your company's new pricing page or last week's incident postmortem and it either guesses or refuses.

Second, it hallucinates. When the model lacks the facts, it generates fluent text that sounds right and isn't. In a demo this is amusing. In a customer-facing support bot it is a liability.

Fine-tuning doesn't fix either problem well. It teaches the model behavior and format, not facts, and updating facts means retraining. RAG takes a different approach: give the model a cheat sheet at query time. Retrieve the relevant passages from your documents, stuff them into the prompt, and let the model answer from that context. The knowledge stays in your database where you can update it in seconds, and the model stays generic.

## The three stages

Every RAG system, from a weekend prototype to enterprise infrastructure, runs the same pipeline.

**1\. Index.** Your documents get split into chunks, each chunk gets converted into an embedding (a vector of numbers capturing its meaning), and those vectors go into a vector store built for similarity search. This happens offline, before any user asks anything.

**2\. Retrieve.** A user asks a question. The question gets embedded with the same model, the vector store returns the most similar chunks, and optionally a reranker reorders them by true relevance.

**3\. Generate.** The top chunks are inserted into the prompt alongside the question, with an instruction to answer from the provided context. The model generates a grounded answer, ideally with citations to the chunks it used.

If you want a slower walkthrough of each stage with diagrams, this [RAG explainer](https://www.searchenginebasics.dev/ai-search/rag-explained/) covers the architecture in more depth.

## How retrieval actually works

Retrieval is where RAG systems are won or lost, so let's open it up.

**Chunking.** Documents get split into pieces, typically a few hundred tokens each, with some overlap between consecutive chunks. Too large and the retriever returns diluted passages full of irrelevant text. Too small and a single chunk rarely contains a complete answer. Overlap exists so that an answer spanning a boundary still appears whole in at least one chunk. There is no universal best size; it depends on your documents and queries, and it's worth experimenting.

**Embeddings.** An embedding model converts each chunk into a dense vector, usually a few hundred to a couple thousand dimensions. Chunks with similar meaning end up close together in this space. The choice of embedding model matters more than most people expect: a model trained on general web text performs differently on legal contracts or medical notes. Match the embedding model to your domain.

**Vector search.** At query time, the question is embedded and the store finds the nearest chunk vectors, typically by cosine similarity. Nobody does exact search at scale; production systems use approximate nearest neighbor indexes like HNSW that trade a tiny bit of accuracy for massive speed. This is the component [vector database implementation](https://www.esaholic.com/services/ai-data-engineering/vector-database-implementation/) gives you: indexing, filtering, and scaling that a prototype with in-memory numpy won't survive in production.

**Reranking (optional but recommended).** Vector similarity is fast and decent. A cross-encoder reranker is slow and much better: it reads the query and each candidate chunk together and scores true relevance. The standard pattern is retrieve 50 candidates by vector search, rerank to the top 5, and generate from those. This single addition fixes a large share of "the answer was in the docs but the bot missed it" complaints.

![RAG retrieval flow diagram](https://cdn.hashnode.com/uploads/covers/69941a713d473a38a8c5bb41/04bb0e9f-635b-40e8-a00f-f69bbdb6699e.webp align="center")

A minimal sketch of the whole flow looks like this:

```python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")

# 1. Index: chunk docs and embed them
chunks = [
    "RAG retrieves relevant passages before generating.",
    "Chunk size controls the tradeoff between precision and context.",
]
index = [(c, model.encode(c)) for c in chunks]

# 2. Retrieve: embed the query, find nearest chunks
query_vec = model.encode("How does RAG work?")
ranked = sorted(index, key=lambda item: cosine_sim(query_vec, item[1]), reverse=True)
top_chunks = [c for c, _ in ranked[:3]]

# 3. Generate: answer from retrieved context only
prompt = (
    "Answer using only the context below.\n\n"
    f"Context:\n{chr(10).join(top_chunks)}\n\n"
    "Question: How does RAG work?"
)
answer = llm.generate(prompt)
```

Prototypes skip the vector database and reranker. Production systems don't.

## RAG vs fine-tuning: when to use which

This is the most common architecture question I hear, and the answer is less either-or than people expect.

|  | RAG | Fine-tuning |
| --- | --- | --- |
| **What it changes** | The knowledge available at query time | The model's behavior, tone, and format |
| **Updating facts** | Instant: update the documents | Slow: retrain or use adapters |
| **Best for** | Q&A over your docs, support bots, research assistants | Style, structured output, domain jargon |
| **Failure mode** | Bad retrieval poisons good generation | Stale facts baked into weights |
| **Cost profile** | Inference-time retrieval cost | Upfront training cost |

![RAG vs fine-tuning comparison](https://cdn.hashnode.com/uploads/covers/69941a713d473a38a8c5bb41/76268c40-9c85-4aef-aa38-5ac162434148.webp align="center")

The practical rule: if the problem is "the model doesn't know X," that's RAG. If the problem is "the model doesn't answer like we need," that's fine-tuning. Most serious systems use both: fine-tune for behavior, retrieve for knowledge. For a deeper comparison of the tradeoffs, [RAG vs fine-tuning](https://www.softbrixai.com/insights/rag-vs-fine-tuning/) is worth reading before you commit to an architecture, and [this Hashnode deep-dive on RAG and fine-tuning approaches](https://hashnode.com/posts/understanding-retrieval-augmented-generation-rag-and-fine-tuning-approaches/66a6a1f2cc791e4ec006fee1) walks through the same decision from another angle.

## Where RAG implementations break

After seeing enough RAG systems in the wild, the failure modes repeat:

**Bad chunking.** The classic: chunks split mid-answer, so no single retrieved passage contains what the user asked. Fix with overlap, semantic chunking (splitting at topic boundaries instead of fixed token counts), or smaller chunks with more of them retrieved.

**Wrong top-k.** Retrieving too few chunks starves the model of context; retrieving too many drowns the answer in noise and burns your context window. Start with 5 after reranking and tune from there.

**No reranking.** Vector search alone returns "similar-ish" chunks. Without a reranker, the generator works from mediocre context and the answers show it.

**Stale index.** Documents updated, index not rebuilt. Now the bot confidently quotes last year's pricing. Your indexing pipeline needs the same freshness SLAs as the rest of your system.

**The prompt lets the model wander.** If your system prompt doesn't firmly instruct the model to answer from the provided context and say "I don't know" when the context lacks the answer, it will happily hallucinate around your retrieval. Constrain it.

**No evaluation.** Teams ship RAG with no measurement: no test set of questions, no tracking of retrieval hit rate or answer faithfulness. You can't improve what you don't measure. Build a small eval set on day one.

## RAG in production: what changes at scale

A prototype fits in one script. Production adds requirements that prototypes ignore.

**The vector store becomes infrastructure.** Filtering by metadata (tenant, date, document type), handling millions of vectors, replicating across regions. This is where dedicated vector databases earn their place over bolt-on solutions.

**Freshness becomes a pipeline.** Change data capture from your source systems, re-chunking, re-embedding, index updates with zero downtime. Treat it like any other data pipeline with monitoring and alerting.

**You monitor retrieval, not just generation.** Log what got retrieved for each query. When users complain about answers, the retrieved chunks usually explain why. Retrieval hit rate on your eval set is the metric that predicts answer quality.

**Security and access control matter.** The retriever must respect document permissions: a user should never get chunks from documents they can't access. Filter at retrieval time, not after generation.

## The checklist

*   Choose an embedding model matched to your domain, not just the default
    
*   Tune chunk size and overlap against real queries, not guesses
    
*   Use approximate nearest neighbor search; don't brute-force at scale
    
*   Add a reranker: retrieve 50, rerank to 5
    
*   Constrain the generator to the retrieved context with a firm system prompt
    
*   Keep the index fresh with a real pipeline, not manual rebuilds
    
*   Enforce document access control at retrieval time
    
*   Build an eval set of questions on day one and track retrieval hit rate
    
*   Start with RAG for knowledge problems, fine-tune for behavior problems
    

RAG looks simple in a diagram and humbles everyone in production. But the pattern is learnable: chunk well, embed well, retrieve generously, rerank ruthlessly, and generate from context. Get the retrieval half right and the generation half mostly takes care of itself.

[*Follow Denebrix AI on Hashnode*](https://hashnode.com/@denebrixai) *for more practical guides on building with LLMs and AI infrastructure.*
