Retrieval-augmented generation (RAG) explained with a working mental model
By TechlyUpUpdated 3 min readDevelopers
Quick answer
RAG retrieves relevant passages from your own documents and gives them to the model as context, so answers can be grounded in your data. Quality depends mostly on retrieval: how you split documents, how you search (keyword, vector, or both), and how many passages you include. Evaluate retrieval and answers separately, and instruct the model to answer only from supplied context and cite it.
The pipeline in four steps
Every RAG system has these stages.
- Ingest: extract text from documents and split it into chunks with metadata.
- Index: store chunks for search (embeddings for semantic search, and often keyword indexes too).
- Retrieve: find the most relevant chunks for a user question.
- Generate: prompt the model with the question and retrieved chunks, asking for a cited answer.
Where RAG fails
Most failures are retrieval failures: the right passage wasn't found, was split badly, or was crowded out by irrelevant chunks. Other failures come from the model ignoring context or answering beyond it.
Design decisions that matter
Chunk by document structure (headings, sections) rather than fixed character counts where possible. Combine keyword and semantic search for names and codes. Rerank results. Keep source metadata so answers can link back.
Evaluate in two layers
Measure retrieval and generation separately.
Retrieval: for 50 questions with known source passages, is the right passage in the top 5? Generation: given the right passage, is the answer correct, complete, and cited? Track both after every change to chunking, embedding model, or prompt.
Common RAG mistakes
Most disappointing RAG systems share these problems.
- Tuning prompts when the real problem is that retrieval returns the wrong passages.
- Chunking by fixed size and splitting tables, lists, or definitions in half.
- Relying only on vector search for queries containing exact names, codes, or IDs.
- Stuffing too many chunks into the prompt, burying the relevant one.
Worked example: an HR policy assistant
A team builds RAG over 40 HR policy documents. Initial answers are often wrong. Measuring retrieval shows the right passage appears in the top five for only about half of test questions. Switching to heading-based chunking and adding keyword search alongside vectors improves retrieval substantially.
Only then do they refine the prompt to require citations and a “not covered by policy” response. Answer quality improves more from those retrieval changes than from any prompt wording — which is the typical pattern in RAG projects.
Try it yourself
Build a small RAG over 20 public documents. Write 20 questions with known answers and measure top-5 retrieval accuracy before tuning anything else.
Frequently asked questions
Is RAG better than fine-tuning?
They solve different problems. RAG adds up-to-date knowledge from your documents; fine-tuning changes behaviour or style. Many applications use RAG first.
Which vector database should I use?
For small projects, a database you already run (with vector support) or a simple library is enough. Choose based on scale, operations, and filtering needs.
How do I stop RAG answers from making things up?
Improve retrieval, instruct the model to answer only from context, require citations, and have it say when context is insufficient.
Want a suggested next step for your situation?
Share a few details and someone from TechlyUp will get back to you. No automated sequences.
Sources and further reading
- Google Cloud: What is retrieval-augmented generation?
- Prompt Engineering Guide: RAG
- Hugging Face LLM Course
Examples are authored practice material, not measured learner outcomes. Tool behavior can change. Found an error? Contact TechlyUp with the page URL and correction.