Retrieval-augmented generation (RAG): giving models your data
Ask ChatGPT "How many vacation days do I get after three years here?" and it has no idea. It never read your company handbook. It might confidently invent an answer, which is worse than saying nothing.
RAGRAGA method that lets an AI model answer using your own documents, retrieving relevant passages before generating a response instead of relying only on training data.View full definition → fixes this. It's the technique behind almost every "chat with your documents" product you've seen: customer support bots, internal HR assistants, legal research tools. By the end of this lesson you'll understand exactly how it works, and you'll see the core of it in about 20 lines of Python.
The problem RAG solves
A large language modellarge language modelA Large Language Model is an AI system trained on vast text data to predict and generate language, enabling tasks like writing, summarizing, and answering questions.View full definition → (LLM) like GPT-5 or Claude knows what was in its training data. It does not know:
- Your company's internal policies
- Last week's meeting notes
- Your private customer records
- Anything that changed after its training cutoff
You could try pasting your entire 200-page handbook into the chat. But models have a context windowcontext windowThe context window is the maximum amount of text (measured in tokens) a language model can process at once, including both the input prompt and the generated output.View full definition →: a limit on how much text they can read at once. Even with today's large windows, dumping everything in is slow, expensive, and makes the model less accurate (it gets distracted by irrelevant text).
RAG's insight: don't give the model everything. Find the few relevant paragraphs and give it only those.
The three steps of RAG
Think of RAG as a smart librarian. You ask a question, the librarian runs to find the three most relevant pages, and hands them to the model along with your question.
- Retrieval: Search your documents for the chunks most relevant to the question.
- Augmentation: Paste those chunks into the prompt.
- Generation: The model answers using that pasted text.
The magic is in step 1. How does a computer know which paragraphs are "relevant"? This is where embeddings come in.
Embeddings: turning meaning into numbers
An embeddingembeddingAn embedding is a numerical vector that represents data (text, images, or items) in a way that captures meaning, so similar items sit close together in space.View full definition → is a list of numbers that represents the *meaning* of a piece of text. Similar meanings get similar numbers.
"How much paid leave do I get?" and "What is the vacation policy?" use almost no words in common. But their embeddings land close together, because they mean nearly the same thing. That's the key: embeddings let you search by meaning, not by keyword.
You create embeddings by calling an embedding model (OpenAI, Google, and Cohere all offer them cheaply). You get back a long list of numbers, often 1,536 of them, for each chunk of text.
Want a visual intuition for this? Google's embedding projector lets you fly through real word embeddings in 3D: projector.tensorflow.org.
Building the handbook chatbot
Let's build it conceptually, then in code.
Step 1: Chop the handbook into chunks
You can't embed a whole 200-page document as one blob. You split it into chunks, usually a few paragraphs each (say 500 words). Each chunk becomes one searchable unit.
So our handbook becomes maybe 400 chunks: one about parental leave, one about expense reports, one about vacation accrual, and so on.
Step 2: Embed every chunk and store it
You run each chunk through the embedding model and store the resulting numbers in a vector databasevector databaseA vector database stores data as high-dimensional numeric vectors (embeddings) and retrieves items by similarity rather than exact matches, powering semantic search and AI applications.View full definition → (a database built to search by similarity). Popular ones in 2026 include Chroma, Pinecone, and pgvector. For small projects you don't even need one: a plain list in memory works.
This step happens once, ahead of time. It's called "indexing."
Step 3: Answer a question
When a user asks something:
- Embed their question.
- Find the chunks whose embeddings are closest to the question's embedding.
- Paste those chunks into the prompt.
- Send it to the model.
What is Retrieval-Augmented Generation (RAG)?
The code sketch
Here's the entire pattern in runnable Python, using OpenAI. Don't worry if you're not a coder; read the comments and you'll follow the logic.
from openai import OpenAI
import numpy as np
client = OpenAI() # uses your API key
# 1. Our "handbook," split into chunks
chunks = [
"Employees accrue 15 vacation days per year. After 3 years, this increases to 20 days.",
"Expense reports must be submitted within 30 days using the Concur portal.",
"Parental leave is 16 weeks of paid time off for all new parents.",
]
def embed(text):
# Turn text into a list of numbers (its embedding)
resp = client.embeddings.create(
model="text-embedding-3-small", input=text
)
return np.array(resp.data[0].embedding)
# 2. Embed every chunk once (the "index")
chunk_vectors = [embed(c) for c in chunks]
# 3. A user asks a question
question = "How many vacation days do I get after three years?"
q_vector = embed(question)
# 4. Find the most similar chunk (cosine similarity)
def similarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
scores = [similarity(q_vector, cv) for cv in chunk_vectors]
best_chunk = chunks[int(np.argmax(scores))]
# 5. Stuff the best chunk into the prompt
prompt = f"""Answer using ONLY the context below.
If the answer isn't there, say you don't know.
Context: {best_chunk}
Question: {question}"""
answer = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": prompt}],
)
print(answer.choices[0].message.content)What happens: the question's embedding is closest to the *vacation* chunk (not expenses, not parental leave). We paste only that chunk in, and the model answers "20 days" straight from your data.
Notice the prompt instruction: "Answer using ONLY the context below." This is what keeps the model honest. It stops it from inventing answers and grounds it in your text.
Knowledge check
1. What is the core insight behind RAG's approach to answering questions using your data?
2. Why does simply pasting a 200-page handbook into the chat window tend to be a poor solution?
3. Why can an embedding-based search match 'How much paid leave do I get?' with 'What is the vacation policy?' even though they share almost no words?
4. Select ALL of the following that are genuine limitations of a base LLM that RAG is designed to address.
Select all the correct answers.
5. Select ALL statements that correctly describe the three steps of RAG (Retrieval, Augmentation, Generation).
Select all the correct answers.
Why RAG beats the alternatives
People often ask: why not just fine-tune the model on my data instead? (Fine-tuningFine-tuningFine-tuning adapts a pre-trained model to a specific task or domain by continuing training on a smaller, targeted dataset, improving accuracy and style for that use case.View full definition → means retraining the model so your data lives inside it.)
RAG usually wins for company knowledge because:
- It updates instantly. Change a policy? Re-embed one chunk. With fine-tuning, you'd retrain.
- It cites sources. You know exactly which chunk produced the answer, so you can show "from page 14 of the handbook."
- It's cheaper and faster to set up.
- It reduces hallucination. The model answers from text in front of it, not from fuzzy memory.
Fine-tuning is better for changing the model's *style or behavior* (always answer in legal language, always use our brand voice). RAG is better for giving it *facts*. Many real systems use both.
Where RAG goes wrong (and how to fix it)
RAG isn't magic. The most common failures:
Bad chunking. If you split a sentence in half, retrieval breaks. Fix: chunk by paragraph or section, and let chunks overlap slightly so context isn't lost at the edges.
Retrieving the wrong chunks. If the question is vague, you might pull irrelevant text. Fix: retrieve the top 3 to 5 chunks instead of just one, giving the model more to work with.
The answer spans many chunks. "Compare our 2024 and 2025 travel policies" needs two sections. Fix: retrieve more chunks and let the model synthesize.
Stale index. You updated the handbook but forgot to re-embed. The bot gives old answers. Fix: re-index on a schedule.
You don't always have to build it
In 2026, you can get RAG without writing the code above:
- ChatGPT and Claude let you upload files to a Project, and they retrieve from those files automatically.
- Custom GPTs and Claude Projects let you build a handbook bot by uploading documents, no code at all.
- Google's NotebookLM is RAG in a friendly wrapper: upload sources, ask questions, get answers with citations back to the exact passage.
Use these for small or personal use. Build the pipelinepipelineAll active sales opportunities across the stages of the sales process, together with their combined potential value and probability of closing.View full definition → yourself when you need control, scale, private hosting, or integration into a product.
Key Takeaways
- RAG = retrieve relevant chunks, paste them into the prompt, let the model answer. It grounds the model in *your* data instead of the open internet.
- Embeddings are the engine. They turn text into numbers so you can search by meaning, matching "vacation policy" to "paid leave" even with no shared words.
- Always instruct the model to answer only from the provided context. This single line dramatically cuts hallucinationhallucinationA hallucination is when an AI model generates output that is fluent and confident but factually wrong, fabricated, or unsupported by its source data.View full definition → and lets you trace answers to a source.
- Start no-code. Use NotebookLM, a Custom GPT, or a Claude Project to prototype your handbook bot today, then move to a real vector database (Chroma, pgvector) when you outgrow it.
- Chunking quality decides retrieval quality. Split by section, allow slight overlap, and re-index whenever your documents change.
What to do, from this lesson
These actions are compiled in the role's Playbook.
- Skip RAG and paste small document sets into the prompt
Related articles
Recent articles from the blog that build on this lesson.
- AIData privacy when everything goes to a model: the blind spots your legal team isn't catchingOrganizations are rushing to deploy LLMs while treating data privacy as a compliance checkbox. The real exposure lies deeper, in architectural choices and behavioral patterns that most governance frameworks haven't caught up with yet.
- AIRight context, wrong assumption: what Morgan Stanley learned about prompting at scaleMorgan Stanley's deployment of an AI assistant for its financial advisors exposed a problem most teams overlook: feeding the model more information does not produce better answers. The real discipline is selecting which context matters, and why that distinction changes how you build prompts entirely.
- AIBloomberg's bet on fine-tuning: what it teaches every enterprise about the RAG-vs-fine-tune decisionBloomberg built a domain-specific large language model from scratch rather than retrieving over generic ones, and the results clarified a decision that still confuses most enterprise AI teams. The logic behind that choice, and where it breaks down for other organizations, is more instructive than the model itself.