How to build a RAG chatbot over your documents

Parse your documents to clean text, split them into overlapping chunks, embed each chunk and store it in a vector database. At question time, embed the query, retrieve the top matching chunks, and ask an LLM to answer only from them with citations. Then evaluate retrieval on real questions.
What is a RAG chatbot and when do you need one?
Retrieval-augmented generation (RAG) means the model answers using passages fetched from your own data at question time, instead of relying on what it memorized during training. The chatbot becomes a search engine plus a writer.
You need RAG when answers must come from documents the model has never seen: internal wikis, contracts, product manuals, support tickets or research notes. It also lets you update knowledge by re-indexing files rather than retraining anything.
You do not need RAG when the whole corpus fits comfortably in the model’s context window and rarely changes. In that case, putting the documents straight into the prompt is simpler and often more accurate.
A useful mental model is a librarian and a writer. The librarian (retrieval) must bring the right pages; the writer (the LLM) can only be as good as those pages. Most effort in a good RAG project goes into the librarian.
Plan for three audiences from day one: end users who want short, cited answers; administrators who need to add and remove documents; and developers who must debug why a particular answer was wrong. Each needs different screens and logs.
What does a RAG pipeline consist of?
Every RAG system has an indexing path that runs when documents change and a query path that runs on every question. Keeping them separate makes the system easier to debug.
| Stage | What it does | Open-source options | Typical failure |
|---|---|---|---|
| Parsing | Turns PDFs, HTML and Office files into clean text | Unstructured, Docling, Apache Tika | Tables and columns turned into word soup |
| Chunking | Splits text into retrievable pieces | LangChain or LlamaIndex splitters | Chunks cut mid-thought or too large |
| Embedding | Turns chunks into vectors | sentence-transformers models, BGE, nomic-embed | Model mismatched to language or domain |
| Vector store | Stores vectors and finds nearest ones | pgvector, Qdrant, Chroma, Weaviate | Missing metadata filters |
| Generation | Writes the answer from retrieved chunks | Any chat LLM, local or API | Answers beyond the retrieved context |
How to build a RAG chatbot step by step
- Collect a small, representative document set first (20–50 files) rather than indexing everything on day one.
- Parse each file to text and keep metadata: source file, page or heading, date and access group.
- Chunk by structure where possible (headings, paragraphs) with a size of a few hundred tokens and a small overlap between neighbours.
- Pick an embedding model, embed every chunk and store the vector, the chunk text and its metadata together.
- At query time, embed the user question with the same model and retrieve the top 5–10 chunks, applying metadata filters such as the user’s permissions.
- Optionally rerank the candidates with a cross-encoder reranker and keep the best few.
- Build the prompt: system instructions, the retrieved chunks with source labels, then the question; tell the model to answer only from the sources and cite them.
- Return the answer plus clickable source references so users can verify it.
- Write 30–100 real questions with expected sources and re-run them after every change to chunking, embeddings or prompts.
Which framework or stack should you pick?
You can write RAG in plain code with an embedding call, a SQL query and a chat call. Frameworks help when you need many loaders, advanced retrievers or agent-style tool use, but they add abstraction you have to learn.
A pragmatic default for a web app is Postgres with pgvector, because your documents, users and permissions already live in one database. Reach for a dedicated vector database when your collection grows large or you need features like built-in hybrid search and sharding.
| Approach | Best for | Trade-off |
|---|---|---|
| Plain code + pgvector | Small teams, existing Postgres apps | You write loaders and evaluation yourself |
| LlamaIndex | Document-heavy retrieval with many data connectors | Framework concepts to learn |
| LangChain / LangGraph | RAG mixed with agents and tool calls | Large API surface, frequent changes |
| Haystack | Explicit, production-style pipelines | More setup than a script |
| Dify, Flowise, AnythingLLM | No-code or low-code RAG apps | Less control over each stage |
How do you stop the chatbot from making things up?
Hallucinations in RAG are usually retrieval failures in disguise: the right passage never reached the prompt, so the model improvised. Fix retrieval first, then tighten generation.
- Instruct the model to say it does not know when the sources do not contain the answer, and test that behaviour explicitly.
- Require citations per claim and drop answers whose citations do not match retrieved chunks.
- Use hybrid search (keyword plus vector) so exact names, codes and IDs are found reliably.
- Keep chunk text readable: include the document title and section heading inside each chunk.
- Log every question, retrieved chunk and answer so you can see where a bad answer came from.
Common mistakes when building RAG
- Evaluating by vibes: without a fixed question set you cannot tell whether a change helped.
- Indexing without permissions: a chatbot that retrieves HR documents for every employee is a data leak.
- Changing the embedding model without re-embedding everything; old and new vectors are not comparable.
- Stuffing 30 chunks into the prompt; more context often means more distraction and higher cost.
- Ignoring document updates: stale chunks keep answering after the source file changed or was deleted.
- Skipping the parsing check: always read a sample of parsed text before you blame the model.
How much does a RAG chatbot cost to run?
Costs split into one-off indexing (embedding every chunk), storage, and per-question inference. Embedding is cheap relative to generation, so the recurring cost is dominated by how many tokens you send to the chat model per answer.
You control that by retrieving fewer, better chunks, caching answers to repeated questions and using a smaller model for simple queries. A fully local stack with an open embedding model and a self-hosted LLM moves the cost to hardware instead of per-token fees.
RepoLoot’s catalog lists RAG projects with notes on difficulty and business value, which is a quick way to find a starting codebase rather than wiring every stage yourself.
How do you keep the index in sync with changing documents?
Documents change, and a RAG index that is rebuilt by hand drifts quickly out of date. Build indexing as a repeatable job from the start, not as a one-off notebook.
Give every document a stable ID and a content hash. When the job runs, it can skip unchanged files, re-chunk and re-embed changed ones, and remove chunks belonging to deleted files.
- Trigger indexing from the source system where possible, such as a webhook when a wiki page is saved.
- Run a scheduled full reconciliation as a safety net for missed events.
- Store the parser, chunker and embedding model versions with each chunk.
- Show the last indexed date in the chatbot’s source citations so users can judge freshness.
- Alert when indexing fails, because a silent failure looks exactly like a working but stale bot.
How do you add permissions and multi-tenant data safely?
The most serious RAG bugs are not wrong answers but leaked ones. If a user cannot open a document in your app, the chatbot must not quote it either.
Store the access information with every chunk and filter at retrieval time, inside the vector query itself. Filtering after generation is too late, because the model has already read the restricted text.
- Attach tenant ID, access group and document ID as metadata on every chunk during indexing.
- Pass the authenticated user’s groups into every retrieval query as a mandatory filter.
- When a document’s permissions change, update the metadata on all its chunks immediately.
- When a document is deleted, delete its chunks in the same transaction or job.
- Log which chunks were shown to which user, so you can audit an incident later.
Frequently asked questions
- What chunk size should I use for RAG?
- Start with chunks of a few hundred tokens with a small overlap and split along headings or paragraphs. Then test with your question set: if answers miss context, grow chunks or retrieve neighbours; if retrieval returns loosely related text, shrink them. There is no universal best size across document types.
- Do I need a vector database for RAG?
- Not necessarily a separate one. For small and medium collections, pgvector inside Postgres or even an in-process store like Chroma is enough. Dedicated vector databases such as Qdrant, Weaviate or Milvus become worthwhile at larger scale or when you need their filtering, sharding and hybrid search features.
- Can a RAG chatbot run fully offline?
- Yes. Use an open embedding model through sentence-transformers or Ollama, a local vector store, and a local chat model served by Ollama or llama.cpp. Quality depends mostly on the chat model size your hardware can run, so test answers against your question set before committing.
- Is RAG better than fine-tuning?
- They solve different problems. RAG injects up-to-date facts at question time and can cite sources, which suits changing documents. Fine-tuning changes style, format or narrow behaviour but is poor at adding reliable facts. Many teams use RAG first and only fine-tune when output format or tone still falls short.