RAG: a complete overview

How retrieval-augmented generation works: eight steps, RAG compared with fine-tuning and a long prompt, pros and cons, three steps in code, mistakes and rules of a reliable system.

Stack and technologies Updated

In short

RAG (retrieval-augmented generation) is a way to make a language model answer from your own materials instead of its general memory. Documents are cut into pieces, each piece gets a vector — a numeric fingerprint of its meaning — and is stored in a database. When a question comes, the closest pieces are found by meaning and handed to the model together with the question, and the model answers only from them, citing the sources. It is how AI assistants answer about prices, terms and products without inventing them. If nothing relevant is found, a good RAG system says so and hands over to a person.

RAG at a glance

The main facts in one table — what RAG is, what it is made of and where it is used.

What it is
Search through your materials plus an answer from the model based on what was found
The term
Introduced in a 2020 paper by Facebook AI Research
Parts
Chunking, an embedding model, a vector index, a language model
Where vectors live
PostgreSQL with pgvector, SQLite with an extension or a separate vector database
Models
Claude, GPT or local models — chosen for the task and the data
Updates
A changed document is re-indexed — no retraining
Main risk
Poor search: the model answers confidently from the wrong pieces

How RAG works: 8 steps

The first four steps are done once and repeated when documents change; the last four run on every question.

StepWhat happensWhat decides quality
1. Collect documents, catalogue, terms, answers to frequent questions up-to-date and non-contradictory sources
2. Cut pieces along headings and paragraphs one thought per piece, with its heading
3. Vectorise each piece gets a vector of meaning an embedding model that knows the language
4. Store pieces and vectors go into the database source and date next to each piece
5. Find the closest pieces by meaning and by words a similarity threshold, filters by section
6. Rerank a second model reorders the candidates optional, but improves precision
7. Answer the model writes only from the numbered pieces a strict instruction and citations
8. Log question, pieces found, answer, rating reviewing the log and filling gaps

RAG, fine-tuning or a long prompt

Three ways to give a model your knowledge. They solve different problems and are often combined.

CriterionRAGFine-tuningEverything in the prompt
Good for facts that change style and format of answers a small fixed set of texts
Updating knowledge re-index a document train again edit the prompt
Sources in the answer yes no partly
Volume of knowledge practically unlimited limited by training data limited by the context window
Cost per question low: only the pieces found low high: the whole text each time
Access rights filters at search time impossible a separate prompt per role
Start days weeks and a dataset hours

Pros and cons of RAG

RAG makes the answer only as good as the search. Both the strengths and the weaknesses come from that.

Pros · 5

  • Answers from your data

    Prices, terms and specifications come from your documents, not from the model’s memory.

  • Sources you can check

    Every statement carries a link to the piece it came from.

  • Fresh without retraining

    A changed price list is re-indexed in seconds.

  • Rights at search time

    A customer and a manager see answers from different sets of documents.

  • Any model

    The model can be replaced without rebuilding the knowledge base.

Cons · 4

  • Search decides everything

    If the right piece is not found, even the best model answers badly.

  • Garbage in the base

    Outdated and contradictory documents turn into confident wrong answers.

  • Questions across the whole base

    “How many orders were there in March” is a query to a database, not a search through texts.

  • Needs care

    The log has to be reviewed and gaps in the knowledge filled.

What RAG looks like in code: 3 steps

Cutting, indexing and the answer in TypeScript with PostgreSQL and pgvector. The pipeline was run end to end on a test knowledge base of a furniture shop.

Cutting into pieces

Pieces follow headings and paragraphs, so one piece holds one thought and knows where it came from.

chunk.ts
// Split a document into pieces of up to ~800 characters along paragraph borders.
// Every piece keeps its source and heading — the answer will cite them.
export type Chunk = { source: string; heading: string; text: string };

export function chunk(source: string, markdown: string, maxChars = 800): Chunk[] {
  const chunks: Chunk[] = [];
  let heading = '';
  let buffer: string[] = [];

  const flush = () => {
    const text = buffer.join('\n\n').trim();
    if (text) chunks.push({ source, heading, text });
    buffer = [];
  };

  for (const block of markdown.split(/\n{2,}/)) {
    if (block.startsWith('#')) {
      flush(); // a new heading always starts a new piece
      heading = block.replace(/^#+\s*/, '');
      continue;
    }
    if (buffer.join('\n\n').length + block.length > maxChars) flush();
    buffer.push(block);
  }
  flush();
  return chunks;
}

Indexing

Vectors from any OpenAI-compatible API; re-indexing a document replaces its pieces in one statement, without duplicates.

index.ts
// Indexing: every piece gets a vector and goes into PostgreSQL with pgvector.
// CREATE TABLE chunks (id bigserial PRIMARY KEY, source text, heading text,
//                      text text, embedding vector(384));
import pg from 'pg';
import { chunk } from './chunk.ts';

const db = new pg.Pool({ connectionString: process.env.DATABASE_URL });

// Any OpenAI-compatible embeddings API — a cloud provider or a local model
export async function embed(texts: string[]): Promise<number[][]> {
  const res = await fetch(`${process.env.EMBEDDINGS_URL}/v1/embeddings`, {
    method: 'POST',
    headers: {
      'Content-Type': 'application/json',
      Authorization: `Bearer ${process.env.EMBEDDINGS_KEY}`,
    },
    body: JSON.stringify({ model: process.env.EMBEDDINGS_MODEL, input: texts }),
  });
  if (!res.ok) throw new Error(`embeddings: HTTP ${res.status}`);
  const { data } = (await res.json()) as { data: { embedding: number[] }[] };
  return data.map((d) => d.embedding);
}

export async function indexDocument(source: string, markdown: string): Promise<number> {
  const pieces = chunk(source, markdown);
  const vectors = await embed(pieces.map((p) => `${p.heading}\n${p.text}`));
  // One statement: the old pieces of the document are replaced atomically
  await db.query(
    `WITH old AS (DELETE FROM chunks WHERE source = $1)
     INSERT INTO chunks (source, heading, text, embedding)
     SELECT $1, h, t, e::vector FROM unnest($2::text[], $3::text[], $4::text[]) AS u(h, t, e)`,
    [source, pieces.map((p) => p.heading), pieces.map((p) => p.text), vectors.map((v) => JSON.stringify(v))],
  );
  return pieces.length;
}

The answer with sources

In the test, questions about delivery, returns and warranty found the right piece with a similarity of 0.52–0.60; a question about the weather scored 0.14 and never reached the model.

ask.ts
// The answer: find the closest pieces, give them to the model numbered,
// require citations — and do not call the model at all if nothing was found
import pg from 'pg';
import { embed } from './index.ts';

const db = new pg.Pool({ connectionString: process.env.DATABASE_URL });
const API = process.env.ANTHROPIC_BASE_URL ?? 'https://api.anthropic.com';

export async function ask(question: string) {
  const [q] = await embed([question]);
  const { rows } = await db.query(
    `SELECT source, heading, text, 1 - (embedding <=> $1) AS score
       FROM chunks ORDER BY embedding <=> $1 LIMIT 4`,
    [JSON.stringify(q)],
  );
  // The threshold is chosen on your own questions; below it — not an answer
  const found = rows.filter((r) => r.score > 0.5);
  if (found.length === 0) return { answer: null, sources: [] }; // hand over to a person

  const context = found.map((r, i) => `[${i + 1}] ${r.heading}\n${r.text}`).join('\n\n');
  const res = await fetch(`${API}/v1/messages`, {
    method: 'POST',
    headers: {
      'content-type': 'application/json',
      'x-api-key': process.env.ANTHROPIC_API_KEY ?? '',
      'anthropic-version': '2023-06-01',
    },
    body: JSON.stringify({
      model: 'claude-sonnet-5',
      max_tokens: 600,
      system:
        'Answer only from the sources below and cite them as [1], [2]. ' +
        'If the sources do not contain the answer, say so.\n\n' + context,
      messages: [{ role: 'user', content: question }],
    }),
  });
  if (!res.ok) throw new Error(`model: HTTP ${res.status}`);
  const data = (await res.json()) as { content: { type: string; text: string }[] };
  return { answer: data.content[0].text, sources: found.map((r) => r.source) };
}

Common mistakes in RAG projects

  1. Pieces of a fixed length

    Cutting every 500 characters splits a rule in half, and neither half answers the question.

  2. No threshold

    The nearest piece is always found, even for a question about the weather — and the model builds an answer on it.

  3. Only vector search

    Article numbers, model names and codes are found better by words — the two searches are combined.

  4. An English-only embedding model

    Questions in other languages find nothing; the model must know the languages of your customers.

  5. Indexing once and forgetting

    The price list changed, the base did not — the assistant quotes old prices.

  6. No set of test questions

    Without 30–50 real questions with expected answers, any change is a guess.

7 rules of a reliable RAG

  1. 01

    Clean the sources first

    Remove outdated versions and contradictions before indexing.

  2. 02

    Pieces by meaning

    Headings and paragraphs as borders, the heading inside each piece.

  3. 03

    Two searches together

    By meaning and by words; the results are merged.

  4. 04

    A threshold and an honest “I don’t know”

    Below the threshold the model is not called — the question goes to a person.

  5. 05

    Citations are mandatory

    An answer without a source is not shown.

  6. 06

    Rights in the query

    Filters by role are applied at search time, not in the instruction to the model.

  7. 07

    Measure on real questions

    A fixed set of questions is run after every change to the base or the settings.

Questions about RAG

What is RAG in simple words?

First find the answer in your materials, then let the model formulate it — with a link to the source.

Does RAG stop hallucinations?

It reduces them sharply when there is a threshold, a strict instruction and citations; without them it does not.

Do I need a vector database?

Usually not a separate one: pgvector in PostgreSQL holds millions of pieces next to the rest of the data.

Is my data used to train the model?

Not with business API access to Claude or GPT; for the strictest cases there are local models.

What documents can be used?

Texts, tables, PDFs, pages of the site, the catalogue; scans need text recognition first.

How is quality measured?

On a set of real questions: was the right piece found, and is the answer correct and backed by a source.

How does RAG relate to MCP?

RAG searches through texts; MCP gives the model tools — orders, stock, CRM. Assistants often use both.

Online form

An assistant
on your data

I build AI assistants that answer from your documents, catalogue and rules — with links to sources and a handover to a person when the answer is not there. Tell me about the task — I answer within one working day.

Or write to [email protected]