Tech & AI

RAG

How AI tools look things up before they answer

Stands for
Retrieval-augmented generation
Part of speech
noun, also used as an adjective
Register
Technical
Tone
Neutral
Seen on
Developer docs, product pages, engineering blogs
In use since
Named in a 2020 research paper, common from 2023

What it means

RAG, short for retrieval-augmented generation, is a technique where an AI system first searches a set of documents for relevant information, then writes its answer using what it found.

A large language model learns from a vast amount of text, and then its knowledge is frozen. It doesn’t know your company’s refund policy or a price that changed yesterday. Asked about them, it either admits it doesn’t know or, worse, guesses.

RAG adds a lookup step. Ahead of time, documents are cut into short chunks. Each chunk becomes an embedding, a long list of numbers that captures roughly what the text means, so similar passages get similar numbers. Your question gets the same treatment. A vector search then finds the chunks whose numbers sit closest to it. Those chunks are pasted in next to your question, and the model answers from them.

The easiest comparison is an exam. A plain model sits a closed-book exam, relying on patchy memory. RAG makes it open-book, so the model can find the right page first. But open-book students still turn to the wrong chapter, misread a sentence or pad an answer with guesses. RAG has the same weak spots, which is why it cuts down hallucination without ending it.

Where it came from

The name comes from a 2020 paper by researchers at Facebook AI Research, the lab now part of Meta. It was titled “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, and its first author was Patrick Lewis. The paper paired a language model with a searchable index of Wikipedia, and looking passages up first helped on factual questions.

The basic idea wasn’t brand new, since question-answering systems had paired search with language models before. The paper supplied a clear recipe and a name, descriptive rather than catchy.

The usual account is that RAG stayed mostly a research term until 2023. Then businesses wanted chatbots that could answer questions about their own documents. Retraining a model on private data was slow, expensive and soon out of date. RAG offered a cheaper route: leave the model alone and change what it reads. “RAG pipeline” quickly became everyday developer vocabulary.

Today the term covers almost any setup that fetches information for a model before it answers, including the lookups AI agents make as they work. Here’s how it compares with other ways of getting knowledge into an AI.

Approach Where answers come from Keeping it current Can it cite sources?
Plain model What it learned in training Wait for a newer model Not reliably
Fine-tuned model Training, plus extra lessons on your data Train it again Not reliably
RAG Documents fetched when you ask Update the documents Yes, if built to

How people actually use it

  • “Our assistant uses RAG to answer from your own help centre articles.” Product page. The bot reads your content rather than winging it.
  • “The RAG pipeline keeps pulling the 2022 pricing doc. Can someone archive it?” Work chat. Used as an adjective, and a reminder that retrieval is only as good as the documents.
  • “Is this actually RAG, or did they just paste the whole PDF into the prompt?” Developer forum. Separating real retrieval from brute force.
  • “We tried fine-tuning first, then switched to RAG because our docs change every week.” Engineering blog. The classic trade-off.
  • “New: answers now come with sources, thanks to RAG.” Product update. The term as a selling point.

In a sentence

Hana: Why did the help bot tell a customer we do free returns? We stopped that in March.
Raj: The RAG index still had the old policy page. It found it and believed it.

Ellie: Couldn’t we just train the model on our wiki?
Tom: It’d be out of date by Friday. With RAG it reads the current version every time.

Priya: Is this chatbot any good?
Chris: Better than most. It’s RAG-based and links each answer to its source, so you can check.

Common misconceptions

  • “RAG means the answer is correct.” It means the answer is based on retrieved text. If that text is outdated or off-topic, so is the answer, and the model can still misread a good source.
  • “RAG retrains the model.” It doesn’t touch the model at all. It only changes what the model is shown when you ask.
  • “A citation proves the claim.” A cited page might not say what the answer claims. When it matters, click through.
  • Embedding: a list of numbers representing what a piece of text means.
  • Vector database: a store built to search embeddings quickly by closeness of meaning.
  • Chunking: splitting long documents into short passages that can be retrieved one at a time.
  • Fine-tuning: training an existing model further to change how it behaves.
  • Grounding: tying an AI’s answer to specific sources rather than to its general memory.

Questions people ask

Does RAG stop AI from making things up?

No. RAG reduces made-up answers by giving the model real sources to work from, but the model can still retrieve the wrong passage, misread the right one or fill gaps with guesses.

What is the difference between RAG and fine-tuning?

Fine-tuning changes the model itself through extra training. RAG leaves the model unchanged and hands it relevant documents each time a question is asked, so the information is easy to update.

What is a vector database?

A vector database stores embeddings, the lists of numbers that represent what each chunk of text means. It finds the chunks closest in meaning to a new question quickly, which is why many RAG systems use one.

Do longer context windows make RAG unnecessary?

Longer context windows let a model read more text at once, measured in tokens, but most document collections are still far too large to paste in whole. RAG remains useful for picking out the relevant parts and keeping costs down.

How do you pronounce RAG?

Most people say RAG as a single word, like the cloth, rather than spelling out the letters.

The short version

RAG is how many AI tools look things up before answering: find the relevant passages, add them to the prompt, then reply based on them. It keeps answers current and checkable without retraining the model, and it makes mistakes less likely rather than impossible.