RAG explained simply: how an AI assistant reads your documents

Retrieval Augmented Generation for decision-makers: how an AI assistant reads your documents, where RAG reaches its limits, and what the costs consist of.

Hand-drawn sketch: a magnifying glass hovering over a stack of documents, the lens filled in teal.

RAG stands for Retrieval Augmented Generation and describes how an AI assistant reads your own documents before it answers. It first searches for the matching passages, feeds them to the language model, and lets it formulate an answer only on that basis. This allows the assistant to say where an answer came from. Almost every tool that lets you “chat with documents” works this way, from Copilot to a self-built knowledge agent. This article explains the principle for decision-makers, without code, and shows where RAG reaches its limits. As of: September 2026, information on Microsoft according to the linked pages on Microsoft Learn.

In short

RAG means: search first, then answer, with a source. Quality depends less on the language model than on your documents, their currency, and the permissions.

What is Retrieval Augmented Generation?

Retrieval Augmented Generation is a method in which a language model looks up a defined knowledge source before every answer. The term comes from a research paper by Lewis and others from 2020, created in Facebook's research division. The idea behind it: a language model only knows what was in its training data. It doesn't know your price list, your quality manual, or last month's work instruction. RAG provides it with these documents at the moment of the question.

The difference from a plain chatbot is practical: a model without RAG answers from memory and always sounds convinced while doing so. A RAG assistant answers from the documents you give it and can show the passage.

How does RAG work? A picture for the mind

RAG works like a good caseworker with a filing cabinet. She hears the question, pulls the two or three matching files, reads the relevant passages, and then answers with a reference to the file. Microsoft describes the same three steps in its Learn documentation on RAG (as of September 2026):

  • Search (Retrieve): The system searches your documents for passages that match the question. To do this, the documents are first broken down into sections and prepared so that not only identical words are found, but also identical meaning.
  • Augment: The passages found are passed to the language model together with the question.
  • Generate: The model formulates the answer based on these passages and cites the source.

An example: an employee asks “How long is the return period for custom-made items?”. The system finds the paragraph in the terms and conditions and the exception in the customer service manual, the model summarizes both in two sentences and links both spots. If the answer is nowhere to be found, a well-configured assistant should say so instead of guessing.

Where is RAG already in use today?

RAG is not a product of its own but the pattern behind many tools. Microsoft names it explicitly: the search of company knowledge that Copilot agents use is described by Microsoft as RAG over the Microsoft Graph (Billing rates, Microsoft Learn, as of September 2026). Dataverse as a knowledge source in Copilot Studio also works, according to Microsoft Learn, with a RAG method. Google's NotebookLM also answers questions only from the uploaded sources and shows the citations, more on this in NotebookLM in the enterprise.

For you as a decision-maker, this means: the question is rarely “RAG, yes or no”, but rather “which tool, on which documents, for whom”.

Where does RAG reach its limits?

RAG is only as good as what it finds. In our projects, it rarely fails because of the model, but because of four things:

  • Poor documents: Scanned PDFs without text, tables as images, five versions of the same manual in the same folder. The system then finds the wrong passage and answers fluently regardless.
  • Outdated versions: If the old price list sits next to the new one, the system doesn't know which one applies. Tidying up and clear titles help more than any setting.
  • Permissions: An assistant may only find what the person asking is allowed to see. Microsoft names security and access control as one of the core challenges of RAG (Microsoft Learn on RAG with Azure Search, as of September 2026). With tools in the Microsoft 365 environment, the assistant inherits the SharePoint permissions. If these are too generous, the problem becomes visible, see Oversharing and SharePoint permissions.
  • Questions outside the documents: RAG does not replace a database query. “How many orders did we have in March?” cannot be answered by a manual. For that, the assistant needs a connection to the relevant system.

What does a RAG assistant cost?

The cost of a RAG assistant consists of three parts: the one-time setup, the ongoing operation per question, and the maintenance of the documents. There is no serious figure without looking at your data, but the logic is always the same.

  • Setup: Define the knowledge area, tidy up documents, set up the assistant, write test questions with target answers and check them. Tidying up is often the biggest item.
  • Operation: Every question costs computing power for search and answer. Microsoft points out that the passages provided lengthen the input to the model and thus increase costs (Microsoft Learn). In Microsoft 365, this runs via licenses or Copilot Credits; with your own technology, via the use of the language model and the server.
  • Maintenance: Someone has to upload new versions and remove old ones. Without this role, every assistant becomes outdated.

For your technical team: build RAG yourself

If your team wants to set up a RAG assistant itself, for example with n8n on your own server, we have three technical guides:

Try RAG in your own operation

The fastest way to find out whether RAG works for you is with a knowledge area and real questions. This is exactly how the knowledge agent pilot is set up: a knowledge area, for example a manual, product data, or policies, implementation on SharePoint with Copilot Studio or with n8n on your own server, test questions with target answers, test log, and handover. The pilot costs 5,000 euros net fixed price for five days, the expansion afterward 1,000 euros net per day. All prices net, plus travel costs and expenses for on-site appointments. No hidden hours. What an internal knowledge assistant looks like in everyday use is described in the article Internal knowledge assistant. Whether your documents are already good enough for this, we discuss in the free initial consultation, 30 minutes.

Frequently asked questions about RAG

What does RAG mean in AI?

RAG stands for Retrieval Augmented Generation. An AI assistant searches your documents for matching passages before answering and formulates the answer on this basis, citing the source.

Does a RAG assistant no longer invent answers?

Less often, but not never. If the system finds the wrong passage or none at all, the model can still formulate an answer. That's why test questions with target answers and a clear instruction to the assistant belong to every setup: if the answer is not in the documents, it says so openly instead of guessing.

Do I have to train my own model for RAG?

No. RAG uses an existing language model and provides it with your documents at runtime. Your own training is not necessary; new documents take effect as soon as they are in the search index.

Do I need Microsoft 365 for RAG?

No. Microsoft 365 brings RAG via Copilot and Copilot Studio. A setup of your own works just as well, for example with n8n on your own server. If the data must not leave the premises at all, the language model also runs locally.

Simon Glowik, founder of NordFlux
About the author

Founder of NordFlux. Spent four years automating processes at enterprise scale at Dräger, and now brings that depth to the mid-market — pragmatic and with full data sovereignty.

Certifications

  • Microsoft certified — PL-900 and AZ-900
  • UiPath certified — Automation Developer Associate
  • UiPath zertifiziert — Automation Developer Associate
All articles
Read more
→

Related posts

Free initial call

Are your documents good enough for an AI assistant?

Whether RAG works for you is decided above all by the state of your documents. We'll look at a knowledge area together and tell you what needs to be tidied up before the start.

  • Knowledge area and typical questions defined
  • Document status and permissions checked
  • Platform chosen to match data and data protection
RAG explained simply: how AI reads your documents | NordFlux