RAG and Enterprise Knowledge — Executive Mental Model
A model knows the world as it was when training stopped; it has never seen your company. Retrieval-augmented generation closes that gap by fetching the relevant, current and permitted pages at the moment of each question. Retrieve, don't retrain: your systems own the facts, and the model writes the answer.
After this chapter you can
- Explain retrieval-augmented generation as retrieve first, generate second.
- Distinguish model knowledge from enterprise knowledge, and retrieval from fine-tuning and training.
- Explain why permissions and version control decide whether an answer can be trusted.
- Match each knowledge gap to the right fix - retrieval, a live query or fine-tuning.
- Name the questions that decide whether a knowledge assistant deserves approval.
In 2020, researchers from Facebook AI Research, University College London and New York University ran a small experiment worth knowing for any executive who funds an AI assistant. Their system answered questions by first searching a copy of Wikipedia, then writing an answer from what it found. They drew up a list of 82 world leaders who had changed between December 2016 and December 2018 and asked the same kind of question about each: “Who is the President of Peru?”
With a December 2016 copy of Wikipedia, the system named the 2016 leaders correctly 70 percent of the time. Before reading on, predict how it did on the 2018 leaders with that same 2016 copy. The answer was 4 percent. Then the researchers did something simple. They did not retrain anything. They swapped the 2016 copy for the 2018 one, and accuracy on the 2018 leaders rose to 68 percent1.
That paper named the pattern retrieval-augmented generation, or RAG. Its lesson fits in three words that this chapter will keep returning to: retrieve, don’t retrain. The knowledge that changes should live where it can be changed, and the model should be handed it at the moment of the question.
Two kinds of knowledge
Every enterprise assistant combines two kinds of knowledge, and keeping them apart makes most decisions about it easier.
Model knowledge is what the model learned in training: language, general facts and patterns of reasoning. It lives inside the model’s parameters. It is current only to the date training stopped, it is the same for every user, and the only way to change it is to train again. Lewis and colleagues called this parametric memory. Enterprise knowledge is what your organization keeps: policies, manuals, contracts, product documentation, customer records. It lives in your systems, it changes every week, much of it is confidential, and the right answer often depends on who is asking. You change it by editing the source. The researchers called this non-parametric memory, and their point was that the second kind can be updated without touching the first1.
RAG is the discipline of bringing the second kind to the first, one question at a time. The system retrieves relevant information, places it in the model’s context beside the question, and only then asks the model to generate. As Prompting and Context Engineering showed, a model can only be as good as what it receives; retrieval is how much of an enterprise’s context arrives. A useful assistant is useful not because the model became the archive, but because retrieval brought the archive to the question.
How an answer is assembled
A RAG system works in two phases. The first happens before anyone asks. Teams collect the documents, clean them, split long ones into passages and store each passage with its metadata: owner, date, version, department and access rules. Splitting matters because retrieval should return the relevant part, not a five-hundred-page manual. Lewis’s system cut Wikipedia into 21 million passages of 100 words each1.
The second phase runs for every question. The question arrives with the identity of the person asking. The system searches only what that person may see, picks the best few passages, puts them into the model’s context and asks for an answer drawn from them, with references back to the passages it used.
The search itself is a design choice, not a product. Vector search finds passages by meaning, using the embeddings that Tokens, Context and Embeddings explained. Keyword search catches exact strings such as product codes, clause numbers and error codes, and on a broad benchmark of 18 datasets plain keyword ranking proved a robust baseline that embedding-based retrievers often failed to beat on unfamiliar material2. Metadata filters narrow by version, region or department, and a reranking step puts the best passage first. Good systems combine these. You do not need to design the search, but someone must, and the quality of that design shows in every answer.
Permissions come before search
One of the most important design decisions in an enterprise assistant is also the least visible: what the model is allowed to receive. A frontline employee must not be handed an executive-only memo because it happens to match the question. The filter has to run before retrieval, using the identity of the person asking, so that the model only ever sees what that person could have opened directly.
Hiding a sensitive document from the final answer is not enough. Once a passage is in the model’s context, it can be paraphrased, summarized or leaked by a cleverly worded follow-up. The OWASP list of the top risks for applications built on language models treats this as a category of its own, warning that in a shared index, content from one group can be retrieved for another group’s questions, and recommending permission-aware stores with strict partitioning3. The practical test for a technology leader is blunt: does the search index carry each source system’s permissions, and is it filtered on the user’s identity at query time? If a document’s permissions cannot be enforced in retrieval, it should not be indexed yet.
Retrieved text is also data, not instructions. A planted document can contain sentences that try to steer the model, which researchers have demonstrated against real applications4; Security and AI Attacks covers that threat in full.
Old and new versions in the index
The hot-swap experiment has a mirror image that causes more damage in practice. Lewis’s system gave the wrong leaders because it held only the old pages. An enterprise index usually holds both: last year’s travel policy and this year’s, the withdrawn price list and the current one, a draft procedure and the approved version. To a search engine looking for passages that match the question, they look almost identical. The old one may even match better, because it has been quoted, copied and discussed more often.
The remedy is unglamorous. Every document carries an effective date, an owner and a status. Superseded versions are retired from the index or flagged, not left to compete. Each topic has one authoritative source, and the official policy outranks a forum thread about it. Deciding which source is authoritative is a knowledge-management question, which AI in Knowledge Management takes up; the retrieval design simply has to enforce the answer. When the assistant cannot find a current source, it should say so rather than answer from whatever it found.
Evidence, not truth
RAG improves the evidence a model works from. It does not guarantee the answer. A model cannot use a passage that retrieval failed to find, and handed the wrong passage, it writes a fluent, wrong answer. Even with the right passage, it can misread it, merge two conflicting sources, or add claims the source never made. The research literature separates answers that contradict their source from answers that go beyond it, and grounded systems can still produce both5.
Even purpose-built legal research tools that retrieve from curated case law still hallucinated in 17 to 33 percent of answers in an independent Stanford study6. Accuracy, Hallucination and Reliability treats the broader problem. More context is not a cure either: models use information at the start and end of a long input better than information in the middle7, so stuffing in fifty passages can bury the one that matters.
Two habits follow. Treat a citation as a claim to be checked, not as proof; a well-designed system verifies that the cited passage supports the sentence and never shows a source that does not. And evaluate retrieval and answers together, on a fixed set of real questions, every time the documents, the search or the model change.
Retrieve, don’t retrain
Three terms are routinely confused. Training sets a model’s parameters. Fine-tuning trains those parameters further on your examples. Retrieval adds information at the moment of the question and changes no parameters at all. As Training vs Inference explained, only the first two change the model.
When an assistant gets facts wrong, the instinctive proposal is to train a model on the company’s documents so that it “knows the business”. The evidence points the other way. In a direct comparison, retrieval consistently beat unsupervised fine-tuning at supplying knowledge, both for facts the model had partly met in training and for facts that were entirely new8. Fine-tuning on new facts is also slow and risky: models learn such examples much more slowly than familiar ones, and as they do, they become more prone to making things up9.
The business case is just as clear. A policy that changes on Friday is re-indexed on Friday; a fine-tuned model would need another training run, another round of testing, and would still hold the old version somewhere in its parameters. Anything trained into a model is visible to everyone who uses it, so it can no longer be restricted by role. And no fine-tuned answer can point to the page it came from.
Not all enterprise knowledge sits in documents. A customer’s balance or an order’s status lives in a system of record, and the right move is to query that system at question time and put the result in context, not to copy records into a document index. Fine-tuning keeps a real role, for behavior rather than facts: a required format, a house tone, a specialized task the model handles poorly. The two can be combined. None of it is free, either: indexing, search, evaluation and extra tokens per question all cost money, so judge an assistant by total business value, not by the novelty of its architecture.
Story: the government chatbot that was not launched
The United Kingdom’s Government Digital Service (GDS), the team that runs GOV.UK, built an experimental chatbot called GOV.UK Chat. It was a textbook RAG system: it retrieved passages from published GOV.UK pages, with pages containing personal data excluded, and asked a language model to answer from them. After testing with a dozen users, the team invited 1,000 people into a private pilot.
The results, published in January 2024, put the team in front of a real dilemma. Nearly 70 percent of users surveyed found the answers useful and just under 65 percent were satisfied. But the answers “did not reach the highest level of accuracy demanded” for a site like GOV.UK, and some contained hallucinations: incorrect information stated as fact. Some users dismissed the risk of error precisely because the answers carried the GOV.UK brand. Some questions failed because the relevant page was too long for retrieval to use well10.
Before reading on, decide. Would you launch, with a clear warning that answers may be wrong, while most users already like it? Or hold the launch and fix the retrieval underneath?
GDS held the launch. The team worked on how pages were split into passages, on search relevance, and on helping users ask clear questions10. Over the next two public pilots, 10,136 people asked 23,838 questions on the web, and 641 people asked 2,670 questions in the GOV.UK app. Accuracy, counted only when an answer met all the standards of published GOV.UK content, rose from 76 percent at the first benchmark to 90 percent. GDS credits that gain to its data scientists’ work and to advances in the underlying models; by the end, the service ran on a different provider’s models. When the assistant was unsure, it learned to ask a clarifying question instead of guessing, and it answered 88 percent of in-scope questions, a separate measure from accuracy: one counts how often it gave an answer, the other how often that answer met the standard. In March 2026, GDS announced that rollout would begin in the GOV.UK app, with website testing later in 2026; it still does not handle personal information11.
The model was not the only lever, and it was the part that got swapped. What the team kept and built was the retrieval, the scope and the measurement around it. And the usefulness score barely moved: nearly 70 percent among survey respondents in the first private pilot, 73 percent among app users in the last. The two groups differ, so the comparison is rough, but measured accuracy climbed from 76 to 90 percent over the same period. Users could not see the difference that mattered most. That is why a leader cannot approve a knowledge assistant on satisfaction scores alone.
What this means for leaders
Treat a knowledge assistant as an enterprise knowledge architecture, not as a model purchase. The model is the most replaceable part. The durable assets are authoritative sources with owners and dates, a search that respects permissions, and an evaluation set of real questions that is rerun after every change. When an answer is wrong, look first at what was retrieved, then at the model.
Be wary of two proposals that sound like progress. The first is to train a model on every company document so that it knows the business; it buys stale facts, blurred versions and secrets that no permission rule can reach. The second is to launch because users like the answers; GOV.UK Chat’s users liked it before it was accurate.
Check yourself
- RAG gives the model access to the whole enterprise.
- Swapping an old index for a new one can update what a RAG system knows, with no retraining.
- If an answer cites a source, it must be correct.
- Fine-tuning a model on company documents is the most reliable way to keep its facts current.
- High user satisfaction shows that a knowledge assistant is accurate.
- Permissions should be applied after the model has drafted its answer.
Reflection: whose pages, which version?
What comes next
You now have the core architecture of an enterprise assistant: a model, a prompt, a designed context, and enterprise knowledge retrieved at the moment of each question. So far, though, the assistant only answers. The next chapter, From AI Assistants to AI Agents, asks what changes when it must decide what to do, call a tool, check the result and keep going until the task is done.
References
- Patrick Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020 (arXiv:2005.11401). 2020.
- Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava and Iryna Gurevych. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. NeurIPS 2021 Datasets and Benchmarks Track. 2021.
- OWASP Foundation. OWASP Top 10 for LLM Applications 2025. OWASP GenAI Security Project. 2024.
- Kai Greshake et al. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (arXiv:2302.12173). 2023.
- Ziwei Ji et al. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12). 2023.
- Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Journal of Empirical Legal Studies 22(2), 216-242. 2025.
- Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 12. 2024.
- Oded Ovadia, Menachem Brief, Moshik Mishaeli and Oren Elisha. Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024), 237-250. 2024.
- Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart and Jonathan Herzig. Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?. EMNLP 2024 (arXiv:2405.05904). 2024.
- Government Digital Service (UK). The findings of our first generative AI experiment: GOV.UK Chat. Inside GOV.UK blog. 2024.
- Sam Dub and Sharon McDonald (Government Digital Service). 5 things we learned testing GOV.UK Chat, an AI assistant for government. Inside GOV.UK blog. 2026.
Further reading
- Patrick Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020 (arXiv:2005.11401). 2020.
- Oded Ovadia, Menachem Brief, Moshik Mishaeli and Oren Elisha. Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024), 237-250. 2024.
- Sam Dub and Sharon McDonald (Government Digital Service). 5 things we learned testing GOV.UK Chat, an AI assistant for government. Inside GOV.UK blog. 2026.
Sources last verified 2026-10-08.