You will be able to explain when retrieval is needed and how a retrieval system answers a question.
Daniel is the only HR executive at a 200-person engineering firm in Paya Lebar. Half his week goes on answering the same questions. How many days of childcare leave do I have? Can I claim a taxi home after working late? What is the process for a work-from-home day? The answers are all in the staff handbook and a folder of policy documents, about 300 pages in total. Nobody reads them. They message Daniel instead.
He wants a tool that answers those questions from the documents, with the page it found the answer on. His first idea was to paste the whole handbook into a chat assistant with every question. This lesson explains why that idea runs out of road quickly, and what replaces it.
In AI fundamentals, lesson 5.1, The context window is the model's only working memory, you met the context window, the limit on how much text a model can take in at once. Some models now accept very long inputs, and Daniel's 300 pages might fit into one of them. That does not make it a good design, for three reasons.
The first is cost. You pay for every input token on every request. If each question carries the whole handbook, Daniel pays to send 300 pages each time someone asks about taxi claims. Run the arithmetic from lesson 1.3, Tokens, pricing and limits you have to plan around, and the per-question cost is dominated by text that has nothing to do with the question.
The second is speed. A longer input takes longer to process, so every answer is slower.
The third is quality. Models do not read a very long input with equal care from start to finish. Details in the middle of a long document can be missed, as lesson 5.2 of AI fundamentals, Why long chats drift and big files get skimmed, described. A short input containing only the relevant pages tends to get a more accurate answer than a long one containing everything.
And document sets grow. Daniel's handbook is 300 pages today. Add the employee benefits guides, the safety manual and three years of HR circulars, and it no longer fits anywhere.
The alternative is to search first. When a question comes in, your code looks through the documents for the few passages most likely to contain the answer. It sends only those passages to the model, together with the question and an instruction to answer from them. The model never sees the other 295 pages.
For the taxi question, the search might return the paragraph on late-night transport from the claims policy, the overtime section of the handbook and a circular that updated the claims limit last year. That is perhaps a page of text instead of 300, and it is the right page.
This is the same thing you do yourself with a long document. You do not reread the whole handbook to answer one question. You search for "taxi", skim the hits and read the relevant section closely.
The pattern has a name. Retrieval-augmented generation, usually shortened to RAG, means retrieving relevant text from your own sources first, then having the model generate an answer from what was retrieved. The retrieval part is the search, and the generation part is the model writing the answer. A RAG system has two halves. The preparation half runs once, and again whenever documents change. It splits the documents into passages and stores them in a form that can be searched by meaning. The answering half runs on every question. It searches the stored passages, picks the best few and sends them with the question to the model. Lesson 3.2 explains how searching by meaning works, and lesson 3.3 covers splitting and storing.
There is one more advantage, and it is the one Daniel cares about most. The firm updates its policies every year. With retrieval, he swaps the old document for the new one and reruns the preparation step on that file. The next question uses the new policy.
The alternative people sometimes suggest is training or fine-tuning a model on the documents. That is slower, costs more, and is poor at storing exact facts like claim limits and leave entitlements. It also cannot tell you which page an answer came from. Retrieval keeps the facts in documents you control, which you can read, correct and point to.
Retrieval does not make the model infallible. If the search returns the wrong passage, the model will answer from the wrong passage. Lesson 3.4, Make answers cite their sources and admit gaps, deals with that.
Before any of that, you need a document set and a clear idea of what it should be able to answer. The activity below asks you to pick both.
Pick a document set you would like to question, such as policies or notes, and write five questions it should answer.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).