Build a question answering tool over your notes

You will build a small tool that answers questions over your own documents with sources and honest gaps.

You now have every part of a retrieval system: documents chosen and five questions written in lesson 3.1, embeddings tried in lesson 3.2, chunks stored with metadata in lesson 3.3, and an answering prompt that cites and admits gaps from lesson 3.4. In this project you put them together into one tool and find out how well it really works.

Allow about an hour. Use your own document set. Daniel's version, used as the example throughout, answers staff questions over his firm's HR handbook and policies.

The brief

Build a small tool that takes a question, finds the most relevant passages in your documents, and returns an answer that cites its sources or says plainly that the documents do not cover it. It can run from the terminal. A web page is not required.

Then test it on ten questions and record two things separately for each: whether retrieval found the right passage, and whether the answer used it correctly. Keeping those apart is the most useful habit in this module, because it tells you which half of the system to fix.

Steps

Start with preparation. Ask your coding assistant for a script that loads your documents, splits them into chunks of a few paragraphs with a little overlap, embeds each chunk with your provider's embedding model, and stores the text, vector and metadata. Use whatever storage you chose in lesson 3.3. For a practice project a local file or a free-tier vector database is enough. Have the script record which embedding model it used.

Next, the answering script. For each question, it embeds the question with the same model, retrieves the top five chunks, builds the prompt from lesson 3.4 with each chunk numbered and labelled, calls the model and prints the answer with citations turned into readable sources. Add a debug option that prints the retrieved chunks too. You will use it constantly.

Then build your test set of ten questions. Use the five from lesson 3.1. Add two questions your documents cannot answer, so you can check the tool says so. Add three questions phrased in unusual words, the way real people ask, without the vocabulary the documents use. Daniel's three were "my kid is sick can I stay home", "what's the deal with claiming a cab after OT" and "do we get anything for our birthday".

Before running the tests, write down the correct answer and the source passage for each answerable question. Do this first, so the tool's output does not shape what you think the right answer is.

Finally, run all ten questions with debug on and fill in the results table.

The results table

One row per question, with five columns. The question. Did retrieval find the right chunk: yes, partly or no. Was the answer correct: yes, partly or no. Did it cite the right source. Notes.

For the two unanswerable questions, retrieval will always return something, since search returns the closest chunks even when none are relevant. Mark those rows on whether the answer correctly said it could not find the information.

Here is how Daniel's first run came out, as a worked example. Of his eight answerable questions, retrieval found the right chunk for six. Of those six, five answers were correct, and one mixed up two leave types that sat in the same chunk. Of the two retrieval misses, one was "what's the deal with claiming a cab after OT", where the casual wording landed nearer the overtime pay policy than the transport claims section. Both unanswerable questions got the fixed "could not find this" reply.

That breakdown told Daniel what to do next. Two retrieval misses meant his chunking and search needed attention before his prompt did. He split the long claims section into smaller chunks, added keyword search, and the cab question started working.

What done looks like

Your finished project has a preparation script and an answering script that run without errors, with the key read from the environment. It has your ten questions with the expected answer and source for each. And it has the results table, filled in for all ten, with retrieval quality and answer quality marked separately.

It does not need to score ten out of ten. A tool that answers seven correctly, admits it does not know the other three and tells you why is in better shape than one that answers all ten confidently and gets two badly wrong. Lesson 7.2, Build a test set from real cases, will reuse these ten questions as the start of a proper eval set, so keep them.

When the table is complete, read down the retrieval column before the answer column. That order tells you where to spend your next hour.

Build the tool, run ten test questions, and save a table recording retrieval quality and answer quality for each.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).