You will build a small agent that uses tools to answer a question and stops within your limits.
Everything in this module comes together in one small build. You have the loop from lesson 5.1, the limits from lesson 5.2, the failure checks from lesson 5.3 and the read, draft and act rules from lesson 5.4. In this exercise you build a research agent that answers a question using a few tools, stops within limits you set, and leaves a log you can read step by step.
Allow about 40 minutes. Daniel's version, used as the example throughout, researches questions over the HR documents he indexed in module 3.
Keep the toolset small. Daniel's agent has three.
The first is search_documents, which runs the retrieval you built in lesson 3.5 and returns the top few chunks with their sources. The second is calculate, the calculator from lesson 4.4, so any arithmetic comes from code. The third is write_note, which appends a line to a notes file for the run. That lets the agent keep track of findings without relying only on the growing conversation, and it is a draft action in lesson 5.4's terms, since it changes nothing outside the run.
None of these tools sends, pays or deletes anything, so this agent needs no approval step. If you add one that does, put it behind approval first.
Write each tool definition as lesson 4.2 described, with clear descriptions of when to use and not use each.
Ask your coding assistant to write the loop with three limits you set: a step cap, a cost budget for the run, and a time limit for the run, plus a timeout on every call. Daniel used a cap of eight steps, a budget worked out from his measured token prices, and two minutes in total. When any limit fires, the loop stops and writes a clear message saying which limit, what was done and what is left.
Add the loop check from lesson 5.3: if the agent tries a tool call with the same name and arguments it already used, return a message saying so instead of running it.
Log every step: the step number, the tool and arguments, a short version of the result, the tokens used and the running cost. Write the log to a file per run.
Include the goal in every request, as lesson 5.3 recommended against drift, and tell the agent in its system prompt to answer only from what the tools return and to say plainly if it cannot find an answer.
Each question tests a different behaviour.
The first is a question the agent can answer, ideally one that needs more than one step. Daniel's was "If an employee takes all their annual leave and the maximum days of childcare leave in the firm's policy, how many working days off is that in total?" It needs two searches and a calculation. In his made-up example handbook, annual leave is 14 days and childcare leave is 6, so the expected answer is 20.
The second is a question the documents cannot answer. Daniel asked, "What is the firm's policy on pets in the office?" The handbook says nothing about pets. A good run searches once or twice, finds nothing relevant and says so. A bad run invents a policy or keeps searching.
The third is a question designed to make it loop. Daniel asked about a form, HR-12, that does not exist, phrased as if it certainly does: "Find form HR-12 and tell me its approval steps." A model told the form exists may keep searching with different words. This is the run that should hit your loop check or your step cap.
Run each question and read each log in full before you look at the final answer.
For each step, note whether the choice was sensible. Was a search needed? Were the arguments good? Did the agent use what came back? Then note any wasted steps: repeated searches, detours from the goal, a calculation done in its head instead of with the tool.
Daniel's runs, as a worked example. The leave question took four steps, two searches, one calculation and one note, and answered 20 with sources. The pets question took two searches and gave an honest "not covered" answer. The HR-12 question tried four differently worded searches, then the loop check caught a repeat, and the step cap stopped it at eight with a clear message. His note: the limits worked, but the agent wasted three searches before giving up, so he added "If two searches find nothing relevant, say you could not find it" to the system prompt.
You have a working agent with two or three tools, a step cap, a cost budget and a time limit, a log file for each of three runs, and a short note on each run saying where the agent chose well and where it wasted steps. One of the three runs should have stopped early because of a limit or the loop check. If none did, your limits are probably too loose.
These three logs show you more about agents than any amount of reading. Go and generate them.
Build the agent, run three test questions, and save the step logs with your notes on each run.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).