Give an assistant a calculator and a lookup tool

You will build a small assistant that calls two tools you wrote and uses their results in its answer.

Ask a model what S$289 is in ringgit, or what 17 percent of S$4,380 comes to, and it will usually give you a number that looks right. Sometimes it is. A model predicts text, as AI fundamentals explained in lesson 3.2, It writes by predicting the next token, again and again, and a calculation in the middle of a sentence is just more text. For anything your app relies on, arithmetic should come from code.

In this exercise you build a small assistant with two tools: a calculator, so numbers come from code, and a lookup over a small table, so facts come from your data. Then you test it with questions designed to find out when it uses each one. Allow about 35 minutes.

Step 1: build the two tools

The calculator takes two numbers and an operation: add, subtract, multiply or divide, as an enum. Keep it that simple. Do not build a tool that evaluates any expression the model writes, because that means running model output as code, which lesson 8.2 explains is dangerous. Your code checks both numbers are numbers, refuses division by zero with a clear error message, and returns the result.

The lookup tool searches a small table you create. Marcus used his product list: ten rows with product code, name, price in Singapore dollars and units in stock. A class timetable, a price list or a list of office rooms works just as well. The tool takes a product name or code and returns the matching row, or a clear message if nothing matches.

Write both tool definitions with the care lesson 4.2 described: plain names, descriptions that say when to use and not use each tool, typed parameters and required fields.

Step 2: have the assistant write the loop

Ask your coding assistant for a script that sends a user question with your two tool definitions, checks whether the reply contains tool calls, validates the arguments, runs the matching function, sends the results back and repeats until the model gives a text answer. Ask it to stop after five rounds, whatever happens, so a confused model cannot loop forever. Lesson 5.2 will make that limit more careful.

Ask for a log line for every tool call with the tool name, the arguments, whether validation passed and the result.

Read the code before running it. Make sure validation runs before each function does, because checking afterwards is too late. Check the key comes from the environment.

Step 3: write eight test questions

Your questions should cover four kinds of case, two of each. Here are Marcus's, as a worked example.

Questions that need no tool: "What does a burr grinder do?" and "Do you ship to Malaysia?" The second is a trap when the answer is in no tool or prompt, and a good assistant admits it does not know rather than inventing a shipping policy.

For one tool: "How many hand grinders do you have in stock?" needs the lookup. "What is 289 times 3?" needs the calculator.

For both tools: "How much would two of the gooseneck kettles cost?" needs the lookup for the price and then the calculator. In Marcus's table the example kettle price is S$68, so the right answer is S$136. "If I buy the grinder and the kettle, what is the total?" needs two lookups and an addition.

For bad input: "Look up product code XJ-9999" for a code that does not exist, and "What is 50 divided by 0?" Both should produce a clear explanation, not an error message or an invented answer.

Step 4: run them and read the log

Run all eight questions and read the log for each before you read the answer.

For every question, check four things in order: the tools called (or none, when none was needed), the arguments, what each tool returned, and whether the final answer used that result correctly.

Marcus's first run went mostly as planned. One surprise: for the kettle question, the model looked up the price and then did the multiplication itself instead of calling the calculator. The answer was right, but only by luck. He added a line to the calculator's description, "Use this for every calculation, including simple ones", and the next run called it. That kind of finding is the reason you read the log and not just the answers.

What done looks like

You have a script with two working tools, argument validation and a round limit. You have eight test questions covering no tool, one tool, both tools and bad input. And you have the tool call log from all eight, with a short note beside each saying whether the tool choice, arguments and final answer were right.

The note matters more than a perfect score. A log that shows you exactly where the model chose badly is the raw material for the agent you build in module 5.

Build the two-tool assistant, run eight test questions, and save the tool call log with a note on each.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).