You will be able to explain how a language model uses tools such as search, calculators and code without doing those things itself.
Ask an assistant what S$18,750 grows to over three years at 3.1 percent a year, compounded monthly, and two very different things can happen. An older or simpler setup writes out a figure that looks right and is a few dollars off, or sometimes a lot more. A newer one pauses, shows a small note such as "Analysing" or "Running code", and comes back with the exact amount. The second assistant did not get better at arithmetic. It asked something else to do the sum.
That something else is a tool, and the way a language model uses tools explains a good deal about what modern assistants can and cannot do. The figures above are just an example to try.
Lesson 5.1 described the hidden instructions that sit in the context window. One thing they often contain is a list of tools: a web search, a calculator, a code runner, a calendar, a way to read files. Each comes with a short description of what it does and what input it needs.
When the model is working out a reply and a tool would help, it does not write an answer. It writes a request instead, in a fixed format the app recognises. For the interest question, that might be a short block of code that does the compound interest sum. For a question about today's weather in Ang Mo Kio, it might be a search request with the words to look up.
This is still next-token prediction, as in lesson 3.2. During fine-tuning, covered in lesson 4.2, the model was trained on many examples where a tool request was the right next step, so for a question with a hard sum in it, the likeliest next tokens are a tool request rather than a guess.
From here the model steps aside and ordinary software takes over. The app spots the request and passes it to the right program: a search engine runs the search, a code runner executes the code, a calculator does the sum. These are the same kinds of programs that have been reliable for decades, and they compute or look things up without predicting anything.
The result comes back as text, such as the search results or the exact figure, and the app pastes it into the context window. Only then does the model continue writing, with that result now in view. If one tool call is not enough, it can make another, and another, before it writes the final answer.
This arrangement plays to each side's strengths. Lesson 3.1 showed that a model sees numbers as tokens, which is why it can stumble on long sums or counting letters. A calculator or a few lines of code have no such problem. And lesson 4.4 showed that a model knows nothing after its training cutoff, while a search tool can fetch a page published this morning.
So when a tool is used well, the arithmetic is done by something that is good at arithmetic, and fresh facts come from a fresh source. The model's job shrinks to what it does well: understanding the question, deciding what to ask the tools, and turning the results into a readable answer.
Tool use cuts down certain errors, but it adds a few points where things can go wrong, and the model is involved at every one of them.
The first is the choice of tool. The model might answer from memory when it should have searched, or skip the calculator because the sum looked easy. The second is the input it sends, where a search for an interest rate might use last year's date, or the code might have the right formula with 1.3 percent typed instead of 3.1. The third is reading the result: a search might return a bank page with several rates for different periods, and the model picks the one for the wrong term. In each case the final answer can look as solid as a correct one.
Many assistants show you when a tool was used, with a note that says "Searched the web", a list of sources, or a panel you can open to see the code that ran, and it pays to open them, since you can check the search words, the inputs to the sum and the page the figure came from, which is far quicker than redoing the whole thing yourself.
A good test is a question that needs both a fresh fact and a calculation, such as how much interest a sum would earn at a rate a local bank is offering this month. Ask one like that, watch for signs of which tools it used, and then check whether the final answer actually used what those tools returned.
Ask an assistant a question that needs a calculation and a recent fact, then note which tools it appeared to use and whether the final answer used their results correctly.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).