Set limits on steps, time and spend

You will be able to add hard limits that stop an agent before it runs away.

Marcus left his first agent running overnight on a batch of customer emails. One of them asked about a product he had stopped selling. The agent searched for it, found nothing, tried a different spelling, found nothing, searched by category, found other products, searched for the original again, and kept going. By morning it had made hundreds of calls on that one email, and his usage page showed a spend he would have noticed much sooner if he had not set a monthly limit in lesson 1.1.

The model was not malfunctioning. It was trying hard to finish a task that could not be finished. Limits are how your code makes the decision the model cannot make for itself: this has gone on long enough.

Cap the number of steps

The simplest limit is a step counter. Every time the loop runs, add one. When the count reaches your maximum, stop the loop, whatever the model wants to do next.

Base the maximum on runs you have actually seen. Marcus's longest legitimate run, the combined delivery question from lesson 5.1, took five tool calls. He set the cap at ten, about double the longest real case, which leaves room for a retry or two without allowing a runaway. Look at the runs you have logged and set the cap at a similar margin above your longest good one.

When the cap is hit, stop cleanly. Return a message that says the task was not completed, what was done so far and what is still open. For Marcus, that becomes a reply to himself rather than the customer: "Stopped after 10 steps. Could not find product 'Aeropress Go' in the catalogue. Customer email moved to manual review." A stop that explains itself is useful. A stop that just ends is another mystery to investigate.

Set a cost budget for the whole run

Steps are a rough proxy for cost, because some steps are much bigger than others. A step that pulls in a long document costs far more than a calculator call. So track cost directly as well.

Every response includes token usage, as you saw in lesson 1.4. Add the input and output tokens from every call in the run, convert them to money using the current prices, and keep a running total. Before each new step, check the total against a budget for the run. If the next step would exceed it, stop.

Set the budget from measured runs, in the same way as the step cap. If a typical run in your logs costs a few cents, a budget of a few times that catches runaways without cutting off normal work. Remember that later steps cost more than earlier ones, because each resends the growing history, so do not estimate from the first step alone.

The run budget is your own control, inside your code. Keep the provider spending limit as well. That one protects your account if your code has a bug.

Limit the time for each call and the whole run

Users wait for agents, so set two time limits. The first, from lesson 2.3, Handle a bad response without crashing, is a timeout on each model call and each tool call, so one slow request cannot hang the run. The second is a limit on the whole run. Record the start time, and before each step check whether the total has passed your limit.

Pick the run limit from what the user can tolerate. A support agent answering a customer in a chat window has a short limit, perhaps under a minute. A research agent working on a report that someone will read tomorrow can have much longer.

Stop the history outgrowing the context window

Every step adds to the conversation: the model's tool call, then the tool's result. Some results are long, like a page of search hits or a full document. After enough steps, the conversation can approach the model's context window, which lesson 1.3, Tokens, pricing and limits you have to plan around, described. At that point the call fails, and well before it, each step is slow and expensive.

Keep the history short on purpose. Before you add a long tool result, cut it down to the part the model needs. For older steps, replace the full result with a short summary, such as "Searched catalogue for Aeropress Go: no match." Keep the original goal and the most recent results in full. Some frameworks and providers offer tools that do this compaction for you. Check what yours offers before writing your own.

Decide what happens at each limit

For every limit, write down in advance what the user sees and what gets logged. Marcus's rules are short. Step cap: stop, mark for manual review, log the steps. Cost budget: stop, mark for manual review, alert him if it happens more than twice a day. Time limit: tell the customer someone will reply within a working day.

The activity below asks you to write the same set of limits for an agent you might build, with what happens when each one fires.

Write the limits for an agent you might build: maximum steps, cost budget, time limit and what happens at each.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).