What an agent is and where it goes wrong

You will be able to define an AI agent and judge where a human should approve its steps.

Kumar, an office manager at a small accounting firm, asks an AI agent to organise the team's year-end lunch. "Find a restaurant near Raffles Place that can seat twelve on the third of December, book it, and send everyone a calendar invite." Twenty minutes later it reports that everything is done, and on the day twelve people arrive at a restaurant that has them down for the twelfth of March. Somewhere early on, the agent read 03/12 the American way, as month then day, and every step after that built on the mistake.

The booking, the invites and the confirmation email were all carried out correctly. Only the date was wrong, and it was wrong everywhere. That is the shape of most agent failures, and this lesson explains why.

A model in a loop

In lesson 7.3 a language model asked for a tool once, got the result and carried on with its answer. An agent takes that one step further and repeats it. Given a goal, the model plans a first step, calls a tool, reads the result, decides what to do next, calls another tool, and keeps going until it judges the goal is met or it gets stuck.

For Kumar's lunch, that loop might run like this. Search for restaurants near Raffles Place. Open three results and read their pages. Check which ones take group bookings. Open the booking form of the best one and fill it in. Read the confirmation. Open the calendar and create an event. Add twelve people from the address book. Send.

Each step is the same thing you learned in lesson 3.2, a model predicting likely next text, which in this case is the next tool request. What makes it an agent is the loop, and the fact that its tools reach out into the world: a browser that clicks and types, access to your files, your email or your calendar.

Mistakes become actions

When an ordinary assistant gets something wrong, you get wrong text. You can read it, spot the problem and ignore it. When an agent gets something wrong, the mistake has often already happened by the time you see it: the booking made, the form submitted, the email gone to a client or the file overwritten.

Agents can browse websites, fill in forms, edit spreadsheets, move files and send messages. Each of those is useful for the same reason it is risky: it changes something outside the chat. A wrong summary costs you a minute, while a wrong payment, a deleted folder or a message sent to the wrong group chat can cost a great deal more, and some cannot be undone.

Small errors compound

In a ten-step task, each step depends on the ones before it. If the model misreads something at step two, steps three to ten carry the error forward, and none of them is designed to go back and question it. Lesson 3.2 made the same point about single answers, where everything after a token has to fit it, and an agent does the same thing with actions instead of words.

The longer the chain, the more chances for a slip, and the further a slip travels. The final report often reads as a clean success, because from the agent's own point of view, every step did what the step before it set up. Nobody, including the agent, notices that the date in step two was wrong until a person checks the outcome against what they actually wanted.

There is a further risk worth one line here. An agent that reads web pages or emails can be misled by text written to trick it, for example hidden instructions on a page telling it to do something you never asked. The course AI risks: privacy, bias, deepfakes and scams covers this in more depth.

How much to let it do

None of this means agents are useless. They can save real time on tedious multi-step work, such as collecting prices from several suppliers' websites into one table. The answer is to design the job so that mistakes are caught before they turn into actions you cannot reverse.

Two habits do most of the work. First, give an agent the least access it needs. If it only needs to read your calendar to find a free slot, do not give it permission to send email. If it is collecting supplier prices, it does not need your company card. Second, require your approval before any step that spends money, sends a message to anyone, or deletes or overwrites something. Many agent tools have a setting for this, and some ask by default. Approval works best at the points where an error would stop being text and start being an action.

In Kumar's case, one approval step before the booking was submitted, showing the restaurant, the date and the number of people, would have caught the error in seconds. Think about a task in your own job that has several steps, and picture where in that chain you would want the same pause.

Pick one multi-step task from your job and mark which steps you would let an agent do alone, which need your approval and which you would never hand over.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).