Test the AI step on 20 real examples

You will measure how often your classification step gets it right before you rely on it.

Open your inbox, past form responses or the sheet your flow reads from, and copy out 20 real messages of the kind your AI step will handle. The exercise takes about half an hour. You'll use those 20 messages for this test and again every time you change the prompt later.

In lesson 5.2 you wrote a classification prompt and tried it on a couple of messages. Two messages coming back right shows the step can work. It does not show how often it works, or which kinds of message it gets wrong. Running it on 20 real messages turns your feeling about the step into a number.

Choosing and labelling the 20 messages

Use real messages because invented ones are tidier than what people actually send. Aim for a spread: mostly ordinary messages, a few awkward ones, and at least one or two you would hesitate over yourself.

Before you use them, remove personal details such as names, phone numbers and email addresses, unless your workplace has approved sending that data to the model provider. Lesson 2.4 of AI at work, "Redact a real document so you can still get help", shows how to remove these details without losing the meaning of the message.

Set up a sheet with five columns: message number, message text, your label, AI label, and match. Fill in your own label for all 20 messages before the AI sees any of them. Once you have seen the AI's answer, you will be tempted to agree with it, and your labels will no longer be a fair check.

Running the step and counting the results

You can run the messages in two ways. You can trigger the flow once per message with test data, or you can use your tool's option to test a single step with different inputs. Record each AI label in its column.

Then mark every row with one of three results:

Match: the AI label is the same as yours. Mismatch: the AI gave a different real label. Unsure: the AI answered unsure.

Count each result. Mismatches are the dangerous ones, because they send a message down the wrong path without anyone noticing. Unsure results are safe, since they go to a person, but if there are too many of them the step is not saving you much work.

Mei Ling runs the enquiries for a tuition centre, and her labels are trial, schedule and unsure. On her first run of 20 enquiries she got 14 matches, 3 mismatches and 3 unsure. That is 14 out of 20, or 70 percent right, with 15 percent sent down the wrong path and 15 percent sent to her to sort. These are illustrative figures, and your numbers will differ.

Finding out why each mismatch happened

When you see the totals, you may want to start adding instructions to the prompt. Read each mismatch first and work out why it happened. A mismatch usually has one of three causes.

The first is overlapping label definitions, where the message really fits two labels. The fix is to tighten the definitions so each message has one home. The second is a kind of message that is missing from your labels entirely. Either add a label, if that kind of message needs its own action, or say explicitly which existing label it belongs to. The third is that your own label was wrong or inconsistent, which happens more than people expect. Correct the label in your sheet and note the correction.

Mei Ling's mismatches came from two of these causes. Two messages asking "is there still space in the Saturday class?" were labelled trial by the AI, while she had labelled them schedule. Her definitions did not say where availability questions belonged. The third mismatch was a parent asking about make-up classes for a missed lesson. It fitted none of her labels and should have come back unsure.

Making a small change and rerunning the same set

Change the label definitions or examples to deal with what the mismatches showed. Change as little as you can each time, so you know which change made the difference. Adding general advice such as "be careful and think step by step" rarely helps, and it makes the prompt harder to maintain.

Mei Ling made two changes. She added "including whether a class has space" to the schedule definition. She also added "questions from current students about their own classes" to the unsure definition, because she wants to handle those herself.

Rerun the same 20 messages through the revised prompt. If you used different messages, you could not tell whether a better result came from the prompt or from easier messages. Record the new results next to the old ones. If the result got worse, undo the change. If it got better, keep it.

Mei Ling's second run on the same 20 gave 17 matches, 1 mismatch and 2 unsure. That is 85 percent right, 5 percent misrouted and 10 percent sent to her.

You can repeat the cycle, but the target for this exercise is one careful revision. Twenty messages is a small sample, so treat your percentages as a rough guide rather than a precise rate. Most of the value comes from reading the mismatches. Keep the 20 messages and your labels, and test every future prompt change on the same set before it goes live.

You are finished when you have a table of 20 real messages with your labels and the AI's labels for two runs, a count of matches, mismatches and unsure for each run, and a one-line note of what you changed in the prompt and why. Gather your 20 messages first, and label them before you open the tool.

Run 20 labelled examples through your AI step, record the results in a table, revise the prompt once and rerun.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).