You will be able to explain why systematic evaluation is needed for AI features.
A week after Priya's app went live, a team lead complained that summaries of Malay feedback were coming out in Malay, even though the system prompt said English. Priya added a stronger line to the prompt, pasted in three Malay examples, and all three came back in English. She shipped the change. Two days later another team lead noticed that short complaints like "late again" were now summarised as "The customer did not state a problem", which had never happened before. Her fix for one case had broken another, and she had no way of knowing until someone complained.
This is what testing by hand looks like with AI features. It feels like testing, but it misses most of what matters.
When you try a prompt by hand, you pick the inputs. You choose ones you thought of, which tend to be typical, tidy and similar to each other. Real users send what they actually have: half-sentences, three languages in one message, pasted email signatures, typos, sarcasm, a single emoji, a 2,000-word rant.
The failures live in that variety. Priya's three Malay tests passed. The failure was in a type of input she had not thought to try, very short complaints, and a handful of hand-picked tries had no way to reach it.
So a few successful tries tell you that the feature can work, which you knew already. They do not tell you how often it fails on real traffic, or on which kinds of input.
There is a second problem. As you saw in AI fundamentals, lesson 3.3, Why the same question gets a different answer, models produce varied output even for identical input. Lowering the randomness setting, as lesson 2.3 suggested, reduces this, but rarely removes it completely. Providers also update models behind the same name from time to time, and that can shift behaviour too.
That means one good result does not guarantee the next. If an input produces the right summary today, it may produce a slightly wrong one on the fifth run. For testing, a single pass on each case is weak evidence. Running important cases a few times, and looking for any that flip between pass and fail, tells you which cases are fragile.
The third problem is the one that caught Priya. A prompt is not a set of separate rules that each affect one case. Every line influences every output. Adding "always reply in English, even for very short or non-English feedback" changed how the model treated short feedback in general, and some of those outputs got worse.
Ordinary software has the same issue, and programmers solved it long ago with automated tests. They keep a set of tests that check the behaviour that matters, and they run all of them after every change. If a change breaks something that used to work, a test fails before users see it. That habit, more than any clever technique, is what lets software teams change code confidently.
For AI features the equivalent is an eval, short for evaluation: a fixed set of test cases, each with an input and a description of a good output, that you run against your app every time you change something. A prompt edit, a new model, a different temperature, a change to your retrieval or a new tool: each gets the same eval run before it ships.
The key word is fixed. The cases stay the same between runs, so differences in the results come from your change, not from different inputs. Over time the set grows, because every bug a user reports becomes a new case, but you never quietly swap cases out.
An eval gives you a number, such as 18 of 20 cases passing. More usefully, it gives you a list of which cases pass and which fail, so after a change you can see exactly what improved and what broke. Had Priya had one, her English fix would have shown a new failure on the short complaints the moment she ran it, before any team lead saw it.
You already have the start of one. The ten hand-checked emails from lesson 2.4, Build an extractor that turns emails into JSON, and the ten test questions from lesson 3.5 were small evals. This module makes the habit systematic: lesson 7.2 builds the test set, lesson 7.3 scores it automatically, and lesson 7.4 uses it to make a real decision.
Before building anything, it helps to know what you are protecting against. The activity below asks you to think through the changes you are likely to make to your app and what each could break.
List three changes you might make to your app, such as a new prompt or model, and what could break with each.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).