You will run your test set against two versions of your prompt and decide which to ship from the results.
This is the moment the eval earns its keep. You have a test set from lesson 7.2 and a scoring method for each case from lesson 7.3. In this exercise you use them to make a real decision: whether a change to your prompt should ship.
Allow about 40 minutes. Priya's run is used as the worked example, testing the "always reply in English" change from lesson 7.1.
Ask your coding assistant for a script that does four things. It reads your test cases from their file. It runs each input through your app's real code path, the same prompt, schema and model the live app uses. It scores each output with the method you chose: exact checks, code rules or the model grader with your rubric. And it saves a results table with one row per case, showing the case ID, the tag, the result of each check and an overall pass or fail.
Make sure the script names each results file with the date and a prompt version label, such as v3 or v4, so results never overwrite each other. Store the prompt itself in version control, as lesson 2.1 recommended, so each label points to an exact text.
The eval calls the model for every case, and the grader adds more calls, so check the cost with the arithmetic from lesson 1.3 before your first run. For a few dozen cases it is usually small, but you will run it often.
Run the full set on your current prompt, without changing anything. This is your baseline. Save the results table.
Priya's baseline on version 3 of her prompt: 20 of her 24 cases passed. The four failures were two Malay messages answered in Malay, one Chinese message answered in Chinese, and one pasted email signature summarised as if it were a complaint.
If any case failed in a way you did not expect, look at it now. Sometimes the expected output in the test case is wrong, and you should fix the test case before going on. Do not fix the app yet.
Make a single change. Only one, so that any difference in results comes from it. Priya changed one line of the system prompt to "Always reply in English, even when the feedback is very short or not in English", saved it as version 4 and committed it.
Run the full set again, every case, not just the ones you hoped to fix. Save the new results table.
Priya's version 4: 22 of 24 passed. At first glance, a clear improvement.
Now put the two tables side by side and sort every case into one of four groups. Passed both times. Failed both times. Fixed: failed before, passes now. Broken: passed before, fails now.
Priya's comparison, as a worked example. 19 cases passed both times. One failed both times, the email signature case. Three were fixed, all three non-English cases. And one was broken: "late again", which had passed on version 3 and now came back as "The customer did not state a problem."
The arithmetic checks out. 20 passing before, plus 3 fixed, minus 1 broken, gives 22. If your counts do not add up like this, something in the comparison is wrong.
The headline went from 20 to 22, but that alone is not enough to decide. Read every case that went from pass to fail, in full, with its output, before deciding anything.
For Priya, the broken case was exactly the kind of feedback her team leads see every day. A short angry message summarised as "no problem stated" is worse than a summary in Malay, because a team lead might skip it entirely. So she did not ship version 4. She wrote version 5, which kept the English rule and added "Very short feedback still describes a problem; summarise what it implies". She ran the full set again: 23 of 24, with nothing broken compared with version 3. Version 5 shipped.
Also check the cases that failed both times. They tell you what your change did not touch. Priya added the email signature case to her list for the next round.
If you ran important cases several times, as lesson 7.1 suggested, check whether any broken case is just flipping between runs. A case that fails one run in three is fragile, and it needs attention whichever version you ship.
You have an eval script, two saved results tables, one for each prompt version, and a short decision note. The note says how many cases passed each time, which cases were fixed and which were broken, what you read in each broken case, and which version you will ship and why.
Keep both tables in your repository next to the prompt versions they came from. The next time someone asks why the prompt says what it says, the answer will be in those files.
Run your eval on two prompt versions, save both result tables, and write a short decision note on which to ship.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).