Score automatically: exact checks, rules and model graders

You will be able to choose a scoring method for each test and use a model as a grader with care.

Priya had 24 test cases and a spreadsheet. The first time she ran them, she checked each output by eye. It took 40 minutes, and she was not sure she had judged the last ten as carefully as the first ten. If she was going to run the set after every change, it had to score itself.

Different parts of an output need different kinds of checking. Picking the right method for each one is what makes an eval fast enough to run often and trustworthy enough to act on.

Exact checks for structured fields

When a field has one right answer, compare it directly. Labels, categories, dates, amounts, yes or no flags and IDs all work this way. If the expected sentiment is negative, the output either says negative or it does not. If the expected date is 2025-03-14, anything else fails.

This is the cheapest and most reliable kind of scoring. It runs in milliseconds, costs nothing, and never disagrees with itself. It is also one of the reasons module 2 pushed you towards structured output with enums. A field that can only be positive, negative or mixed is easy to check exactly. A free-text description of sentiment is not.

Lists need a small decision. For Priya's issues field, she checks that the expected issues are present, and allows extra ones, because an extra minor issue is not wrong. Write that choice down for each list field, so the check is consistent.

Code rules for format and content

Plenty of requirements can be checked by code that has no idea what the right answer is. Priya's app has a word limit of 25 on the summary, three required fields, an English-only rule and a short list of words that must never appear. Her code also flags any reply that repeats a line from the system prompt, and validates the JSON against the schema from lesson 2.2.

Your coding assistant can write each of these rules in a few lines. They cost nothing to run and give the same verdict every time, so use one wherever a requirement can be written down as a rule.

Model graders for open-ended text

What a rule cannot check is meaning: whether a summary captures the main problem, or whether a reply is accurate and polite. For these you can ask a model to do the grading, an approach often called LLM-as-a-judge.

The grader is a second model call. It gets the input, the output and a written rubric, which lists the criteria a good output must meet in words clear enough that two careful people would apply them the same way. Priya's rubric for summaries asks three questions, each answered pass or fail with a one-line reason. The first is whether the summary states the customer's main complaint. The second is whether it avoids any claim that is not in the original feedback, and the third is whether a team lead could act on the summary without opening the feedback.

Narrow yes or no questions like these work far better than a request to rate the summary from 1 to 10. Scores out of ten drift from run to run and are hard to interpret, while "did it add anything that was not in the feedback?" has an answer you can check yourself.

Graders get things wrong as well. They tend to be lenient, they can prefer longer answers, and they may prefer text that reads like their own. Priya tested hers before relying on it. She graded 20 outputs herself without looking at the grader's verdicts, compared the two sets, and found a summary that invented a refund request which the grader had passed. Her first rubric only asked about the main complaint, so the grader had no reason to object. The second question in her rubric was added because of that case.

Repeat a small hand check now and then, and always after you change the grader's model or rubric.

Look at pass rates by case type

A single headline score hides the pattern you most need to see. Suppose 20 of 24 cases pass. That sounds healthy until you notice that all four failures are among the six non-English cases, which means two in three of your non-English users get a bad result.

So report a pass rate for each tag alongside the total. This is why lesson 7.2, Build a test set from real cases, asked you to tag every case. Priya's first scored run had normal cases at 15 of 15 and other-language cases at 3 of 6, which pointed her straight at the problem.

Each scoring run should produce a table with one row per case, showing the case ID, its tag, the result of every check and a final pass or fail. Save each table with the date and the prompt version it tested, because lesson 7.4 puts two of them side by side.

The activity below asks you to choose a scoring method for each of your 20 cases, and to write the rubric for any that need a model grader.

Choose a scoring method for each of your 20 cases and write the rubric for any case a model will grade.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).