Run a fair side-by-side test

You will compare two or three assistants on your own tasks with the same prompts and a simple scoring sheet.

Most people choose an assistant the way they choose a phone: they try one, get used to it, and assume the others are roughly the same or roughly worse. Lesson 8.1, What actually differs between assistants, explained why reviews cannot settle the question for you. This exercise settles it with your own work, using the same prompts in each assistant and a simple scoring sheet.

Allow about thirty-five minutes. You need access to at least two assistants, and three if you can. Free plans are fine for most tasks, though if one of your tasks needs a feature like file upload, check that each assistant offers it on the plan you have. If you are testing for work, use only assistants your employer allows for the data involved, as lesson 8.2, Data settings and the rules at your workplace, covered.

Step 1: pick five real tasks

Choose five tasks that represent what you actually do, in roughly the proportions you do them. If most of your use is writing emails, two of the five should be emails. Include at least one task with a document, such as summarising a report or answering questions about a PDF, and at least one with a fact you can check, such as a question about a rule whose answer you can confirm on an official website.

Grace, the recruitment team lead from this module, chose: a job ad for an accounts executive role, a rejection email to a candidate who reached the final round, a summary of three CVs against a job description (with the candidates' details removed, as her policy required), a question about the current rules for a particular work pass whose answer she could check on the Ministry of Manpower website, and a set of interview questions for a sales role.

Step 2: run identical prompts, fairly

Write each prompt once, using everything from this course: context, task, constraints and format. Then use exactly the same prompt in each assistant. Do not improve it as you go, even if you notice a weakness, because then the later assistants get a better prompt than the earlier ones.

Run every task in a fresh chat, so no earlier conversation affects the answer. Use the same mode in each assistant where possible, standard against standard or reasoning against reasoning. If one assistant has custom instructions or memory switched on and another does not, either switch them off for the test or set up equivalent instructions in each, using your setup record from lesson 6.4, Set up your assistant for your working week.

Paste each output into a document, labelled only with a letter, such as A, B and C. Keep a separate note of which letter is which assistant.

Step 3: score without looking

Score each output on three things, each from one to five. Accuracy: is everything in it correct, checked against the source for the document and fact tasks? Usefulness: does it do the job you needed? Edits needed: five means usable as it is, one means you would rewrite it.

If you can, score blind. Ask a colleague to shuffle the outputs and hide the letters, or simply score from the lettered document without checking your note. Most people have a favourite assistant before they start, and knowing which one wrote an answer tends to tilt the score.

Check the fact task properly. Grace went to the Ministry of Manpower website for the work pass question. Assistant A's answer matched the current rules. Assistant B gave a figure that had been correct in an earlier year and presented it as current. That single result told her more about how much to trust each one on regulatory questions than any review.

Step 4: decide

Add up the scores for each assistant, and look at the pattern as well as the totals. Out of a possible 15 per task, in the order listed in step 1, Grace's scores for assistant A were 12, 13, 9, 12 and 14, a total of 60. Assistant B scored 11, 12, 13, 6 and 10, a total of 52. These are her scores, given as an example of how a test can come out.

The totals favoured A, but the pattern told her more. A was ahead on the three writing tasks her team does most, and B's outdated work pass answer cost it heavily on accuracy. On the CV summary, though, B scored four points higher, a gap large enough to take seriously when a difference of one point on a task could easily be chance.

Grace's decision was to make A the team's main assistant, provided it was the one her agency already had a business plan for, and to note that document-heavy summaries might be worth sending to B if the agency ever approved it.

What a finished version looks like

Your scoring sheet has a row for each task and, for each assistant, three scores and a total. Below it, write a three-sentence decision: which assistant you will use as your main one, which task, if any, you would send elsewhere, and what would make you run the test again, such as a new version of either assistant.

Keep the five prompts. They are a ready-made test you can run again in a few months, when the products have changed, as they will. Start by listing the five tasks, and write the prompt for the first one.

Complete a scoring sheet for five tasks across at least two assistants and write a three-sentence decision on which to use for what.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).