You will be able to test whether an assistant treats people differently by changing one detail at a time.
Your manager asks you to use an assistant to help write year-end reviews for your team of six. You paste in your notes for each person and ask for a polished paragraph. The results read well. But would they read the same if the names were different? You have no way of knowing just by looking, because you only ever see one version.
The cases in lesson 3.2, Documented cases: hiring, faces and risk scores, were found by comparing outcomes across groups. A swap test is a small version of that comparison, and you can run it yourself on any assistant in a few minutes.
The idea is simple. You run the same prompt twice, and the only thing you change between the two runs is one personal detail: a name, a gender, an age, a nationality, a race. Everything else stays word for word identical. If the outputs differ in ways that matter, the detail you swapped is the likely cause, because it's the only thing that changed.
Say Farah, the HR executive from lesson 3.1, wants to test an assistant she uses to summarise interview notes. She writes a set of notes for a fictional warehouse supervisor candidate: eight years' experience, led a team of twelve, improved picking accuracy, sometimes direct with colleagues. She runs it once with the candidate named Tan Wei Ming and once with the candidate named Nurul Aisyah. Then she tries an older and a younger candidate, where the age and graduation year are the only things she edits.
Pick swaps that matter to the task. For hiring, gender, race, age and nationality are the obvious ones. For customer service replies, you might swap a local name for a foreign one, or an English-sounding message for one with Singlish phrasing. For loan or insurance notes, age and family situation.
Don't stop at the headline answer. Two summaries can both say "recommend for next round" and still treat the candidates very differently. Read the two outputs side by side and look at these things.
Tone: does one sound warmer, more cautious, more enthusiastic? Recommendations: does one get suggested for a more senior role, more training, a lower salary band? Adjectives: this is where stereotypes often hide. Watch for one person being "caring" and "supportive" while the other is "confident" and "driven", or one person's directness becoming "abrasive" while the other's becomes "decisive". Omissions: did one summary leave out an achievement that the other kept?
In Farah's test, both summaries recommended the candidate. But for Nurul Aisyah, the summary described the direct manner as "may need coaching on communication", while for Tan Wei Ming it said "clear and assertive communicator". The notes were identical and only the name had changed.
Assistants don't give the same answer every time. AI fundamentals: what it is, how it works, where it fails explains why in lesson 3.3, Why the same question gets a different answer. That matters for a swap test, because a single difference might be chance rather than bias.
So repeat each version a few times, ideally in fresh chats so earlier answers don't influence later ones. If you run each version three times and the "needs coaching" phrasing turns up only for one name, every time, that's a pattern. If it shows up for both names on different runs, it may be noise, though it's still worth noting that the assistant sees "direct" as a problem at all.
Keep the test fair in other ways too. Turn off memory or use a temporary chat, as described in lesson 1.3, History, memory, temporary chats and deletion, so the assistant isn't drawing on what it knows about you. And use fictional people, never real colleagues or candidates, since the point is to test the tool, not to paste personal data into it.
The rule is straightforward. If the output changes with the swap in a way that would affect how someone is treated, don't use that output for decisions about people.
That doesn't always mean abandoning the tool. You might rewrite the prompt to remove names entirely and refer to "the candidate". You might use the assistant only for structure and grammar, and write the judgements yourself. You might keep it for tasks that don't involve people at all. Lesson 3.4, Bias-test an AI task you rely on, asks you to make exactly this kind of decision.
Farah changed her process. She now removes names from interview notes before using an assistant, and she writes the final recommendation herself.
A swap test won't prove that a tool is fair. It can only show you a problem when there is one, and a clean result on five swaps says nothing about the sixth. But it turns a vague worry into something you've actually checked, and that's more than most people ever do.
For your first test, choose a prompt about people that you might really use at work, such as a performance review paragraph or a job description, and decide which single detail you'll swap.
Run a swap test on a performance review or job description prompt and note every difference between the two outputs.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).