Where bias in AI comes from

You will be able to name the main sources of bias in an AI system and give an example of each.

Farah is an HR executive at a logistics company in Jurong. She gets about 300 applications for every warehouse supervisor opening, and a vendor has offered a tool that ranks CVs automatically. The demo looks impressive. The vendor says the tool is objective, because it's a machine and doesn't have opinions. Farah isn't so sure, and she's right to hesitate.

A machine doesn't have opinions, but it does have a history. Everything an AI system knows comes from the data it learned from and the choices the people building it made. Bias gets in through both. This lesson covers four routes it takes, so you can recognise them in the tools you use.

Route one: it learns from the past

Most AI systems learn by finding patterns in past examples, as AI fundamentals: what it is, how it works, where it fails explains in lesson 2.3, A model can only be as good as its examples. Show a system thousands of past hiring decisions and it learns what the people who made those decisions tended to prefer.

If those past decisions were fair, that's fine. If they weren't, the system learns the unfairness along with everything else and repeats it, consistently and at speed. Suppose a company's supervisors were mostly promoted from one group of staff for twenty years, for reasons that had little to do with ability. A tool trained on that history will learn that people who look like past supervisors are good candidates. It has no way of knowing the pattern was a mistake. To the model, it's just a pattern.

Route two: it finds a stand-in

The obvious fix is to remove sensitive details: take out gender, race, religion, age and nationality, and the system can't use them. Unfortunately it often still can, through proxies: other details that happen to track the ones you removed.

A proxy can be almost anything. A postcode can track income or ethnicity in cities where neighbourhoods are divided along those lines. The school someone attended can track family background. A gap in employment can track caregiving, which still falls more often on women. Certain hobbies, clubs or even words in a CV can track gender. The system isn't looking for these on purpose. It finds whatever patterns predict the outcome it was trained on, and if the outcome was biased, the proxies come along.

That's why "we removed the protected details" isn't enough on its own. You have to look at what the system actually does to different groups.

Route three: it saw some people less

A system learns best about the people it saw most often in training. Groups that were rare in the data get less accurate results, even if nobody meant any harm.

A speech recognition tool trained mostly on American and British voices may struggle more with Singaporean English. A tool that reads medical images may be less accurate for skin tones that were under-represented in its training set. A CV tool trained mainly on applications from local graduates may handle a foreign degree or an unusual career path badly. In each case, the tool works well on average and worse for a particular group. And the average is often the only number anyone checked.

Route four: it repeats what it read

Language models such as the ones behind ChatGPT, Claude or Gemini learned from enormous amounts of text written by people. That text contains stereotypes about jobs, gender, race, age and nationality, and the model picks them up along with grammar and facts.

You can see this in everyday use. Ask for a story about a nurse and an engineer and notice which one the model makes a woman. Ask for a reference letter and notice whether the adjectives change with the name: "warm" and "helpful" for one person, "decisive" and "strategic" for another. Providers work to reduce this, and modern assistants are much better than early ones, but nobody has removed it entirely. The words can still lean one way without you noticing.

For Farah, this route matters even though the vendor's tool isn't a chatbot. If she uses a general assistant to write job descriptions or summarise interview notes, stereotypes in its writing can shape who applies and how candidates are remembered.

Why this matters to you

You don't have to build AI to be affected by its bias. You're affected when you use a tool that ranks, scores, summarises or writes about people. Singapore's Tripartite Guidelines on Fair Employment Practices already expect employers to select people on merit, and "the software decided" doesn't change who's responsible for the decision.

Farah's first step wasn't to reject the tool. It was to ask the vendor what data it was trained on, whether it had been tested for different results across groups, and what those results were. The vendor's answers were vague, which told her a lot.

Think about the tools around you at work that make or shape decisions about people, even in small ways. For each of the four routes, there's probably a plausible way it could show up in one of them.

Write one example of each source of bias that could plausibly arise in a tool used at your workplace.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).