Documented cases: hiring, faces and risk scores

You will be able to explain three well-documented bias cases and the lesson each one teaches.

When someone tells you an AI tool is biased, it's easy to shrug. It sounds like a theory, or an opinion about technology in general. So this lesson leaves theory aside and looks at three cases that were investigated, documented and reported in detail. Each one went wrong in a different way, and each teaches you a check you can use.

You met the four routes bias takes in lesson 3.1, Where bias in AI comes from. Watch for them in each case.

Amazon's recruiting tool

In 2018 Reuters reported that Amazon had built an experimental tool to rate job applicants' CVs, and then abandoned it. According to the report, the tool had been trained on CVs submitted to the company over about ten years. Most of those came from men, reflecting who worked in the technology industry at the time.

The system learned that pattern. Reuters reported that it downgraded CVs that included the word "women's", as in "women's chess club captain", and marked down graduates of two all-women's colleges. Engineers edited the system to ignore those specific terms, but they couldn't be sure it wouldn't find other ways to sort candidates that had the same effect. Reuters reported that Amazon said the tool was never used by its recruiters to evaluate candidates.

This is route one, learning from the past, combined with route two, proxies. Nobody told the system to prefer men. It found the preference in the history and then found words that stood in for gender. The lesson for Farah, the HR executive from lesson 3.1, is direct: a CV ranking tool trained on your company's past hires inherits your company's past choices.

Gender Shades: faces and skin tone

In 2018 the researchers Joy Buolamwini and Timnit Gebru published a study called Gender Shades. They tested commercial facial analysis systems from several large technology companies on a task those systems were sold to do: classifying the gender of a face in a photo.

They built a test set of photos that was balanced across gender and skin tone, which the common test sets of the time weren't. The systems performed very well on lighter-skinned men. They were much less accurate on darker-skinned women, with error rates far higher than for any other group. The headline accuracy figures that companies quoted had hidden this, because the usual test sets were dominated by lighter-skinned faces.

This is route three, gaps in the data, and it also shows a problem with testing. If the test data has the same gaps as the training data, a system can score well and still fail badly for a whole group of people. The companies involved later released updated systems, but the study changed how many researchers and companies think about measuring accuracy.

COMPAS and risk scores

In 2016 the US news organisation ProPublica published an analysis of COMPAS, a commercial tool used in parts of the American justice system to estimate how likely a defendant was to reoffend. ProPublica studied scores given to people arrested in one Florida county and compared them with who actually reoffended over the following years.

ProPublica reported that black defendants who didn't go on to reoffend were almost twice as likely as white defendants to have been labelled higher risk. White defendants who did reoffend were more often labelled lower risk. The company behind COMPAS disputed the analysis, arguing that its scores were equally accurate for both groups by its own measure.

Both sides had a point, and that's the deeper lesson here. Researchers later showed that several common definitions of fairness can't all be met at once when groups have different underlying rates. Choosing which kind of fairness matters is a human decision about values. A tool can't make it for you, and a vendor's claim that a system is "fair" means little until you know which definition they used.

What the three cases share

The three cases involved hiring, faces and courts, three very different settings. They share one feature. In each, the problem was found by looking at outcomes broken down by group: who got marked down, who was misclassified, who was wrongly labelled high risk. Overall accuracy didn't reveal it.

That gives you the check. Before an AI tool is used on decisions about people, someone should test what it does for different groups, and compare. After complaints arrive is too late, because by then the decisions have been made and people have been affected.

You won't usually be running large studies like these. But you can ask a vendor the question Farah asked: have you tested results across groups, and can we see them? And for your own everyday use of assistants, the next lesson gives you a small version of the same test.

Look back at the three cases now with one question in mind: at what point, before any harm was done, could a simple check have shown the problem?

For each case, write one sentence on where the bias came from and one check that might have caught it earlier.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).