A model can only be as good as its examples

You will be able to name the main ways training data limits a model and spot them in a real product.

Ask a voice assistant to call your grandmother and it gets it right. Ask your grandmother to use the same assistant and it might not understand her at all. Same phone, same app, same request. The difference is who is speaking, and how many people who sound like her were in the recordings the model learned from.

Lesson 2.2 showed that a model learns by adjusting its dials to fit examples. The flip side follows directly: it can only learn what is in those examples. This lesson covers the four main ways training data sets limits on a model, so you can spot them in products you use.

What is rare in the data, the model does badly

A model gets good at the cases it sees often. Cases it rarely sees get fewer nudges during training, so its dials end up tuned for the common ones.

Speech recognition is the clearest example. If a model was trained mostly on recordings of speakers with American or British accents, it will usually do worse with accents that were rare in its training audio. In Singapore you can hear this when voice typing stumbles over Singlish particles, Malay or Hokkien place names, or an older speaker's English. The model is not choosing to ignore anyone. It simply had less to learn from.

The same thing happens with photos taken in unusual lighting, skin tones that were under-represented in an image dataset, or medical scans from a type of machine the model never saw. Whenever a group or situation is rare in the data, expect worse results for it, and expect the model to give no warning that it is out of its depth.

The labels carry human mistakes

Many models learn from labelled examples: this email is spam, this X-ray shows a fracture, this CV was shortlisted. Those labels come from people, and people make mistakes, disagree with each other and bring their own assumptions.

If the staff who labelled customer complaints often marked angry-sounding messages as urgent and politely worded ones as routine, the model learns that tone means urgency. A polite email about a serious problem then gets sorted to the bottom of the pile. Nothing in the model is broken. It learned exactly what the labels taught it, assumptions included.

Past decisions bring past habits

A special case of this is a model trained on past decisions. Suppose a company trains a CV screening model on ten years of its own hiring records, labelled hired or not hired. The model learns to predict who this company used to hire. If those choices favoured graduates of certain schools, or quietly passed over older applicants, or people who had taken a career break to care for a parent, the model learns that too, and applies it at speed to every new CV.

This is not only a hypothetical. Reuters reported in 2018 that Amazon had dropped an experimental recruiting tool after finding it marked down CVs that included the word "women's", as in a women's chess club, because it had learned from years of applications that came mostly from men.

Builders can correct for this by checking results across groups, removing or rebalancing biased examples, and testing before launch. But none of that happens by default. A model trained on history repeats history unless someone deliberately stops it. For an HR officer, that is worth knowing before trusting any tool that ranks applicants, and Singapore's Tripartite Guidelines on Fair Employment Practices still apply to the decisions you make with it.

Every dataset has a date

Training data is collected up to some point and then frozen. A model knows nothing about anything that happened after its data was gathered, unless it is given new information while you use it.

A model trained to predict polyclinic no-shows from records before a new booking system launched will not know the new system exists, and its predictions may drift as patient habits change. A language model whose text ends before a rule change will describe the old rule with full confidence. Module 4 returns to this as the knowledge cutoff, and lesson 4.4 has you test it on recent Singapore changes.

Putting it together

When you meet a product that uses AI, these four questions get you most of the way. Who or what was rare in its training data? Who made the labels, and what might they have got wrong? Was it trained on past decisions, and were those decisions fair? How old is its data?

You rarely get full answers from a vendor. But you can often make a sensible guess from what the product does and who built it, and that guess tells you where to watch its results most closely. Start with a tool you know well from your own week, because you already know who uses it and where it has let someone down.

Choose one AI feature you use and write down who or what might be under-represented in its training data and how that could show up in its results.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).