Learning means guessing, measuring the miss and adjusting

You will be able to describe the training loop and explain the difference between learning a pattern and memorising examples.

If you went through secondary school in Singapore, you probably met the ten-year series: a thick book of past exam papers. Some classmates worked through every question until they understood the topic. Others memorised the answers. Both groups scored well on the practice papers. Only one group did well when the real exam asked something new.

That difference, between learning a pattern and memorising answers, is the centre of how machine learning works. In lesson 2.1 you saw that a model is a calculation set by millions of dials. This lesson is about how those dials get set.

Guess, measure, adjust

Training is a loop with three steps, repeated a very large number of times.

First, the model makes a guess. Take a model that predicts how many people will miss their polyclinic appointment on a given day, so the clinic can plan staff. You feed it one example from last year's records: a rainy Monday in December, school holidays, 300 appointments booked. Its dials start at random settings, so its guess is poor. Say it predicts 10 no-shows. These figures are an example.

Second, you measure the miss. The records say 45 people did not turn up that day. The guess was 35 too low. That gap is called the error, and training is about making it smaller.

Third, you adjust. An algorithm works out, for every single dial, whether turning it slightly up or slightly down would have made the guess closer to 45. Then it nudges every dial a small step in that helpful direction. Perhaps the dial for rainy days moves up a fraction, and the one for school holidays moves a little too.

Then the loop starts again with the next example, and the next. One example teaches very little, and a small nudge barely changes anything. But repeat the loop over thousands of past days, and then go through all of them several more times, and the dials settle into settings that give small errors on most days. For a large language model, the loop runs over an enormous amount of text and every step nudges billions of dials at once, yet the steps stay the same.

The nudges are kept small on purpose. If the model swung its dials hard to fit each new example, it would lurch about and never settle, like a stallholder who doubles her order after one busy day and halves it after one quiet one.

The real goal is the next exam

A model that does well on its training examples has not proved anything yet. The clinic already knows how many people missed appointments last year. What it needs is a good guess for next month, on days the model has never seen.

Doing well on new examples is called generalisation, and it is the whole point of training. A model generalises when it has picked up the real pattern, such as rain and school holidays pushing no-shows up, rather than quirks of the particular days it was shown.

When a model memorises

Here is the trap. A model with enough dials can fit its training examples almost perfectly by memorising them, in the way that classmate memorised the ten-year series. It might learn that one Tuesday in March had exactly 52 no-shows because of some random event, and twist its settings to reproduce that. On the training data it looks brilliant. On next month's data it does worse than a simpler model would.

That is called overfitting: fitting the training examples so closely that the model fails on new ones. It gets more likely when you have a big model and few examples, or examples that are all very alike.

How builders test fairly

The defence is simple and builders use it everywhere. Before training starts, they hold back a portion of the examples and never show them to the model during training. Those held-back examples act like an exam paper the student has never seen.

After training, the model is scored on the held-back set. If it does well on training examples and also well on the held-back ones, it has probably learned the pattern. If it does well on training examples and badly on the held-back ones, it has memorised, and the builders go back and change something: more varied examples, a smaller model, or less training.

There is one more rule. The held-back set must look like the cases the model will really face. If the clinic tested only on dry weekdays in June, a good score would say little about rainy Mondays in December. Testing fairly means testing on the mix of cases you actually care about.

This loop, with its held-back test at the end, is how almost every model you use was built, from spam filters to the assistants in module 3. You can describe it yourself now in plain steps. Try that with a pair of fruits that even people sometimes mix up, and pay as much attention to the fair test as to the training.

Describe in four steps how you would teach a model to tell durian photos from jackfruit photos, including how you would test it fairly.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).