You will be able to explain what pretraining is, what data it uses and what the resulting model can and cannot do.
Imagine a new hire who has read most of the public internet, a huge library of books and millions of pages of code, and has never once spoken to a customer. They know an enormous amount. Ask them a question, though, and they might reply with three more questions, because that is what the FAQ pages they read looked like.
That is roughly what you get at the end of the first stage of building an assistant like ChatGPT, Claude, Gemini or Copilot. The stage is called pretraining, and it is where almost all of the model's knowledge comes from. This module follows the three stages that turn raw text into the assistant you use, and this lesson covers the first.
Pretraining means training a language model on a vast collection of text to do one thing: predict the next token. It is the same guess, measure and adjust loop from lesson 2.2, applied to the next-token prediction you saw in lesson 3.2.
The text comes from many sources: public web pages, books, articles, forums, reference sites and a great deal of computer code. The builders filter it to remove spam, duplicates and some harmful material, but it remains a broad slice of what people have written. The model reads a passage, predicts each next token, compares its prediction with the token that really came next, and nudges its billions of dials. Then it does the same on the next passage, over and over, across the whole collection.
Nobody labels anything in this stage. The text supplies its own answers, because the right next token is simply whatever the author wrote next. That is what makes it possible to train on so much material.
Predicting the next token sounds like a narrow skill. But to do it well across that much text, the model has to absorb a lot.
To predict the next word of a sentence, it needs grammar. To continue "The capital of Malaysia is", it needs the fact. To continue a legal contract, a recipe or a Python function, it needs to know how each of those is structured. To finish a worked example in a textbook, it needs some grip on the steps of reasoning that lead to the answer. No one teaches it these things directly. They are picked up because they make the predictions better.
So a pretrained model ends up with broad general knowledge, the ability to write in many styles and languages, and patterns of reasoning. It also picks up whatever is wrong or lopsided in its text. If a myth is repeated more often online than the correction, the model may learn the myth. If some places are written about far more than others, it knows those places better. For many Singapore specifics, from CPF rules to the names of smaller neighbourhoods, there is much less text than for, say, American topics, so its knowledge is thinner and less reliable there.
And all of it stops at the date the text was collected. That is the knowledge cutoff, which lesson 4.4 has you test.
Here is the surprising part. At the end of pretraining you have a model that is very good at continuing text, and that is all it does.
Type "How do I renew my passport?" into a pretrained model and it continues the document it thinks it is reading. That might be an answer. It might equally be "How long does it take? What documents do I need? Where do I collect it?", because a list of questions on a government FAQ page is a likely thing to follow your question. It has no idea it is meant to help you. It will also continue harmful or rude text as readily as polite text, since all of that was in what it read.
Turning this raw predictor into something that answers questions, follows instructions and declines some requests takes two further stages: fine-tuning, in lesson 4.2, and learning from human feedback, in lesson 4.3.
Pretraining a model at the scale of today's leading assistants takes enormous computing power, running for weeks or months in large data centres, plus the work of collecting and cleaning the text. Only a handful of large companies and labs can afford it. That is why there are relatively few of these top models, often called frontier models, and why many other AI products are built on top of one of them rather than trained from scratch.
The later stages are far cheaper. A company can take an existing pretrained model and shape it for its own purpose without paying for pretraining again.
Go back to the new hire from the start of this lesson. You now know where their knowledge came from, so you can predict where it will be strong and where it will be patchy, and that prediction is what the activity is about.
Write three things a model would likely learn from reading most of the public web and two things it would likely get wrong because of what the web contains.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).