You will be able to describe how a language model produces a full answer through repeated next-token prediction.
Type "Your polyclinic appointment is on" into a message and your phone's keyboard suggests the next word: perhaps "Monday", "the" or "Tuesday". Tap one, and it suggests another. Keep tapping and you get a sentence, usually a bland one.
A large language model works on the same basic idea, taken much further. It predicts the next piece of text, adds it, and predicts again. The difference is in how much it has learned and how much of the text it takes into account. That is enough to turn a keyboard trick into something that can draft a report or explain a tenancy clause.
Start with what happens at one step. The model takes in all the text so far, which means your question, any instructions the app added, and whatever it has already written of its reply. It runs that text, as tokens, through its billions of learned settings. Out comes a score for every token in its vocabulary, turned into a probability: how likely each one is to come next.
For "Your polyclinic appointment is on", the model might give "Monday" 18 percent, "the" 25 percent, "Tuesday" 15 percent, "Friday" 12 percent, and tiny amounts to every other token in its vocabulary, including "banana". Those numbers are invented for illustration, but the shape is real. There is no single answer, only a spread of likelihoods learned from all the text the model was trained on.
Next, one token is chosen from that spread. Lesson 3.3 explains how the choice is made and why it is not always the most likely one. The chosen token is added to the end of the text.
Then the whole process runs again. The model now reads "Your polyclinic appointment is on the" and produces a fresh set of probabilities. Dates and numbers now rank highly, because "the" changed what is likely. Another token is chosen and added. And again.
This repeats, one token at a time, until the model produces a special token that means "end of answer", or hits a length limit. A three-paragraph reply is a few hundred of these steps. When you watch an assistant's answer appear on screen bit by bit, you are seeing this loop happen roughly in real time.
This is the part that changes how you read AI answers. The model does not write a finished answer somewhere out of sight and then type it out. Each token is chosen based on everything before it, and once a token is written, it stays. The model cannot go back and revise its opening sentence in light of how the paragraph ended.
Researchers who study the inside of these models have found that they do look ahead a little. When asked to write a rhyming line, for example, a model can settle on the rhyme word before it writes the words that lead up to it. But whatever it anticipates, it still commits to its answer one token at a time, with no step where it reads the whole thing back and corrects it.
That explains a familiar behaviour. If an assistant starts with "Yes, you can claim that", everything after has to fit a yes, even if the honest answer turns out to be no. Some assistants now offer a thinking or reasoning mode, where the model first writes out working notes and then the answer. That is still next-token prediction. It simply gives the model room to work through the problem in writing before it commits to the part you read.
It is tempting to conclude that the model is only autocomplete and therefore produces fluent filler. That undersells it. To predict the next token well across billions of examples of human writing, the model had to pick up a great deal: grammar, facts that appear often, how arguments are built, how code fits together, how a summary relates to the text it summarises.
So when you ask it to summarise a meeting transcript, the most likely next tokens are ones that summarise it, because that pattern was all over its training data. When you ask it to work out the GST on a S$480 invoice, likely next tokens follow the steps a person would write. Prediction is the method, and the skill comes from what was learned in order to predict well.
The same mechanism also explains the failures you will study in module 6. The model produces what is likely, which is usually what is true, but not always. A plausible-sounding regulation can be a likely sequence of tokens whether or not it exists.
You can get a feel for this without any software. Playing the model yourself for a few steps, with nothing more than a pen, shows how much each early choice decides what follows.
Write the first five words of a sentence, list three likely next words with rough probabilities you would assign, and note how your choice changes the sentence after it.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).