You will be able to explain hallucination as a direct result of next-token prediction rather than a rare glitch.
In 2023, two lawyers in New York filed a court brief in a personal injury case against the airline Avianca. The brief cited earlier court decisions to support their client's case. When the airline's lawyers went looking for those decisions, they could not find them, and neither could the judge. Several of the cases did not exist. One of the lawyers had used ChatGPT to research the brief, and the chatbot had produced case names, court details and quotes from judgments that were entirely made up. When he asked it whether one of the cases was real, it told him it was. The court sanctioned the lawyers, and the case, Mata v. Avianca, is now a standard warning in law schools and firms.
It is tempting to treat that as a rare glitch, or as a mistake only a careless person would make. This lesson argues the opposite. Given how language models work, invented answers like those are a normal output, and once you see why, you will know where to look for them.
A hallucination is output that reads as fact but is false or not supported by any source. The fake court cases are a dramatic example, but most are smaller: a statistic with a precise figure that nobody ever measured, a quote put in the mouth of a real person who never said it, a book title that sounds right for the author but was never written, or a regulation that is close to the real one with a single detail changed.
The word is a little misleading, because it suggests the model normally sees clearly and occasionally imagines things. That is not how it works. The model is doing exactly the same thing when it is right as when it is wrong.
Lesson 3.2 showed that a language model writes by predicting the next token, again and again, choosing from what is likely given everything before it. Lesson 4.1 showed that it learned those likelihoods from a vast amount of text.
Think about what a legal citation looks like in that text. There is a pattern: two party names, "v.", a volume number, the name of a law report, a page number, a court and a year. The model has seen that pattern many thousands of times. Asked for cases that support a point, the likeliest next tokens are ones that fit the pattern, with names and numbers of the kind that usually appear there. Whether that particular combination exists in any law report is not something the prediction takes into account.
The same goes for academic papers, news articles, URLs and figures. A plausible author, a plausible journal and a plausible year make a reference that looks just as likely as a real one. Often the model will produce a real reference, because famous papers and well-known cases appear often in its training data. For anything less famous, it fills the gap with something shaped right.
You might expect the model to check a fact before writing it down. On its own, it has no step that does that. There is no lookup against a database of cases, no comparison with the original report and no moment where it reads its claim back and asks whether it is true. It predicts, adds the token and moves on.
Some products add checking around the model. Lesson 5.3 showed how retrieval puts real sources in the window, and lesson 7.3 will show how a model can call a search tool. Those steps cut down invented answers a good deal. But the model writing the final text is still predicting, and it can still produce a claim its sources do not support.
Here is the part that catches people. A person who is unsure usually sounds unsure. They hedge, they pause, they say "I think". You have spent your whole life reading those signals, and you probably use them without noticing.
A language model does not give those signals reliably. The confident, well-organised style comes from how it was trained, including the human feedback in lesson 4.3, where answers that sounded complete and certain were often rated higher. It uses the same style whether the content is correct or invented. An invented CPF rule arrives in the same tidy paragraphs, with the same calm phrasing, as a correct one.
So tone tells you nothing about accuracy. Neither does detail. A made-up answer is often more specific than a real one, because specific-sounding text is what the pattern calls for. A paper with three named authors, a journal and a page range feels more trustworthy than a vague mention, and that is exactly the feeling to distrust.
The quickest way to believe this is to watch it happen on a subject where you can tell the difference. Pick a narrow topic you know well from your studies or your job, the kind where you could name a few real papers or reports yourself, and have that topic ready for the activity.
Ask an assistant for three academic papers on a narrow topic you know, then search for each title and record which exist and which details are wrong.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).