You will be able to explain what tokens are and predict when text will use more tokens than you expect.
Ask an AI assistant how many letter r's are in "strawberry" and, for a long time, many of them got it wrong. The same assistant could summarise a fifty-page report and draft a polite reply to an angry customer. It seems absurd that something so capable would fail a task a primary school child can do. The reason is that the assistant never saw the letters in the first place.
Before a language model reads anything you type, your text is cut into pieces and turned into numbers. Those pieces explain a surprising amount of the model's behaviour, from what it costs to run to some of its oddest mistakes.
A token is the chunk of text a language model works with. Often it is a whole common word such as "the", "and" or "meeting", while longer or rarer words get split into several tokens, and punctuation marks and spaces usually count as tokens or as part of one.
The splitting is done by a separate program called a tokenizer, before the model itself gets involved. The tokenizer has a fixed vocabulary of chunks, built by finding which strings of characters appear most often in a huge amount of training text. Frequent words earn a token of their own. Less frequent ones are built from smaller pieces.
So "appointment" might be a single token, while "Kallang" might be split into two or three pieces, perhaps something like "K", "all" and "ang", although that split is only an illustration because each tokenizer cuts text its own way and different assistants use different tokenizers.
The model cannot calculate with text. So each token is swapped for a long list of numbers that the model learned during training. Think of it as a set of coordinates for that token, where tokens used in similar ways end up with similar coordinates. "Clinic" and "hospital" sit closer together than "clinic" and "durian".
Everything after that point is arithmetic on those lists. The model never looks at individual letters unless a letter happens to be a token on its own. When it reads "strawberry", it may see two or three chunks, each one a list of numbers, with nothing in them that directly says how many r's are inside.
Because the vocabulary was built from training text, and much of that text was English, common English words tend to be cheap: one word, one token. Other kinds of text split into many more pieces.
Rare words and names, including many Singapore place names, Malay and Hokkien terms, and brand names, often break into several tokens. Text in other scripts, such as Chinese, Tamil or Thai, often uses more tokens for the same meaning, though how many depends on the tokenizer. Numbers split into chunks of digits. A long bank account number or an NRIC-style string becomes a run of separate pieces. Computer code, with its symbols, brackets and spacing, tends to use more tokens than ordinary prose.
This matters for two practical reasons. Assistants measure their limits in tokens, so a document in Chinese or full of figures fills those limits faster than the same length of plain English. Module 5 covers that limit, the context window. And services that charge by usage, which you meet if your company builds on these models, usually count tokens, so token-heavy text costs more.
Once you know about tokens, several strange failures make sense.
Counting letters is hard because the letters are hidden inside the chunks. The model has to have picked up, from its training, which letters each chunk contains, and that knowledge is patchy. Ask it to spell the word out letter by letter first and it often does better, because each letter then becomes its own token that it can see.
Long numbers are risky for a similar reason. A sixteen-digit card number may be split into uneven groups of digits, and the model can slip when copying it, reversing it or doing arithmetic with it. If a figure from an assistant really matters, such as a total on a bank statement or a reference number, check it against the source.
Rhymes, word puzzles and anagrams fail in the same way. These tasks depend on letters and sounds, while the model works in chunks of meaning. Newer assistants often handle such tasks better, sometimes by writing the word out in pieces or by running a small program to do the counting, a trick you will meet in lesson 7.3. But the underlying cause has not gone away.
Seeing your own text cut into tokens makes all of this concrete in a couple of minutes. Several free online tokenizers show the split in colour as you type, and most people are surprised by which of their everyday sentences turn out to be the expensive ones.
Paste three short texts into a free online tokenizer, one in English, one with numbers and one in another language you know, and note which used most tokens.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).