Why the same question gets a different answer

You will be able to explain sampling and temperature, and why variation between answers is a feature you can use.

Two colleagues sit side by side and ask the same assistant the same question about how annual leave carries over under the Employment Act. One gets an answer that says unused leave can be carried forward. The other gets one that says it depends on the employment contract. Same words typed, same tool, same afternoon. Neither of them did anything wrong.

The difference comes from one step in the loop you met in lesson 3.2, It writes by predicting the next token, again and again. At every step the model produces probabilities for the next token, and then one token has to be chosen. How that choice is made decides whether you get the same answer twice.

The model rolls the dice, a little

The simplest way to choose would be to always take the single most likely token. That sounds safe, but in practice it tends to produce flat, repetitive text that can get stuck saying the same thing in circles. So most assistants do something else. They sample: they pick a token at random, weighted by the probabilities.

Go back to the example from lesson 3.2. If "the" has a 25 percent chance and "Monday" has 18 percent, then over many attempts, "the" gets chosen about a quarter of the time and "Monday" a bit less than a fifth. The likely tokens win most often, but not always. Over a few hundred tokens, those small random choices add up, and two answers to the same question can end up worded differently, structured differently, and sometimes saying different things.

Temperature turns the randomness up or down

Builders control how adventurous the sampling is with a setting usually called temperature. At a low temperature, the most likely tokens get an even bigger share of the chances, so the model behaves predictably and gives much the same answer each time. At a high temperature, the chances are spread more evenly, so less likely tokens get picked more often and the writing becomes more varied, more surprising and more prone to wander.

Most chat apps do not show you this setting. The company picks a sensible middle value for you. If your workplace uses a developer playground or builds its own tools on these models, you may see a temperature slider, and other sampling settings next to it, but the idea is the same.

Variation is a tool and a warning

Whether this randomness helps you or hurts you comes down to the job.

For brainstorming, it is exactly what you want. Ask for ten names for a hawker stall's new set meal, then ask again, and you get a fresh batch. Regenerating three or four times gives you a wider pool to pick from than any single answer would.

For facts, variation is a warning light. If you ask the same factual question twice in fresh chats and get two different answers, at least one of them is wrong, or the honest answer depends on details the question left out. Either way, you have learned that you cannot rely on a single reply. Check both against the source, such as the Ministry of Manpower's website for that leave question. Agreement between attempts is weaker evidence. Two answers can agree and still both be wrong, because the model can be consistently mistaken on something it learned badly.

One unlucky token can steer the rest

Because each token depends on everything before it, an early choice can send the whole answer down a path. Suppose you ask whether a particular expense is tax deductible. The probabilities for the first word might be split between "Yes" and "It". If sampling happens to land on "Yes", the rest of the answer will tend to justify a yes. If it lands on "It", the answer is more likely to explain that it depends on your circumstances.

The model cannot go back and reconsider its first word, so a slightly unlucky draw early on can produce a confidently wrong reply. That is why pressing regenerate sometimes fixes a bad answer completely, with nothing changed in your question. It is also why one strange reply does not prove the tool is useless, and one good reply does not prove it is reliable.

The practical habit that comes out of this is to treat a single answer as one draw from a range of possible answers. For creative work, take several draws and choose. For anything factual, look at how much the draws disagree, because that spread tells you how far to trust any one of them. Pick a question where you can look up the right answer afterwards, because seeing the spread on a fact you can verify is far more convincing than reading about it.

Ask an assistant the same factual question three times in fresh chats and record where the answers agree, where they differ and which differences matter.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).