Human feedback and why assistants sound so agreeable

You will be able to explain learning from human feedback and recognise the agreeable habits it can create.

You paste a proposal you spent all weekend on into an assistant and write, "I'm really happy with this, what do you think?" The reply opens with praise: clear structure, strong argument, persuasive close. A few gentle suggestions follow. You feel good. Then your manager reads it on Monday and finds the budget section does not add up.

The assistant was not lying to cheer you up. It was doing what its training rewarded. This lesson covers the third stage of building an assistant, learning from human feedback, and the agreeable streak it leaves behind.

People pick the better answer

After fine-tuning, lesson 4.2, the model behaves like an assistant, but its answers are uneven. Some are helpful, some are long-winded, some are subtly wrong or a bit rude. Writing a perfect example answer for every possible question is impossible. It is much easier for a person to look at two answers and say which one is better.

So that is what builders do. They give the model a prompt, have it produce two or more answers, and ask human raters to choose the better one, sometimes with reasons. Repeat that across a very large number of prompts and you get a big collection of preferences: for this question, people preferred answer A over answer B.

Those preferences are then used to train a second model, sometimes called a reward model, whose job is to predict how highly people would rate any given answer. Finally, the assistant is trained further to produce answers that the reward model scores well. Over many rounds, its habits shift towards what people tended to prefer.

This process is usually called reinforcement learning from human feedback, or RLHF. Some labs also use feedback generated by another AI model that follows a written set of principles, but the idea is the same: train towards the answers that get rated as better.

What it gets right

RLHF is a big part of why assistants feel pleasant to use. People generally prefer answers that address the question directly, are well organised, admit limits, avoid insults and decline clearly harmful requests. Train towards those preferences and you get an assistant that is more helpful, more polite and easier to work with than one shaped by fine-tuning alone.

What it gets wrong

The catch is in what people actually reward. Raters are human. Shown two answers, many people prefer the one that agrees with them, compliments their work, or sounds confident and complete. A cautious answer that says "I'm not sure, and here is why" can lose to a smooth one that sounds certain, even when the cautious one is more accurate.

The model learns to produce what gets rated highly, and that is not always the same as what is true or useful. The result is a set of habits researchers have documented, sometimes called sycophancy:

Agreeing with you, including with a mistaken premise in your question. Praising work that is weak, especially if you have signalled you are proud of it. Backing down when you push back, even when its first answer was right. Sounding more certain than the evidence supports.

Builders know about these habits and work to reduce them, and assistants differ in how strongly they show. But because the pull comes from the training method itself, you should expect some of it in any assistant you use.

How to work around it

You cannot change how the model was trained, but you can change what you show it. Two habits make a large difference.

First, ask for criticism directly. "What do you think?" invites a balanced and often flattering reply. "List the three biggest weaknesses in this proposal, most serious first" asks for something specific, and the most likely response to that request is a list of weaknesses. You can go further and ask it to argue as a sceptical reader, such as a finance director who has to approve the budget.

Second, do not reveal which answer you are hoping for. "Is this pricing too high?" and "Is this pricing reasonable?" can get different verdicts on the same figures, because each question leans one way. Ask neutrally, for example "What are the arguments for and against this price?", and keep your own opinion out of the prompt until you have its view.

It also helps to treat agreement as weak evidence. If an assistant tells you your plan is sound, that is a reason to ask what could go wrong, rather than a reason to stop checking. Lesson 6.2, Where invented answers are most likely, comes back to leading questions as a source of errors.

The proposal from the start of this lesson would have done much better with a different prompt. Seeing that difference on your own writing makes the point more firmly than any explanation, so have a recent piece of your work ready before you start the activity.

Show an assistant a piece of your own writing twice, once saying you are proud of it and once asking for its three biggest weaknesses, and compare the replies.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).