You will be able to explain what fine-tuning does and why companies fine-tune models for specific jobs.
Think about the first week of a new customer service officer at a telco. They already speak English well and know what a phone plan is. What they learn in that first week is how to behave in the job: greet the customer, find out what they need, answer the actual question, keep it short, and never promise a refund they cannot authorise. Nobody reteaches them English that week, because the point of the training is how they use what they already know.
That is a good picture of fine-tuning. Lesson 4.1 left you with a pretrained model that knows a lot but behaves like a text continuer and may answer a question with more questions. Fine-tuning is the stage that teaches it to act like an assistant.
Fine-tuning means taking a model that has already been trained and continuing to train it on a smaller, carefully chosen set of examples. The method is the same guess, measure and adjust loop from lesson 2.2. What changes is the data.
Instead of raw web text, the examples are written to show the behaviour you want. For a general assistant, many of them are a question or instruction paired with a good answer. A typical pair might be: "Summarise this email in two sentences", followed by an actual email and a clean two-sentence summary. Another might be a request to explain a term simply, followed by a short plain explanation. These examples are often written or checked by people hired for the job, and there may be many thousands of them, which is tiny next to the pretraining text.
Because the model now trains on text where a question is always followed by an answer, the likeliest continuation of a question becomes an answer, and most of what feels different about an assistant follows from that.
Fine-tuning mostly teaches format and manner. From the examples, the model picks up habits like these:
Answer the question that was asked, rather than continuing the text in some other direction. Follow instructions about length, format and tone, such as "use bullet points" or "write it for a ten-year-old". Take turns in a conversation, as the assistant replying to a user. Decline some requests, such as instructions for causing harm, and do it politely.
This stage, sometimes called instruction tuning, is why the assistant you use feels like a helpful colleague rather than an autocomplete. Lesson 4.3 covers the next stage, learning from human feedback, which refines these habits further.
The same technique is used for narrower purposes. A company can take an existing model and fine-tune it on its own examples so that it writes in the house style, uses the right terms for its industry, or produces a particular format every time.
An insurer might fine-tune a model on hundreds of well-written claim summaries so that every new summary follows the same structure. A law firm might fine-tune on examples of its own memo style. A software company might fine-tune a model to write code in the way its engineers prefer. Because fine-tuning costs far less than pretraining, this is within reach of many organisations, and some AI providers offer it as a service on top of their models.
Here is the limit that people most often get wrong. Fine-tuning is very good at shaping how a model behaves. It is a poor way to teach it new facts.
Say your HR team wants an assistant that knows the company's leave policy. The obvious idea is to fine-tune a model on the policy document. The trouble is that a small number of examples barely shifts what the model knows, so it may still give a generic answer drawn from pretraining. Worse, it may blend the new policy with the many other leave policies it read about, and produce a confident mixture of both. And when the policy changes next year, the fine-tuned model is out of date until someone trains it again.
The better approach for facts is usually to give the model the document at the moment it answers, so the policy sits right in front of it. Module 5 covers how that works, in lesson 5.3, Retrieval: how an assistant looks things up. Many real products combine the two: fine-tuning for the manner and format, and a document lookup for the facts.
So when you think about fine-tuning, think about teaching manners and routines, the things you would show a new officer by example in their first week. Your own team has habits like that, the way it phrases replies, the format it expects, the requests it always turns down, and those make good material for the example pairs you write in the activity.
Write five example question-and-answer pairs you would use to fine-tune an assistant for your team, and note what behaviour each one teaches.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).