You will be able to handle malformed, refused, truncated and slow responses in a predictable way.
Priya's summariser ran for a week on real feedback before anything went wrong. Then, on one afternoon, three things happened. A long complaint from a corporate client produced JSON that ended halfway through the issues list. A message containing a rude word came back as a polite refusal instead of a summary. And the provider was slow for ten minutes, so her team's page sat on a loading spinner until people gave up and refreshed it.
None of these were bugs in her code. They are normal behaviour for a model API under real traffic. What matters is whether your app has already decided what to do about each one.
Start by reducing how often things go wrong. Most models have a randomness setting, usually called temperature. Low values make the model pick its most likely wording, so the same input gives very similar output each time. High values let it pick less likely words, giving more variety.
Extraction and classification want low values. When the job is to pull a delivery date out of an email or label a complaint, you want the same answer on every run, and variety is just noise. Brainstorming, such as suggesting ten names for a product, wants higher values, because variety is the point. Priya set her summariser low and saw fewer odd labels straight away.
Some models fix this setting or ignore it, so check your model's documentation. And low randomness reduces variation without removing it, so the rest of this lesson still applies.
In lesson 2.2 you set up validation. Now decide what happens when it fails.
The pattern that works is one retry, then a fallback. On the retry, send the request again with the validation error included, for example: "Your previous reply failed validation: sentiment must be one of positive, negative or mixed. Reply again with valid JSON." Models are good at fixing a specific, named problem. A blind retry with the same request may repeat the mistake.
If the second attempt also fails, stop. Do not loop, because a request that fails twice will often keep failing and keep costing money. Send the case to a fallback. For Priya, the fallback marks the feedback "needs review" in the sheet and leaves it for a person. For other apps the fallback might be a simpler default value, an apology to the user, or a queue for later.
Log every failure and every retry, with the input that caused it. Patterns in that log tell you whether to fix your prompt, your schema or your validation.
Every response includes a field that says why the model stopped generating. Providers name it differently, often finish reason or stop reason. A normal ending says the model finished. A different value says it hit your maximum output length.
That second case is the one that broke Priya's corporate complaint. The reply was cut off mid-object, so the JSON was not even valid. Worse, in plain text mode a cut-off reply can look finished, ending at a full stop that happens to fall just before the limit.
So check the stop reason on every call, before you parse anything. If the output hit the limit, you have choices. Raise the maximum output length if the task genuinely needs more room. Ask for less, for example a cap of five issues. Or shorten the input. Do not retry the same request unchanged, because it will be cut off at the same place.
Refusals deserve a check too. Some providers return a refusal as a separate field or a distinct stop reason, others as ordinary text. Find out how yours does it and route refusals to your fallback rather than treating them as a summary.
A model call can take a long time when the provider is busy or the output is long. If your code waits forever, so does your user. Set a timeout on every call. Most official libraries let you pass one, and many have a default that is far longer than any user will wait.
Decide in advance what the user sees when the timeout fires. Priya chose a message that says "The summary is taking longer than usual. Your feedback has been saved and will be summarised shortly", and a background job that tries again later. The worst option is an endless spinner, because the user cannot tell whether to wait or give up.
Rate limit errors from lesson 1.3 sit in the same family. Retry after a pause, give up after a small number of attempts, and show a clear message.
The activity below asks you to do what Priya did after that afternoon: list every way a call can go wrong for your app, before it happens.
List every failure your feedback app could meet and write the action your code takes for each one.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).