Cut cost and latency without hurting quality

You will be able to apply and measure common ways to make an AI feature cheaper and faster.

Three months in, Priya's summariser had spread beyond her team. Customer service in two other regions were pasting in their feedback too, and the monthly bill had grown with them. Users had a second complaint: a batch of 50 items took long enough that people switched tabs and forgot to come back. Her manager asked if it could be cheaper and faster. Priya's first instinct was to switch everything to the cheapest model she could find.

That might work. It might also quietly make every summary worse. This lesson covers the common ways to cut cost and waiting time, and the rule that goes with every one of them: measure before and after, and check quality with your eval.

Measure first

Before changing anything, find out where the cost and time actually go. Run your eval set from module 7 and record, for each case, the input tokens, output tokens, cost and response time. Your token counts come from the usage field, as in lesson 1.4, and the prices from your provider's current pricing page.

Priya found that her system prompt, with its examples and schema, made up most of the input tokens on every call. The feedback itself was usually short. Output was small. Time was dominated by running 50 calls one after another.

That told her where to look. Without measuring, she would have guessed.

Use the smallest model that does the job

Providers offer models of different sizes. Smaller ones cost less per token and usually respond faster. Larger ones handle harder reasoning, longer documents and subtle judgement better.

Many steps in an app are simple: classify this ticket into one of five labels, extract a date, check whether a message is in English. A small model often does these as well as a large one. Hard steps, such as summarising a long, messy complaint or answering a tricky policy question, may need the larger model.

So assign models per step, not per app. And never switch on faith. Change the model for one step, rerun your eval, and compare case by case exactly as in lesson 7.4, Run an eval before and after a prompt change. If quality holds, keep the cheaper model. If some case types drop, keep the larger model for those, or for the whole step.

Priya tried a smaller model for her summaries. The eval showed normal cases unchanged but two of her six non-English cases slipping. She kept the smaller model and added a rule that sends non-English feedback to the larger one.

Send less, and cap what comes back

Every token you send costs money and processing time, so trim what each call carries. Shorten the system prompt: remove instructions that the eval shows make no difference, and keep only the examples that earn their place. Send only the context a step needs. A retrieval step from module 3 might send the top three chunks instead of the top ten, if your eval shows the answers hold.

Cap output length too. Set the maximum output to what the task needs plus a margin, and ask for concise formats, such as a 25-word summary rather than a paragraph. Output tokens usually cost more than input tokens and take longer to generate, so shorter outputs help on both counts. Keep checking the stop reason, as lesson 2.3 explained, so a cap that is too tight shows up as a failure rather than a silently cut-off reply.

Caching and batching

Two provider features target specific patterns. Check your provider's documentation for whether and how it supports each, and for the current pricing.

Prompt caching helps when many requests start with the same long text, like Priya's system prompt and examples. The provider stores the processed form of that repeated beginning, and later requests that start the same way are charged at a lower rate for that part and often process faster. Some providers do this automatically above a certain length. Others need you to mark which part to cache. To benefit, put the unchanging content first and the varying content, like the feedback, last.

Batch processing helps with work that does not need an answer straight away. You submit many requests together, the provider processes them within a stated window rather than immediately, and charges less. Priya's nightly summary of the day's website feedback suits this. Her team's live paste-and-summarise screen does not.

Make waiting feel shorter

Some delay cannot be removed. You can still make it easier to sit through.

Streaming sends the reply to the user piece by piece as the model generates it, instead of all at once at the end. The total time is the same, but the user sees text appear almost immediately, which feels far quicker than staring at a spinner. Streaming suits text a person reads as it arrives. It is less useful for JSON your code must parse whole.

For Priya's batches, the bigger win was running calls in parallel rather than one after another, within her rate limits, and showing each row as soon as it was ready.

The activity below has you measure your own app, apply two of these changes and check with your eval that quality held.

Measure the cost and response time of your app on your test set, apply two changes, and rerun the eval to check quality held.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).