Tokens, pricing and limits you have to plan around

You will be able to estimate token counts and monthly cost for a feature, and name the limits that can stop a request.

Priya's manager liked the feedback summariser and asked the obvious question: what will it cost if the whole customer service team uses it every day? Priya had no idea. She knew the provider charged per token, but she did not know how many tokens a piece of feedback was, how long the summaries ran, or whether a busy morning would hit some limit and start failing.

Those are the three things you plan around before you build anything people depend on: how many tokens a task uses, what that costs, and which limits can stop a request.

Count tokens from real samples

A token is a chunk of text, and how text splits into tokens depends on the tokenizer each model uses. Common English words are often one token each. Long or unusual words split into several. Chinese, Tamil and other scripts often use more tokens for the same meaning than English does, and so do code, URLs and long strings of numbers. A message full of order references and Singlish abbreviations will not count the same as a tidy paragraph from a textbook.

That is why guessing from word counts goes wrong. Take real samples instead. Most providers offer a token counting tool or an endpoint that counts tokens without running the model, and every response includes a usage field with the exact input and output counts. Send ten typical inputs, read the counts, and work from the average. Include a few of the longest inputs too, because they set the size of your worst case.

Remember what goes into the input count. It is not only the user's text. It is the system prompt, any examples you include, any retrieved documents and, in a chat, the whole history you resend. In many apps the fixed instructions are bigger than the user's message.

Turn token counts into money

The cost of one request is simple arithmetic. Multiply the input tokens by the input price, multiply the output tokens by the output price, and add the two. Providers quote prices per million tokens, and output tokens usually cost more than input tokens. Prices change and differ by model, so take them from the provider's current pricing page on the day you plan, never from memory or an old blog post.

Here is a worked example with made-up prices, purely to show the method. Say Priya measures an average of 1,200 input tokens per feedback summary, most of it her system prompt, and 300 output tokens. Suppose the pricing page said S$3 per million input tokens and S$15 per million output tokens. One summary would cost 1,200 times 3 divided by a million, which is S$0.0036, plus 300 times 15 divided by a million, which is S$0.0045. That is S$0.0081 per summary, or S$8.10 for 1,000 summaries. Real prices are usually quoted in US dollars, so convert at the current rate.

Two habits keep this estimate honest. First, a feature that makes three model calls for each thing a user does costs three calls, so price the whole chain. Second, add a margin for retries and for the long inputs you found in your samples.

Limits that can stop a request

Every model has a context window. That is the most tokens it will handle in one request, and your input and the reply both count towards it. There is also a separate cap on how long the output can be, which you can lower with the maximum output setting. Send more input than the model accepts and the call fails with an error. Ask for a reply longer than the output cap, or set your maximum too low, and the reply simply stops mid-sentence. Lesson 2.3, Handle a bad response without crashing, shows how to detect that, because a cut-off reply is easy to mistake for a finished one.

Look up the context window and output cap for your chosen model on the provider's model documentation page. They differ widely between models and they change with new releases.

Rate limits and retries

Providers also cap how much you can send per minute. These rate limits usually count both requests and tokens, and they apply to your account or project as a whole, not to each user of your app. New accounts usually start on lower limits that rise as you spend more. Your provider's documentation explains the tiers.

When you go over, the provider rejects the request, usually with an HTTP 429 "too many requests" error. Your code should treat that as a reason to wait, not a reason to give up. The standard pattern is to retry after a short pause, and double the pause each time it fails again, up to a small number of attempts. Many official libraries do this for you, so check before writing your own.

For Priya, this means a busy Monday morning does not need special code, just a retry with a wait and a clear message to the user if every attempt fails.

In the activity below you will do Priya's homework for your own idea: real inputs, real token counts, and a cost estimate for 1,000 runs.

Take five real inputs for an idea of yours, count their tokens with the provider's tool or usage field, and estimate the cost of 1,000 runs.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).