You will be able to describe the parts of a model API request and response and explain why the model keeps no memory between calls.
When you type into a chat assistant, a lot happens that you never see. When you build with a model API, you have to handle all of it yourself. The good news is that the core of it is simple: your program sends a message to the provider's server, and the server sends text back.
An API call is a request from your code to a provider such as OpenAI, Anthropic or Google, sent over the internet in the same way a browser asks a website for a page. The request carries four things. The first is your API key, which proves who you are and tells the provider whose account to bill. The second is the name of the model you want. The third is the conversation so far, written as a list of messages, each marked as coming from the system, the user or the assistant. The fourth is a set of settings, such as the longest reply you will accept.
The response comes back as structured data. It contains the model's reply and a usage count: how many tokens you sent and how many the model generated. A token is a chunk of text, often a short word or part of a longer one. Providers charge per token, and they usually price input tokens and output tokens differently. The prices change often, so you read them off the provider's pricing page when you plan a project rather than remembering a number.
Here is the part that surprises most people. The model remembers nothing between calls. Each request starts from a blank slate, and the only thing the model knows about your conversation is what you include in that request. A chat app feels like it remembers because the app resends the whole conversation every time you send a message. That has a cost. By the twentieth message, you are paying for the first nineteen again, and the request takes longer to process.
This one fact shapes a lot of design decisions. If you are building a tool that summarises customer feedback from a form, each summary can be a fresh call with only that piece of feedback in it. You do not need history, so each call stays small and cheap. If you are building a chat assistant for your team, you have to decide how much history to send, and when to trim or summarise the older parts.
The API key deserves its own warning. It works like a password tied to your credit card. Anyone who has it can run requests on your account until you notice. So the key lives on a server or in your own machine's environment settings, never inside a web page, a mobile app or a file you upload to a public code repository. If your app runs in a browser, the browser talks to your server, and your server talks to the model provider. Most providers also let you set a monthly spending limit, and you should set one before you write anything else.
You do not need to be a programmer to follow this course, but you do need to be willing to read and run code that an AI coding assistant writes for you. Every step in the course is explained in plain words first, so you know what the code is supposed to do before you see it.
Your task: open the developer console of one model provider, create an account, set a low monthly spending limit, and create an API key. Save the key in a password manager, not in a note or a chat. Then find the provider's pricing page and write down, in your own words, how it charges for input and output tokens.
Create a provider account, set a low monthly spending limit, create a key, store it in a password manager, and note how input and output tokens are priced.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).