You will send a working request from your own machine and read every part of the response.
You have an account, a spending limit, a key stored as an environment variable and a rough sense of what a call costs. You have not yet sent a request yourself. This exercise closes that gap. By the end you will have a short script on your own machine that calls a model, and you will be able to point at every part of what comes back.
Set aside about half an hour. You need your provider account from lesson 1.1, your key set up as in lesson 1.2, Keep your API key on the server, never in the browser, and an AI coding assistant.
Every major provider has a playground or workbench in its developer console, a web page where you type a prompt and adjust the settings without writing code. Start there. Pick a model, write a short system message such as "You summarise customer feedback in one sentence", and paste in a piece of feedback as the user message. Run it.
Now look at the settings panel. You will see the model name, a maximum output length and usually a randomness setting called temperature. Many playgrounds also have a button that shows the request as code. Open it. That code is the same request you are about to send from your own machine, and seeing it first means nothing in your script will be a surprise.
Open your AI coding assistant in an empty project folder and ask for something specific. Priya's request read roughly like this: "Write a short Python script that calls the model API using the provider's official Python library. Read the key from the environment variable MODEL_API_KEY. Send one system message and one user message. Print the full raw response, then print the reply text and the input and output token counts on separate lines."
Two details in that request matter. The official library is the one the provider publishes and documents, so it handles details like retries and current request formats. Ask for it by name if you are unsure which one it is. And naming the environment variable means the assistant has no reason to write the key into the file. If it does anyway, refuse the change.
Install the library with the command the assistant gives you, then run the script from your terminal.
The first run prints a block of structured data. Do not skip to the reply. Work through it field by field.
You will find an ID for the response, the model that actually answered, the reply itself (often nested inside a list, because some settings let you ask for more than one), a reason the reply stopped and a usage section. Priya's run, as a worked example, showed 41 input tokens and 23 output tokens, with a stop reason saying the model had finished naturally.
Now open your provider console's usage page. It can take a few minutes to update. Find your request and check that the token counts match what your script printed. This matters, because those usage numbers are what you are billed on, and you now know exactly where to read them in your own code.
Experiment, but change only one thing per run so you know what caused each difference.
First, set the maximum output length very low, for example 10 tokens, and run it again. The reply stops mid-sentence and the stop reason changes to say it hit the limit. Remember that field. It is how your code will notice a cut-off reply later in the course.
Second, put the maximum back and run the same request three times. The wording usually changes a little between runs. Then lower the temperature, where your model offers it, and run three more. The replies should be more alike. Some newer models do not accept a temperature setting at all, in which case the documentation will say so.
Third, make your system message much longer, say a full paragraph of instructions, and watch the input token count rise even though your user message has not changed.
Write down what you saw after each change.
A finished exercise has three things. A script that runs from your terminal and reads the key from the environment, with no key anywhere in the file. A saved copy of one full response. And notes beside that response that say in your own words what each field means, plus what happened when you changed each setting.
If the script fails, read the error before changing anything. An authentication error usually means the environment variable is not set in the terminal you ran it from. A 429 error means a rate limit, as lesson 1.3, Tokens, pricing and limits you have to plan around, explained, and you should wait and try again. Paste any other error in full into your assistant and ask what it means before asking it to fix anything.
Once your notes would make sense to someone who has never seen an API response, you are ready for the task below.
Run a script that sends one request, prints the reply and the token usage, and save the output with a note on what each field means.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).