The API: your first real calls
The fastest way to outgrow the ChatGPT window is to make the same model answer from inside your own code, and with OpenAI that takes about ten lines. This lesson is about those ten lines: what is OpenAI-specific, why the Responses APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition → is now the default, and how to stream tokenstokensA token is the basic unit of text that language models process, often a word fragment, whole word, or punctuation mark rather than a single character.View full definition → as they arrive.
Set up once, then forget it
Get a key from the API keys page. Treat it like a password: it is tied to billing, and anyone holding it can spend your credits. Never paste it into code or commit it to git.
Set it as an environment variable so the SDK finds it automatically:
export OPENAI_API_KEY="sk-proj-..."
pip install openaiThe official Python SDK reads OPENAI_API_KEY on its own, so you never write the key in your script. One important boundary: your API account and your ChatGPT subscription are separate. ChatGPT Plus or Pro does not give you API credits, and API usage is billed per token against a different balance. They share models, not wallets.
Your first real call
Here is a complete, runnable script using the Responses API, OpenAI's current primary interface:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-4.1-mini",
instructions="You are a terse assistant. Answer in one sentence.",
input="Explain what an idempotent API request is.",
)
print(response.output_text)Run it and you get a single clean sentence back. Now let's unpack the parts that are specific to OpenAI.
client = OpenAI()
This builds the client and silently picks up your environment key. Everything you do flows through this object: text generation, embeddings, files, audio. You configure it once.
model="gpt-4.1-mini"
The model string is a real, billable choice, not a label. OpenAI ships several families: the GPT-4.1 line for general-purpose work, the smaller mini and nano variants for cheaper and faster calls, and the reasoning models (the o-series, like o4-mini) that think longer before answering. Pick by job. A classifier or a formatter wants mini or nano. A multi-step planning task wants a reasoning model. Check the current roster and pricing on the models page, because the lineup shifts.
instructions vs input
This is the cleanest improvement the Responses API brings. instructions is the system-level steering (personapersonaA semi-fictional, research-based representation of your ideal customer: their goals, frustrations, behaviours and decision criteria.View full definition →, rules, output format). input is the actual user turn. You no longer hand-build a list of role-tagged message dictionaries for simple calls. For a single exchange, two strings is all you need.
response.output_text
A convenience accessor that flattens the response into the final text string. The full response object carries more: token usage, the model that actually served you, tool calls, and structured content blocks. For quick work, output_text is what you want; for production, you will read the richer fields.
Responses or chat completions?
You will see both in tutorials, so know the difference.
Chat Completions (client.chat.completions.create) is the older, widely-copied interface. It takes a messages list of {"role": ..., "content": ...} dicts. It still works and is not deprecated, so existing code is safe.
Responses (client.responses.create) is what OpenAI now recommends for new projects. It is stateful-friendly, has built-in tools (web search, file search, code interpreter) you can switch on without wiring them yourself, and it handles multi-step tool use more cleanly. The Responses API docs are the canonical reference.
Rule of thumb: new code, start with Responses. Only reachreachThe number of unique people exposed to your message in a given period. Unlike impressions, reach counts each person once, no matter how often they see it.View full definition → for Chat Completions if you are extending something that already uses it.
Streaming: tokens as they arrive
Waiting for a long answer to fully generate before showing anything feels broken to users. Streaming sends tokens as the model produces them, which is exactly the typewriter effect you see in ChatGPT.
You opt in with stream=True and iterate over events:
from openai import OpenAI
client = OpenAI()
stream = client.responses.create(
model="gpt-4.1-mini",
input="Write a two-line haiku about slow APIs.",
stream=True,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
print()The stream yields typed events, not just raw text. You filter for response.output_text.delta to catch the incremental text chunks. Other event types announce when the response starts, when a tool is called, and when generation completes. The flush=True forces each chunk to the terminal immediately instead of buffering.
Streaming changes nothing about cost or the final result. It only changes *when* you see the bytes. Use it for any user-facing interface; skip it for batch jobs where nobody is watching.
Three OpenAI-specific things worth knowing early
Structured Outputs
When you need the model to return data your code can parse, do not beg it for JSON in the prompt and hope. Use Structured Outputs, which forces the response to match a schemaschemaA schema is the formal blueprint that defines how data is structured, named, typed, and related within a database, file, or message.View full definition → you define. You pass a JSON Schema (or a Pydantic model in Python) and OpenAI guarantees the shape, so you never write defensive parsing for a stray markdown fence again.
from pydantic import BaseModel
from openai import OpenAI
client = OpenAI()
class Ticket(BaseModel):
priority: str
summary: str
response = client.responses.parse(
model="gpt-4.1-mini",
input="My checkout button has been broken for two days, losing sales.",
text_format=Ticket,
)
ticket = response.output_parsed
print(ticket.priority, "|", ticket.summary)response.output_parsed hands you a typed Ticket object. This is the single biggest reliability upgrade for anyone building real features. The Structured Outputs guide covers the schema rules.
Function calling
The model can decide to call functions you expose, returning the function name and arguments for you to execute. You describe your functions, the model picks one when relevant, you run it, then you feed the result back. This is the primitive under tool-using assistants. Note the boundary: the API's function calling is not the same as GPT Actions in Custom GPTs. Actions are the no-code version inside the ChatGPT product; function calling is the raw mechanism you control in code.
OpenAI Function Calling Explained
The agents SDK
When one function call becomes a loop of planning, calling tools, and checking results, you graduate to the Agents SDK. It is OpenAI's official framework for multi-step agents: it manages the tool-call loop, handoffs between specialized agents, and guardrailsguardrailsRules and controls that keep an AI system inside safe, legal and on-brand boundaries, blocking outputs and actions that cross the line.View full definition →, so you do not hand-roll the orchestration. You do not need it for your first calls, but it is the natural next step once a single Responses call is not enough. It builds directly on top of the Responses API.
Knowledge check
1. Why should you set your OpenAI key as an environment variable rather than writing it directly in your script?
2. A colleague argues that their ChatGPT Plus subscription should cover their API usage. What is the correct clarification?
3. You need to build a fast, cheap text classifier that just labels incoming messages. Which model choice best fits the job?
4. Select ALL correct statements about the Responses API and the OpenAI client as described in the lesson.
Select all the correct answers.
5. Select ALL correct statements about the distinction between 'instructions' and 'input' in the Responses API.
Select all the correct answers.
Reading the response object properly
For anything beyond a demo, stop printing output_text and start inspecting the whole object. Two fields matter immediately.
Usage. Every response reports token counts:
print(response.usage.input_tokens, response.usage.output_tokens)This is how you measure cost. Input tokens (your prompt and instructions) and output tokens (the generation) are priced differently, and output is usually the more expensive side. Logging usage from day one saves you from a surprise bill later.
The served model. The model field tells you which exact model version answered. When you pin a dated snapshot versus an alias that auto-updates, this field is how you confirm what ran. For reproducible production behavior, pin a specific snapshot rather than a floating alias.
Errors you will actually hit
Real calls fail for real reasons. Handle the common ones explicitly:
from openai import OpenAI, RateLimitError, APIError
client = OpenAI()
try:
response = client.responses.create(
model="gpt-4.1-mini",
input="Summarize the API economy in one line.",
)
print(response.output_text)
except RateLimitError:
print("Slow down or upgrade your tier, then retry with backoff.")
except APIError as err:
print(f"OpenAI-side issue: {err}")RateLimitError is the one beginners meet first. New API accounts start on a low usage tier with modest per-minute limits; tiers rise automatically as your account ages and you spend. Do not treat a rate limit as a bug. Treat it as a signal to add retry-with-backoff, which the SDK can do for you via the max_retries setting on the client. APIError and its subclasses cover authentication, bad requests, and transient server errors. Catch them so one hiccup does not crash your app.
Keep latency and cost sane
A few habits pay off immediately:
- Cap output. Set
max_output_tokensso a chatty model cannot run long and run up the bill. A summarizer rarely needs more than a couple hundred tokens. - Right-size the model. Most production traffic does not need your largest model. Route simple calls to a
miniornanoand reserve the heavy or reasoning models for tasks that genuinely need them. - Stream user-facing, batch the rest. Perceived speed comes from streaming; real throughput comes from sending independent calls concurrently rather than in a slow loop.
These three levers (output cap, model choice, concurrency) cover most of the cost and latency tuning you will do early on.
Key Takeaways
- Start new code on the Responses API (
client.responses.create) withinstructionsandinput; keep Chat Completions only for existing codebases. - Set `OPENAI_API_KEY` as an environment variable and remember your API balance is separate from any ChatGPT subscription.
- Use Structured Outputs (`responses.parse` with a Pydantic model) whenever you need parseable data, instead of promptingpromptingPrompt engineering is the practice of designing and refining text inputs to guide large language models toward accurate, relevant, and reliable outputs.View full definition → for JSON and hoping.
- Read `response.usage` from day one and cap
max_output_tokensso cost stays visible and bounded. - Stream for any user-facing interface by filtering
response.output_text.deltaevents, and catchRateLimitErrorwith backoff rather than letting it crash your app.
What to do, from this lesson
These actions are compiled in the role's Playbook.
- Start new code on the Responses API and read response.usage from day one