+200 XP

Function calling and structured outputs

The fastest way to turn a chatbot into a system is to hand the model a set of functions it can call and demand its answers come back as strict JSON your code can trust. This lesson goes deep on both mechanisms in the OpenAI API: tool calling (the loop where the model asks your code to run something) and Structured Outputs (the guarantee that the model's response matches your schema). They solve different problems and they compose beautifully.

We will use the Responses API, OpenAI's current primary interface. Chat Completions still works and uses near-identical concepts, but Responses is where the ecosystem is heading.

Two mechanisms, two jobs

Keep these separate in your head:

  • Function (tool) calling: the model decides it needs external data or an action, and emits a structured request to call one of *your* functions. Your code runs it, returns the result, and the model continues. This is how the model reaches outside its context: databases, your API, a calculator, the weather.
  • Structured Outputs: you force the model's *final answer* to conform exactly to a JSON Schema. No prose, no markdown fences, no "Sure, here's your JSON." Just valid, parseable data, every time.

You often use both: tools to gather facts, Structured Outputs to package the result.

Defining a tool

A tool is a JSON description of a function: a name, a description, and a parameter schema. The description is not decoration. **The model reads it to decide *when* and *how* to call the function**, so write it like documentation for a junior developer.

python
from openai import OpenAI

client = OpenAI()

tools = [{
    "type": "function",
    "name": "get_weather",
    "description": "Get the current temperature for a city in Celsius.",
    "parameters": {
        "type": "object",
        "properties": {
            "city": {"type": "string", "description": "City name, e.g. 'Lisbon'"}
        },
        "required": ["city"],
        "additionalProperties": False
    }
}]

Two details matter. required lists arguments the model must supply, and additionalProperties: False blocks the model from inventing extra fields. Both push toward predictable calls.

The tool-calling loop

Here is the part people get wrong: the model does not run your function. It returns a *request* to call it. You run the function, feed the result back, and call the model again. That round trip is the loop, and you own it.

The flow:

  1. Send the user message plus your tool definitions.
  2. The model replies with one or more function_call items (or a normal answer if no tool is needed).
  3. Your code executes each call and appends a function_call_output with the result.
  4. You call the model again with the updated input. It now writes the final answer.
python
import json

def get_weather(city):
    # Pretend this hits a real API.
    return {"city": city, "temp_c": 19}

input_list = [{"role": "user", "content": "What's the weather in Lisbon?"}]

response = client.responses.create(
    model="gpt-4.1",
    tools=tools,
    input=input_list,
)

# Carry the model's output forward as part of the conversation.
input_list += response.output

for item in response.output:
    if item.type == "function_call":
        args = json.loads(item.arguments)
        result = get_weather(**args)
        input_list.append({
            "type": "function_call_output",
            "call_id": item.call_id,
            "output": json.dumps(result),
        })

final = client.responses.create(
    model="gpt-4.1",
    tools=tools,
    input=input_list,
)
print(final.output_text)

Note call_id. The model can request several tool calls in one turn (parallel tool calling), and each output must be matched back to its call by id. Append every result before you make the follow-up request.

Practical rules for the loop

  • Loop, do not assume one round. After you return tool outputs, the model might call another tool. Wrap the request/execute step in a while loop that continues until no function_call items come back. Cap it (say, 8 iterations) so a confused model cannot spin forever.
  • Validate arguments before executing. The schema constrains the model, but treat tool arguments like any untrusted input. Never pass them straight into a shell, SQL string, or file path.
  • Keep functions narrow. get_order_status(order_id) beats one mega-function with a mode flag. Narrow tools are easier for the model to choose correctly and easier for you to secure.
  • Control choice when needed. Use tool_choice="auto" (default), "required" to force *some* tool, or name a specific tool to force exactly that one. Set parallel_tool_calls=False if your tools must run in sequence.

Structured Outputs: stop parsing prose

Tool calling handles *actions*. Structured Outputs handles the *shape of the answer*. When you set strict: true and supply a schema, the API constrains generation so the output provably matches your schema. This is stronger than the old "respond in JSON" prompt trick, which produced valid JSON most of the time and broke at 2am.

The cleanest path in Python is to define your schema as a Pydantic model and let the SDK parse it for you.

python
from pydantic import BaseModel

class CalendarEvent(BaseModel):
    name: str
    date: str
    participants: list[str]

response = client.responses.parse(
    model="gpt-4.1",
    input=[
        {"role": "system", "content": "Extract the event details."},
        {"role": "user", "content": "Standup with Ana and Rui on Friday."},
    ],
    text_format=CalendarEvent,
)

event = response.output_parsed
print(event.participants)  # ['Ana', 'Rui']

output_parsed hands you a typed object, not a string you have to json.loads and pray over. If you are not in Python, you pass a raw JSON Schema in text.format with "type": "json_schema" and "strict": true, and you get back guaranteed-conformant JSON text.

Schema design that actually works

  • Every property is effectively required. Strict mode treats all keys as required. To make a field "optional," give it a union with null, for example Optional[str] in Pydantic, then check for null in your code.
  • Use enums to constrain choices. If status can only be open, pending, or closed, define it as an enum. The model cannot return in-progress and surprise you downstream.
  • Descriptions guide values. Field descriptions in the schema steer *what* goes in each field, not just the types. Use them.
  • Mind the supported subset. Structured Outputs supports a defined slice of JSON Schema. Patterns like minimum, maximum, and some format constraints may not be enforced. Check the Structured Outputs guide before relying on a keyword.

OpenAI Function Calling and Structured Outputs Explained

Watch on YouTube

Combining both: the realistic pattern

Most production features use the two together. Imagine a support assistant that answers a refund question:

  1. The model calls lookup_order(order_id) (tool calling) to fetch real data.
  2. You return the order record.
  3. The model produces a final answer constrained to a RefundDecision schema (Structured Outputs): a boolean eligible, an enum reason, and a customer_message string.

Your application code never parses free text. It reads decision.eligible and branches. That is the whole point: the model handles language and judgment, your code handles control flow, and the boundary between them is a typed contract.

A subtlety worth knowing: tool *parameters* and Structured *Outputs* are separate schema slots. One shapes the call going in, the other shapes the answer coming out. You can use either alone or both at once in the same request.

Knowledge check

1. What is the fundamental difference between function (tool) calling and Structured Outputs?

2. In the tool-calling loop, what actually happens when the model 'calls' a function?

3. Why does the lesson emphasize writing a tool's 'description' like documentation for a junior developer?

MULTIPLE CHOICE

4. Select ALL correct statements about the tool parameter schema fields discussed in the lesson.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL scenarios that correctly reflect the intended purposes described in the lesson.

Select all the correct answers.

Errors, costs, and failure modes

Real systems break in specific ways. Plan for them.

  • The model hallucinates a tool that does not exist, or bad arguments. With strict tool schemas this is rare, but if a function genuinely fails, return an error *as the tool output* (for example {"error": "order not found"}) rather than throwing. The model can read that and recover, asking the user for a correct order id.
  • Refusals. Structured Outputs can still refuse on safety grounds. The SDK surfaces this so you check for a refusal field before trusting output_parsed. Handle it explicitly instead of treating a refusal as malformed data.
  • Truncation. If the model hits the output token limit mid-JSON, you get incomplete data even with strict mode, because the constraint guarantees *valid structure*, not *completion*. Check the response status for incomplete and raise your max_output_tokens for large objects.
  • Latency on first use of a new schema. The very first request with a novel strict schema may be slower while the API prepares the constraint. Repeated schemas are fast. Reuse schemas rather than generating them dynamically per request.
  • Tokens. Tool definitions and schemas live in the context window and count as input tokens on every call. Long, verbose schemas across many tools add up. Keep descriptions tight and only attach tools the model might plausibly need for this request.

Where this sits in the ecosystem

You have now seen the raw machinery. Higher-level OpenAI products are built on exactly these primitives:

  • GPT Actions in Custom GPTs are function calling driven by an OpenAPI spec. You describe your API once, and the GPT calls it through the same request/execute pattern, just managed by ChatGPT instead of your code.
  • The Agents SDK wraps the tool loop, retries, and handoffs so you stop writing the while loop by hand. When your tool orchestration gets complex (multiple agents, guardrails, tracing), graduate to it rather than maintaining bespoke loop code. See the Agents SDK docs.
  • Built-in tools like web search and file search are function calls the API runs server-side. You enable them in tools without writing the executor.

Understanding the bare loop first means none of these feel like magic. They are conveniences over the contract you just learned.

Key Takeaways

  • Tool calling is a loop you own. The model *requests* a function; your code runs it, returns the output keyed by call_id, and calls the model again. Loop until no tool calls remain, with a hard iteration cap.
  • Use `strict: true` Structured Outputs for any answer your code parses. It guarantees the shape, eliminating brittle prompt-and-pray JSON. In Python, define a Pydantic model and read output_parsed.
  • Design schemas defensively. Treat all fields as required (use null unions for optional), constrain choices with enums, write field descriptions, and confirm your keywords are in the supported subset.
  • Return tool errors as data, and check for refusals and truncation. Let the model recover from failed calls; verify response status before trusting the payload.
  • Reach for the Agents SDK or GPT Actions when the loop grows. They are built on these exact primitives, so the mental model transfers directly.

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Start new code on the Responses API and read response.usage from day one
See the full action playbook →