+160 XP

The OpenAI model family: GPT, reasoning models, and when to use each

OpenAI ships two fundamentally different kinds of behavior under one roof: fast general answering, and deliberate reasoning that thinks before it responds. Picking the wrong mode wastes either your money or your accuracy. This lesson is about routing: matching the task to the model (or the reasoning level) so you stop overpaying for trivial work and stop under-powering hard work.

Two behaviors, converging into one product

Historically OpenAI split these into two named lineups: the GPT "chat" models (GPT-4o and friends) and the "o" series reasoning models (o3, o4-mini). With the GPT-5 family, that split has largely collapsed into a single model that decides how hard to think, but the underlying tradeoff is exactly the same, so the routing logic below still holds.

Fast general answering is optimized for breadth and latency. You send a prompt, tokens stream back almost immediately. This mode is excellent at writing, summarizing, classification, formatting, code completion, and conversation.

Reasoning spends extra hidden compute working through a problem before emitting the final answer. That hidden work is called *reasoning tokens*: internal steps you pay for and never see. It trades latency and cost for accuracy on multi-step problems: math, hard debugging, planning, scientific analysis, and anything where a wrong intermediate step ruins the result.

The family also comes in smaller variants (mini and nano tiers). Same general capability profile, lower cost, lower latency, slightly less depth. A `mini` model with reasoning turned up still "thinks," just less expensively.

The official map of who's who and what each is tuned for lives in the models documentation. Check it when names change; the routing logic below does not.

The routing decision in one question

Ask: does this task have a verifiable chain of steps where an early mistake breaks the outcome?

  • No → fast general answering. Rewrites, tone changes, drafting, extraction, summarization, simple Q&A, casual chat.
  • Yes → turn up reasoning. Multi-constraint analysis, math proofs, ambiguous bug hunts, multi-document synthesis with logic, planning a sequence of actions.

Then ask a second question for cost: is this high-volume or latency-sensitive? If yes, drop to a `mini`/`nano` variant.

Concrete example: rewrite vs analysis

Task A: "Rewrite this paragraph to be more concise and confident."

This is a one-shot transformation. There is no chain to get wrong. Keep reasoning low (or off) and let it answer fast.

python
from openai import OpenAI
client = OpenAI()

resp = client.responses.create(
    model="gpt-5",
    reasoning={"effort": "minimal"},
    input="Rewrite to be concise and confident:\n\n" + paragraph,
)
print(resp.output_text)

Task B: "Here are three vendor contracts. Find every clause where renewal terms conflict, and tell me which contract wins under the precedence rules in section 12."

That is multi-step: extract clauses, compare them, apply precedence logic, resolve conflicts. An early misread cascades. Turn reasoning up and let it spend tokens.

python
resp = client.responses.create(
    model="gpt-5",
    reasoning={"effort": "high"},
    input=contracts_text + "\n\nFind conflicting renewal clauses and resolve per section 12.",
)
print(resp.output_text)

Two things to notice. First, both calls use the Responses API (responses.create), OpenAI's current primary interface. Second, the `reasoning.effort` knob (`minimal` / `low` / `medium` / `high`) trades depth for cost and latency. On the GPT-5 family this is your main lever: same model, different amount of thinking. If you still target an older reasoning model like o3 or o4-mini, the same parameter applies, though those are now legacy choices in most workflows.

In ChatGPT: the picker in plain terms

In the ChatGPT apps the default is now a single GPT-5 model that auto-routes: it decides internally whether to answer fast or think harder. You still get manual control:

  • The default handles roughly 80% of daily work without you touching anything.
  • An explicit Thinking option forces deliberate reasoning. Switch to it when a fast answer keeps coming back shallow or subtly wrong.
  • When auto-routing is on, your real lever is the prompt ("think carefully and check each step") rather than a dropdown, because the model reads that as a cue to spend more reasoning.

Older o-series names (o3, o4-mini) may still appear under legacy or advanced menus for some plans, but the GPT-5 family is the default path forward.

A practical habit: start fast, escalate to explicit thinking only when you catch the model skipping logic. Thinking costs more and is slower, so do not make it your default.

Where this intersects ChatGPT's features

Model choice, or how hard you tell it to think, changes how the rest of the ChatGPT toolset behaves.

Advanced Data Analysis (Code Interpreter). When you upload a CSV and ask for analysis, the model writes and runs Python in a sandbox. High reasoning effort is far better at planning a correct multi-step analysis (clean, join, aggregate, validate) before writing the code. A fast answer is fine for "make a quick bar chart."

Canvas. For long-form writing and code you edit side by side, fast answering keeps the loop tight. **Turn reasoning up only when the *content itself* requires hard logic**, like deriving an algorithm.

Custom GPTs and Projects. Your custom instructions and memory shape behavior either way, but a model thinking harder will follow multi-part instructions more reliably because it can plan around them. If your Custom GPT keeps ignoring constraint #4 of 6, the reasoning level may be the bottleneck.

The ChatGPT agent and scheduled tasks. Autonomous, multi-step work (browse, click, synthesize) leans on heavier reasoning by design, because each step depends on the last.

Choosing Between GPT and Reasoning Models

Watch on YouTube

Cost and latency: the real tradeoff

The pricing details change, so reason qualitatively (the live numbers are on the pricing page):

  • Reasoning costs more per task than the token rate suggests, because it generates invisible reasoning tokens on top of the visible answer. High effort on a hard problem can cost several times a fast answer on an easy one.
  • mini/nano variants exist precisely so you can run reasoning at scale without that cost exploding. For batch classification of 100k support tickets, a mini model is almost always the right call over the flagship.
  • Latency tracks the same axis. High reasoning effort feels slow because the model is thinking. That is the point, not a bug. Do not put it behind a UI that needs sub-second responses unless you cap effort at minimal or low.

A simple internal rule for teams: default to fast answering, allow high reasoning effort on an explicit allowlist of task types, and reserve the flagship at high effort for low-volume, high-stakes work.

Knowledge check

1. What is the fundamental distinction between general GPT models and reasoning models?

2. According to the lesson, what single question best determines whether to route a task to a reasoning model?

3. What are 'reasoning tokens' as described in the lesson?

MULTIPLE CHOICE

4. Select ALL tasks that are best suited to a general GPT model rather than a reasoning model.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL correct statements about the 'mini' and 'nano' variants.

Select all the correct answers.

Reasoning changes how you prompt

Once you turn reasoning up, your prompting style should change, or you waste its strength.

Stop hand-holding the steps. With fast answering you often add "think step by step." A model already reasoning internally does this on its own, and spelling out the steps can actually constrain it. Instead, state the goal, the constraints, and the success criteria, then get out of the way.

Weak prompt when reasoning is high:

First list the clauses. Then compare them. Then apply section 12. Then conclude.

Stronger:

Resolve all renewal-term conflicts across these contracts. Success = every conflict identified with the controlling contract named and the section-12 rule that decides it. Flag anything ambiguous rather than guessing.

Give it room and a clear finish line. A reasoning model rewards a crisp definition of "done" and explicit instructions about uncertainty ("flag, don't guess"). It also honors structured outputs well, so pair it with a JSON schema when the answer feeds another system.

python
schema = {
    "name": "conflict_report",
    "schema": {
        "type": "object",
        "properties": {
            "conflicts": {
                "type": "array",
                "items": {
                    "type": "object",
                    "properties": {
                        "clause": {"type": "string"},
                        "winning_contract": {"type": "string"},
                        "rule": {"type": "string"},
                    },
                    "required": ["clause", "winning_contract", "rule"],
                    "additionalProperties": False,
                },
            }
        },
        "required": ["conflicts"],
        "additionalProperties": False,
    },
    "strict": True,
}

resp = client.responses.create(
    model="gpt-5",
    reasoning={"effort": "high"},
    input=contracts_text + "\n\nReturn the conflict report.",
    text={"format": {"type": "json_schema", **schema}},
)

strict: True guarantees the model returns exactly that shape, which matters most when the output flows into downstream code.

Reasoning inside agents

When you build with the Agents SDK or wire up function calling, the tradeoff maps cleanly onto roles:

  • The planner (decides which tools to call and in what order) benefits from high reasoning effort. Tool selection is exactly the kind of multi-step decision it handles well.
  • The workers (a tool that summarizes a fetched page, formats a result, drafts a reply) can run fast or on a cheaper mini model. They do narrow, single-shot jobs.

This split keeps an agent both smart and affordable: think hard about *what to do*, act cheaply on *each step*. A common anti-pattern is running the flagship at high effort for every tool call, which makes agents both slow and expensive for no accuracy gain.

A quick mental checklist

Before every non-trivial task, run this:

  1. Chain of dependent steps? → raise reasoning effort. Single transformation? → keep it fast.
  2. High volume or needs speed? → drop to mini/nano.
  3. Output feeds code or another tool? → add structured outputs.
  4. Reasoning turned up? → state goal + constraints + "done," not the steps.
  5. Building an agent? → reason hard to plan, act fast to execute.

Key Takeaways

  • Route by structure, not by hype. A single transformation should answer fast; a chain of dependent steps where one mistake breaks the result should reason harder. On the GPT-5 family that is the reasoning.effort knob rather than a separate model.
  • mini/nano variants are your cost lever. For high-volume or latency-sensitive work, a smaller model at the right effort beats the flagship every time.
  • Reasoning costs more per task than the token rate implies because of invisible reasoning tokens. Default to fast, escalate deliberately, and cap effort when latency matters.
  • Prompt differently when reasoning is high: give goal, constraints, and a definition of "done," and tell it to flag uncertainty instead of guessing. Skip the manual "step by step."
  • In agents, split roles: high effort to plan, fast models for individual tool calls. Pair reasoning output with structured outputs whenever it flows into other code.

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Route reasoning models to plan and cheap fast models to execute steps
  • Prompt reasoning models with goal, constraints, and definition of done
  • Cap reasoning effort and default to fast models, escalating deliberately
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.