+140 XP

Tokens, training, and inference: what happens when you hit enter

Type "Paris is the capital of" into ChatGPT, Claude, or Gemini, and it will almost certainly finish with "France." We cover why the model does this in detail in 'What a Model Actually Does: Prediction, Not Understanding.'

Let's open the hood. Three ideas explain the whole pipeline: tokens, training, and inference.

Step 1: tokens (how the model reads)

A large language model (LLM) does not read words the way you do. It breaks text into tokens: small chunks that are often whole words, but sometimes pieces of words.

Here is the classic surprise. The word "strawberry" is not one unit to the model. It is usually read as something like "straw" + "berry." This is why models historically struggled to count the letters in "strawberry": they never saw the individual letters, just the chunks.

A rough rule of thumb in English: 1 token is about 4 characters, or roughly 3/4 of a word. So 100 tokens is about 75 words.

Some quick examples of how text splits:

  • "cat" → 1 token
  • "unbelievable" → "un" + "believ" + "able" (3 tokens)
  • "GPT-4" → "G" + "PT" + "-" + "4" (several tokens)
  • A space before a word is usually part of the token (" France" is different from "France")

You can see this for yourself with OpenAI's free Tokenizer tool. Paste in a sentence and watch it split into colored chunks.

Why tokens matter to you

Tokens are the unit of pricing and limits. When a tool says it has a "200,000 token context window," it means roughly 150,000 words can fit in a single conversation. When an API charges "$3 per million input tokens," it is counting these chunks, not words.

Practical takeaway: shorter prompts cost less and fit more. And odd behavior with spelling, rhyming, or counting letters often traces back to tokenization.

Step 2: training (how the model learned)

Before you ever typed anything, the model went through training: a long, expensive process of learning from enormous amounts of text (books, websites, code, articles).

The core task is shockingly simple. The model is shown a sequence of text with the last token hidden, and it must predict the next token.

Show it "The sky is" and it learns to predict "blue." Show it "2 + 2 =" and it learns to predict "4." Do this across trillions of tokens, billions of times, and the model slowly adjusts its internal settings (called parameters, the numbers it tunes) to get better at guessing.

Nobody hand-coded the rule "Paris pairs with France." The model saw that pattern millions of times in its training text and learned the statistical link on its own.

Two phases you should know

Training has two main stages:

  1. Pre-training: the model reads the giant pile of text and learns next-token prediction. This produces a "raw" model that is knowledgeable but not very helpful or polite.
  2. Fine-tuning and alignment: humans and automated feedback teach the model to follow instructions, be helpful, and avoid harmful answers. This is what turns a text-predictor into ChatGPT.

This is also why models have a knowledge cutoff. Training happened at a fixed point, so a model may not know about events after that date unless it can search the web. In 2026, ChatGPT, Claude, and Gemini all offer live web search to patch this gap, but the base knowledge is still frozen at training time.

Step 3: inference (what happens when you hit enter)

Inference is the moment of use: the model is done learning and is now making predictions for you in real time.

Here is the full sequence when you press Enter:

  1. Your prompt is split into tokens.
  2. The model reads those tokens and predicts the most likely next token.
  3. It adds that token to the sequence, then predicts the next one.
  4. It repeats, one token at a time, until it decides the answer is complete.

That word-by-word streaming you see in the interface? That is literally the model generating one token, then the next, then the next. You are watching inference happen live.

Why you get different answers to the same prompt

If the model always picked the single most likely token, it would be repetitive and robotic. So there is a setting called temperature that adds controlled randomness.

  • Low temperature (near 0): predictable, focused, repeatable. Good for factual tasks, code, data extraction.
  • High temperature (near 1 or above): varied, creative, surprising. Good for brainstorming, story ideas, marketing copy.

This is why asking "give me a tagline for my coffee shop" twice gives two different taglines. The randomness is a feature, not a bug.

Seeing inference in code

If you want to watch the pipeline concretely, here is a short, runnable example using the OpenAI API. (You need an API key; the same idea works with Anthropic's Claude and Google's Gemini APIs.)

python
from openai import OpenAI

client = OpenAI()  # uses your API key

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[
        {"role": "user", "content": "Paris is the capital of"}
    ],
    temperature=0,        # near-zero = most likely token, very predictable
    max_tokens=5          # limit the answer to a few tokens
)

print(response.choices[0].message.content)
# Expected output: France

Two things to notice. temperature=0 tells the model to play it safe and pick the most likely continuation. max_tokens=5 caps the response length in tokens, not words, which ties straight back to Step 1.

Knowledge check

1. Why did language models historically struggle to count the letters in a word like 'strawberry'?

2. A tool advertises a '200,000 token context window.' What does this practically tell you?

3. According to the lesson, what is the core task the model learns during training?

MULTIPLE CHOICE

4. Select ALL correct statements about tokens and why they matter.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL practical takeaways that follow correctly from how tokenization works.

Select all the correct answers.

Putting it together: the full pipeline

Let's trace one real prompt end to end. You type:

"Summarize this email in one sentence: [pastes a long email]"
  1. Tokens: your instruction plus the whole email gets chopped into tokens. A 500-word email is roughly 650 tokens.
  2. Training (already done): the model long ago learned what "summarize" means and how summaries are structured, from countless examples.
  3. Inference: the model reads all those tokens and generates a one-sentence summary, one token at a time, until it hits a natural stopping point.

No lookup. Just very, very good next-token prediction, shaped by training and powered by inference.

Why this mental model makes you better at AI

Once you see the model as a next-token predictor, a lot of practical advice clicks into place:

  • Give context up front. The model predicts based only on the tokens in front of it. If you don't paste the email, it can't summarize it. There is no hidden memory of your files.
  • Examples steer predictions. Showing the model two sample summaries in your preferred style nudges it toward predicting that same style. This is why "show, don't just tell" works so well in prompts.
  • Be specific to narrow the guesses. "Write a tweet" leaves the prediction wide open. "Write a 280-character tweet, friendly tone, one emoji, no hashtags" sharply constrains what comes next.
  • Hallucinations are over-confident predictions. When the model invents a fake source, it is predicting text that *looks* right based on patterns, even though it isn't true. That is why you verify anything that matters.

For a deeper but still readable explainer, the Wikipedia article on large language models is a solid, free reference that stays reasonably plain.

Key Takeaways

  • The model reads in tokens, not words. Roughly 4 characters per token. This drives pricing, context limits, and quirks like miscounting letters in "strawberry." Try the Tokenizer tool once to make it concrete.
  • Training is next-token prediction at massive scale. The model learned patterns like "Paris → France" by guessing the next chunk across trillions of tokens, not by memorizing a fact database.
  • Inference is what runs when you hit Enter: one token predicted at a time, streamed back to you. The live word-by-word output is the prediction happening in real time.
  • Use temperature deliberately. Low for factual, repeatable work; high for creative brainstorming.
  • Feed the model context and examples. It only predicts from the tokens you give it, so specific prompts with clear examples produce sharper, more useful answers.

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Write specific prompts with examples, roles, and format constraints
  • Set temperature low for factual work, high for creative brainstorming
  • Budget token usage against context limits using the tokenizer
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.