Context Window
Also: Context Length, Token Window, Maximum Context Length
The context window is the maximum amount of text (measured in tokens) a language model can process at once, including both the input prompt and the generated output.
What It Is
The context window is the fixed limit on how much text a large language model (LLM) can consider at a single time. It is measured in tokens, not words or characters. A token is a chunk of text, roughly 3 to 4 characters in English, so 1,000 tokens equals about 750 words.
The context window covers everything in a single interaction:
- The system prompt (instructions defining the model's behavior)
- The user input (your question or task)
- Any retrieved documents or pasted content
- The conversation history in a chat
- The model's own response as it is generated
If the combined total exceeds the window, the oldest or least relevant content must be dropped or truncated.
Why it matters
The context window defines the practical limits of what a model can reason about. A model cannot use information it never received, and it cannot remember anything outside the current window unless that information is fed back in.
Key implications:
- Larger windows allow processing long documents, full codebases, or extended chats, but cost more and can slow responses.
- Smaller windows force you to summarize, chunk, or selectively retrieve content.
- Models may show degraded accuracy in the middle of very long inputs, an effect often called "lost in the middle."
How it is used in practice
Professionals manage the context window through several techniques:
- Chunking: splitting long documents into smaller pieces.
- Retrieval Augmented Generation (RAG): pulling only the most relevant passages into the window.
- Summarization: condensing earlier chat turns to free up space.
- Prompt budgeting: reserving enough tokens for the output.
Concrete Example
Suppose a model has an 8,000 token context window. You paste a 6,000 token contract and ask a question using 200 tokens. That leaves roughly 1,800 tokens for the system prompt and the answer. If you paste a 9,000 token contract instead, it will not fit, so you must summarize it or retrieve only the relevant clauses. Planning around this budget is essential for reliable, cost effective AI applications.
See also
Frequently asked questions
What exactly is a context window in an LLM?
The context window is the fixed maximum amount of text a large language model can consider in one interaction, measured in tokens rather than words. It covers the system prompt, your input, any pasted or retrieved documents, the conversation history and the response being generated. Anything beyond that limit must be dropped, truncated or summarized.
How do tokens convert into words?
In English, a token is roughly 3 to 4 characters, so 1,000 tokens correspond to about 750 words. Token counts, not word counts, are what a model measures against its context limit and what most providers bill on. Estimating your prompt in tokens is the only reliable way to know whether a document will fit.
Is the model's answer counted inside the context window?
Yes. The generated response consumes tokens from the same window as the input, which is why prompt budgeting matters. If you fill nearly the entire window with a pasted document, the model has little room left to answer and the output gets cut short. Reserve tokens for the output before deciding how much material to paste.
Does a larger context window always give better results?
No. Bigger windows let you process long documents, full codebases or extended chats, but they cost more and can slow responses. Models also tend to lose accuracy on information sitting in the middle of very long inputs, an effect known as "lost in the middle." Feeding fewer, better-targeted passages often beats dumping everything in.
What do I do when a document is too long for the context window?
Four techniques cover most cases: chunking the document into smaller pieces, using Retrieval Augmented Generation (RAG) to pull only the relevant passages into the window, summarizing earlier conversation turns to free up space, and prompt budgeting to reserve tokens for the output. Concretely, with an 8,000 token window, a 9,000 token contract will not fit, so you summarize it or retrieve only the clauses that matter.