Chain-of-Thought
Also: CoT, Chain of Thought Prompting, Step-by-Step Reasoning
A prompting technique where a language model is guided to produce intermediate reasoning steps before giving a final answer, improving accuracy on complex tasks.
What It Is
Chain-of-Thought (CoT) is a prompting and reasoning technique used with large language models (LLMs) in which the model generates a sequence of intermediate reasoning steps before producing a final answer. Instead of jumping straight to a conclusion, the model is encouraged to "think out loud," breaking a problem into smaller, logically connected parts.
The technique can be triggered in several ways: by adding instructions like "Let's think step by step," by providing worked examples that include reasoning (few-shot CoT), or through model training that makes step-by-step reasoning a default behavior.
Why it matters
Many tasks, such as multi-step arithmetic, logic puzzles, and complex decision making, are difficult to solve in a single pass. Chain-of-Thought improves performance because it:
- Decomposes complexity: Large problems are split into manageable steps.
- Reduces errors: Each step builds on a verified intermediate result.
- Increases transparency: The visible reasoning trail makes it easier to audit and debug model behavior.
- Improves reliability: Studies consistently show higher accuracy on reasoning benchmarks compared to direct-answer prompting.
How it is used in practice
- Few-shot CoT: Include 2 to 5 examples in the prompt that demonstrate reasoning, then ask the new question.
- Zero-shot CoT: Simply append a cue such as "Explain your reasoning step by step."
- Self-consistency: Generate multiple reasoning paths and take the majority answer to boost robustness.
- Hidden reasoning: In production, the reasoning can be generated internally and stripped from the user-facing output to keep responses clean.
A practical caution: exposed reasoning can leak sensitive logic or be verbose, so teams often separate the reasoning from the final response shown to users.
Concrete Example
Prompt: "A store sells pencils at 3 for $2. How much do 12 pencils cost? Think step by step."
Model output:
1. 12 pencils divided into groups of 3 gives 4 groups.
2. Each group costs $2.
3. 4 groups times $2 equals $8.
Final answer: $8.
Without the steps, a model might guess incorrectly. With Chain-of-Thought, the structured path leads to a correct, verifiable result.
Frequently asked questions
What is Chain-of-Thought prompting?
Chain-of-Thought (CoT) is a technique where a large language model produces a sequence of intermediate reasoning steps before giving its final answer, instead of jumping straight to a conclusion. It is triggered by instructions such as "Let's think step by step", by worked examples that show the reasoning, or by training that makes step-by-step reasoning the model's default behaviour. It is also referred to as CoT or step-by-step reasoning.
When does Chain-of-Thought actually improve results?
On tasks that cannot be solved in a single pass: multi-step arithmetic, logic puzzles, and complex decision making. Chain-of-Thought helps because each step builds on a verified intermediate result, which cuts errors, and because the visible reasoning trail makes the model easier to audit and debug. Reasoning benchmarks consistently show higher accuracy than direct-answer prompting.
What is the difference between zero-shot and few-shot Chain-of-Thought?
Zero-shot CoT adds a simple cue to the prompt, such as "Explain your reasoning step by step", with no examples. Few-shot CoT includes 2 to 5 examples in the prompt that demonstrate the reasoning explicitly, then asks the new question. Few-shot costs more tokens but gives you control over the shape of the reasoning you want.
Should the reasoning steps be shown to end users?
Usually not. Exposed reasoning can leak sensitive business logic and tends to be verbose, so teams generate it internally and strip it from the user-facing output, an approach known as hidden reasoning. The reasoning stays available for internal audit and debugging while the user sees only a clean final answer.
What is self-consistency and how does it relate to Chain-of-Thought?
Self-consistency generates several independent reasoning paths for the same question and keeps the majority answer. It builds on Chain-of-Thought: a single reasoning chain can go wrong at one step, whereas agreement across multiple chains makes the final answer more robust. The trade-off is cost, since you run the model several times per question.