The gist
Before a tough move, a chess player thinks for 5 minutes: runs through options, calculates consequences, throws out bad lines. Extended Thinking (Claude's deep-analysis mode) is the same thing for Claude: it "thinks out loud" before answering instead of going with the first thing that comes to mind. For simple tasks it's overkill. For architecture decisions and complex analysis, it makes a fundamental difference in quality.
Terms in this lesson: extended thinking (Claude's deep-analysis mode), API (application programming interface), token (a unit of text for AI), prompt (a request to the AI), prompt caching (saving a prompt for reuse so you don't pay for it again in full).
Key concepts
- Extended Thinking: an API mode in which Claude generates an internal monologue before the final answer
- Thinking tokens: the tokens of that internal reasoning, returned in a separate
thinkingblock - budget_tokens: a parameter capping the maximum tokens spent on thinking (manual mode, only for older models: deprecated on 4.6, returns a 400 error on 4.7 and newer)
- Adaptive Thinking: the automatic mode (
"type": "adaptive"), where the model decides whether to think and how deeply. The depth is set by theeffortparameter inoutput_config. The main approach for all current models (as of October 2026: Opus 5.5, Sonnet 5.5, Fable 5.1) - display: a display parameter, either
"summarized"(a summary of the reasoning) or"omitted"(only a signature, no text). On Opus 5.5, Sonnet 5.5 and Fable 5.1 the default is"omitted" - Interleaved Thinking: thinking between each tool call, not only at the start
- Thinking block: a separate block in the API response with the fields
type: "thinking",thinking: "..."andsignature: "..." - Cost: thinking tokens are billed as output tokens (pricier than input). You're billed for the full thinking tokens, even if display =
"summarized"
Theory
How it works under the hood
Without Extended Thinking, Claude gets a request and immediately generates a response. It's fast, but the thinking is "flat": the model doesn't get a chance to check alternatives.
With Extended Thinking, the request goes through two stages:
Request → [Thinking phase: Claude works through options] → Final answerThe thinking phase is invisible by default; you only see the final answer. Through the API you can get a summary of the internal monologue (the raw train of thought isn't returned under any settings).
What's new as of October 2026: on current models (Opus 5.5, Sonnet 5.5, Fable 5.1), thinking is already on by default, and the standard way to turn it off (thinking: {"type": "disabled"}) returns a 400 error. So the developer's job has shifted: not "turn thinking on," but "choose the depth" (effort) and decide whether to show the reasoning text. The old manual mode with budget_tokens is only needed for legacy models.
Two modes: Manual vs. Adaptive
The API has two ways to control thinking:
1. Manual (a manual budget): you set the token limit yourself. Works only on legacy models:
thinking={"type": "enabled", "budget_tokens": 10000}2. Adaptive (automatic): the model decides how much to think, and you set the depth with the effort parameter, separately from thinking:
thinking={"type": "adaptive"},
output_config={"effort": "medium"} # low / medium / high and above, depending on the modelWhich models support what
| Model (as of October 2026) | Manual ("enabled") |
Adaptive ("adaptive") |
Note |
|---|---|---|---|
| Claude Fable 5.1, Opus 5.5, Sonnet 5.5 | ❌ returns a 400 error | ✅ adaptive only | Thinking is on by default ("disabled" returns 400); display defaults to "omitted" |
| Claude Opus 4.8 and Opus 4.7 | ❌ returns a 400 error | ✅ adaptive only | Without the thinking field, thinking is off; turn it on explicitly |
| Claude Opus 4.6, Sonnet 4.6 | ⚠️ deprecated | ✅ recommended | Manual still works |
| Claude Opus 4.5, Sonnet 4.5, Haiku 4.5 | ✅ the only mode | ❌ returns a 400 error | No adaptive mode |
Older models are gradually being retired from the API; current list: What's current.
The trend: Anthropic has moved to Adaptive. For new projects, use "adaptive" and effort.
An API request with Extended Thinking (Manual, for legacy models)
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-haiku-4-5", # manual mode: legacy models only; on Haiku 4.5 it's the only mode
max_tokens=16000,
thinking={
"type": "enabled",
"budget_tokens": 10000 # up to 10000 tokens for thinking
},
messages=[{
"role": "user",
"content": """Design a system architecture for the following case:
- A SaaS platform for small businesses
- 1000 active users
- Multi-tenancy is required
- Budget: $200/month for infrastructure
- Team: 1 developer
Evaluate at least 3 approaches with real trade-offs."""
}]
)
# Go through the response blocks
for block in response.content:
if block.type == "thinking":
print("=== INTERNAL REASONING ===")
print(block.thinking)
print()
elif block.type == "text":
print("=== FINAL ANSWER ===")
print(block.text)An API request with Adaptive Thinking (recommended for new models)
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
thinking={
"type": "adaptive",
"display": "summarized" # to see a summary of the reasoning
},
output_config={"effort": "high"}, # depth of work: low, medium, high and above (depends on the model)
messages=[{
"role": "user",
"content": "Design the architecture of a multi-tenant SaaS platform..."
}]
)With Adaptive, you don't have to guess at budget_tokens: the model decides how deeply to think, and you only adjust effort. Keep in mind that at low effort, the model may skip thinking entirely on a simple request. If the SDK complains about output_config, update the library: pip install -U anthropic.
What you see in a thinking block
An example of a real internal monologue (shortened):
Hmm, I need to design an architecture for a SaaS on a tight budget... Option 1: Shared database schema - Pros: simple, cheap, one database - Cons: hard to isolate customer data, risky as it grows - OK for 1000 users, but what if it grows to 10000? Option 2: Database per tenant - Pros: full isolation, easy to roll back one customer - Cons: $200/month won't cover it if 1000 customers = 1000 databases - Ruling this out for this budget Option 3: Schema per tenant (Postgres schemas) - A compromise: isolation without an explosion in databases - Row Level Security adds another layer - Cloudflare Workers + PlanetScale serverless = fits in $200/month I'll recommend Option 3 as the main one, and explain when to move to Option 2...
This isn't a staged showcase. It's the real process of working through options.
The budget_tokens parameter (Manual mode)
| budget_tokens | When to use | Cost (example calculation) |
|---|---|---|
| 1,024 (minimum) | Moderately complex tasks | ~$0.005 per request |
| 5,000 | Architecture decisions, analysis | ~$0.025 per request |
| 10,000 | Maximum depth, math | ~$0.05 per request |
| 32,000 | Extremely complex tasks | ~$0.16 per request |
The math: thinking tokens × the output token price. The example uses $5 per 1M output tokens (what Haiku 4.5 costs as of October 2026: manual mode still works on it, and it may be retired from the API no earlier than October 15, 2026); thinking tokens are counted as output. Prices differ by model: current prices and versions: What's current. Check your real usage in the usage.output_tokens_details.thinking_tokens field of the response.
Limits: budget_tokens must be at least 1024 and less than max_tokens (except with interleaved thinking, where it can be more). The budget is a guideline, not a hard ceiling: the model may stop earlier; max_tokens sets the hard ceiling. For budgets above 32,000 tokens, the documentation recommends the Batch API: requests like that run long and hit timeouts.
The display parameter: what the user sees
It controls what comes back in the response's thinking block:
| display | What comes back | What it's for |
|---|---|---|
"summarized" |
A summary of the reasoning (the default on Opus 4.6, Sonnet 4.6 and older) | Debugging prompts, understanding the model's logic |
"omitted" |
An empty thinking field, only signature (the default on Opus 5.5, Sonnet 5.5, Fable 5.1) |
Production (faster time-to-first-token) |
The field name is the same in both modes. The text in the block is always a summary, not the raw train of thought.
# Production mode: faster, no extra data
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
thinking={
"type": "adaptive",
"display": "omitted" # ← only the signature, no reasoning text
},
messages=[{"role": "user", "content": "..."}]
)Important: with "omitted" you still pay for all the thinking tokens. The savings aren't in money but in how fast the answer arrives. When you pass thinking blocks back in a multi-turn conversation (with tools this is required), pass them unchanged: the server decrypts the full reasoning context from the signature field.
Streaming thinking tokens
For long reasoning, streaming makes sense, because you can see the progress:
with client.messages.stream(
model="claude-opus-5-5",
max_tokens=16000,
thinking={
"type": "adaptive",
"display": "summarized"
},
messages=[{"role": "user", "content": "...your complex request..."}]
) as stream:
for event in stream:
if event.type == 'content_block_start':
if event.content_block.type == 'thinking':
print("[Starting to think...]")
elif event.type == 'content_block_delta':
if event.delta.type == 'thinking_delta':
print(event.delta.thinking, end='', flush=True)
elif event.delta.type == 'text_delta':
print(event.delta.text, end='', flush=True)
elif event.type == 'content_block_stop':
print("\n[Block finished]")When streaming, three types of delta events arrive (with display: "omitted", thinking_delta arrives with an empty string and there's no reasoning text):
thinking_delta: the reasoning text (in chunks)signature_delta: the cryptographic signature (for multi-turn)text_delta: the final answer
When you need Extended Thinking
Use it for:
- Architecture decisions (choosing technologies, system structure)
- Multi-step math problems
- Analyzing complex trade-offs with several variables
- Debugging tricky bugs where the cause isn't obvious
- Writing critical algorithms
Don't use it for:
- Writing simple text or a short summary
- Routine API calls and straightforward code
- Tasks where the first answer is already correct
- When you need speed and cost matters
Interleaved Thinking: thinking between tool calls
When Claude uses tools, it can think after each tool call, not only at the very start. That's called Interleaved Thinking.
Request → [Thinking] → tool_use: get_weather("Chicago")
→ tool_result: "23°F"
→ [Thinking: "OK, it's cold in Chicago, need to factor that in..."] ← thinks BETWEEN calls
→ tool_use: get_weather("Miami")
→ tool_result: "78°F"
→ [Thinking: "Miami is warmer, let me compare..."]
→ Final answerModel support (as of October 2026):
- Opus 5.5, Sonnet 5.5, Fable 5.1, plus Opus 4.8 and 4.7: interleaved automatically in adaptive mode, no header needed
- Opus 4.6: only in adaptive mode (not available in manual)
- Sonnet 4.6: automatic in adaptive mode; in manual it still works through the beta header
interleaved-thinking-2025-05-14, but that's deprecated - Opus 4.5, Sonnet 4.5 and other Claude 4 models: through the beta header
interleaved-thinking-2025-05-14 - Haiku 4.5: not supported
Important when working with tools: when sending a tool_result back, always pass all the thinking blocks from the previous assistant response. Don't modify or remove them.
Example: with thinking vs. without
Without Extended Thinking, the request: "Pick a database for my SaaS"
The answer comes back in 2-3 seconds. It will most likely recommend PostgreSQL or MongoDB with boilerplate arguments.
With Extended Thinking (adaptive, effort high; in the old manual mode, budget_tokens: 8000):
Claude will spend noticeably more time on analysis (on the order of tens of seconds, depending on the model and load). In the thinking block you'll see it consider your specific parameters, compare the cost of different cloud databases at a load of 1000 users, take into account that you have one developer, and weigh PlanetScale vs. Supabase vs. Neon on real criteria.
The answer is noticeably better, not because the model is "smarter," but because it had time to think.
Limitations of Extended Thinking
Not every API feature is compatible with Extended Thinking:
| Feature | Compatibility | Note |
|---|---|---|
tool_choice: "auto" |
✅ | Works |
tool_choice: "none" |
✅ | Works |
tool_choice: "any" |
❌ in manual; works in adaptive, except on Opus 5.5, Sonnet 5.5 and Fable 5.1 | On those three models, forcing a tool call always returns a 400 error |
tool_choice: {"type": "tool", "name": "..."} |
❌ in manual; works in adaptive, except on Opus 5.5, Sonnet 5.5 and Fable 5.1 | Same as above |
max_tokens: 0 (cache warming) |
❌ | Incompatible |
| Prompt Caching (system prompt) | ⚠️ | When you change the thinking mode, budget_tokens or effort, the system prompt and tools cache can miss too: treat it as starting over |
| Prompt Caching (messages) | ⚠️ | Invalidated by any change to the thinking mode, budget_tokens or effort |
Multi-turn: pass the thinking blocks along
In multi-turn conversations, pass along all thinking blocks from previous assistant responses unchanged: inside a tool-use loop this is required, and in a regular conversation it's recommended. Otherwise the model can lose the context of its own reasoning:
# Turn 1
response1 = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
thinking={"type": "adaptive"},
messages=[{"role": "user", "content": "First question?"}],
)
# Turn 2: pass response1.content in full (with the thinking blocks), unchanged
response2 = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
thinking={"type": "adaptive"},
messages=[
{"role": "user", "content": "First question?"},
{"role": "assistant", "content": response1.content}, # ← ALL the blocks
{"role": "user", "content": "Follow-up question?"},
],
)Practice
Assignment: an architecture decision with Extended Thinking
Pick a real architecture problem from your own project (or use a practice one: "How should I store user data for a SaaS with 500 customers?")
First, request an answer without Extended Thinking and save it:
python response_basic = client.messages.create( model="claude-haiku-4-5", # for comparison: a model without thinking by default max_tokens=2000, messages=[{"role": "user", "content": YOUR_REQUEST}] )Then the same request with Extended Thinking (Manual: on Haiku 4.5 it's the only thinking mode):
python response_thinking = client.messages.create( model="claude-haiku-4-5", max_tokens=8000, thinking={"type": "enabled", "budget_tokens": 6000}, messages=[{"role": "user", "content": YOUR_REQUEST}] )And the same request with Adaptive Thinking (a current model, for example Opus 5.5):
python response_adaptive = client.messages.create( model="claude-opus-5-5", max_tokens=8000, thinking={"type": "adaptive", "display": "summarized"}, output_config={"effort": "high"}, messages=[{"role": "user", "content": YOUR_REQUEST}] )Print all the final answers and the thinking blocks
Compare: where is the analysis deeper? What did the "fast" answer miss?
Estimate the cost of each option with
response.usage
Goal: get a feel for the difference in quality and understand which tasks justify paying extra.
Tools and resources
- Anthropic API docs: Extended Thinking and Thinking
- Python SDK:
pip install -U anthropic(get a recent version: theoutput_configanddisplayparameters were added recently) - In Claude Code: the effort level is set with the
/effortcommand (or the--effortflag),Ctrl+Oshows the thinking (verbose mode), and the wordultrathinkin a request asks for deeper thinking for one turn; in Claude Code on Opus 5.5, Sonnet 5.5 and Fable 5.1, thinking can't be switched off. More: model documentation - Models: Opus 5.5, Sonnet 5.5, Fable 5.1: adaptive only; Opus 4.6 and Sonnet 4.6: adaptive (manual is deprecated); Opus 4.5, Sonnet 4.5, Haiku 4.5: manual only
- Prices: What's current, claude.com/pricing
Key takeaways
Extended Thinking isn't magic, it's time to work through options. Quality goes up, and so does cost. Thinking tokens are billed as output: to avoid overpaying, lower
effort(in manual mode,budget_tokens), and set a hard ceiling withmax_tokens. For all current models, use Adaptive Thinking ("type": "adaptive") andeffortinoutput_config: the model decides how much to think. A manualbudget_tokenson new models returns a 400 error. Thedisplay: "omitted"parameter speeds up responses in production but doesn't save money: billing is based on the full thinking tokens. Use it for architecture decisions and complex analysis. For simple tasks, it's unnecessary overhead.
What's next
The mark stays in this browser only and is never sent anywhere. My progress