Anthropic CCA-F: Claude Extended Thinking

The approved title uses “Extended Thinking,” but current Claude model behavior has moved toward adaptive thinking. Anthropic’s current documentation says manual extended thinking with thinking.type: "enabled" and a fixed budget_tokens is deprecated on Claude 4.6 models, unsupported on Claude 4.7 and later, and replaced by adaptive thinking on newer model families. Claude Sonnet 5.5, for example, uses adaptive thinking with a per-request effort setting instead of a manually fixed thinking budget.

Within Claude Engineering, the practical lesson is to design around model-specific thinking capabilities rather than hard-code one API pattern. The workload should express how much reasoning it needs, measure quality/latency/cost, and let the selected model use the supported thinking mode.

Know which thinking mode the model supports

Older supported models such as Claude Haiku 4.5 still use manual extended thinking, while current higher-end models use adaptive thinking.

Claude Fable 5.1 uses adaptive thinking only; Claude Sonnet 5.5 uses adaptive thinking with configurable effort.

Model capability should therefore be read from the current model documentation or Models API before request construction.

Manual extended thinking is a compatibility mode

For models that still support it, manual extended thinking allocates a fixed thinking-token budget before the final answer.

This can be useful where predictable cost or latency matters.

But new systems should avoid making a fixed budget central to application logic because the current model line increasingly expects adaptive thinking.

Adaptive thinking should be controlled with effort

Current adaptive-thinking models let the model decide when and how much reasoning is useful while the application sets an effort level.

Use lower effort for routine classification or simple transformations and higher effort for planning, coding, difficult analysis, or long-horizon tool work.

Evaluate each task class rather than setting maximum effort globally.

Thinking blocks are part of the response contract

Current Claude responses can include thinking blocks in addition to text and tool-use blocks.

Application parsers should handle content by block type rather than assuming every response is one text string.

Streaming UIs should also account for models that return text between tool calls inside thinking blocks unless the display behavior is configured differently.

Preserved thinking affects multi-turn conversations

Current model documentation describes cases where thinking blocks can be preserved across turns or tied to a model/conversation.

Do not strip or rewrite structured response blocks casually if later turns or tool interactions depend on the model’s supported conversation format.

Claude Conversation State provides the broader state-management context.

Tool use and thinking must be tested together

Agent workloads often alternate reasoning and tool calls.

New model releases can change whether thinking appears before or between tools and how forced-tool behavior works.

Run regression tests on full tool trajectories, not only final answer quality, whenever moving to a newer Claude model.

Reasoning depth is a cost control

More thinking can improve difficult-task quality but can also increase latency and token usage.

Track thinking-enabled task cost alongside ordinary input/output token usage and downstream tool cost.

Claude Cost Controls should treat effort as a routing decision, not a free quality switch.

Latency-sensitive paths need explicit budgets

An interactive support assistant may have a p95 latency target that makes high effort unacceptable on every turn.

Route easy intents to lower effort or smaller models and reserve deeper reasoning for cases that genuinely benefit.

Escalation can be based on task type, confidence, policy risk, or failure of the simpler path.

Thinking should be evaluated for business outcome

Do not judge a reasoning mode by whether responses look more detailed.

Measure task success, correctness, tool decisions, safety, latency, user completion, and cost.

Claude Evaluation Sets should include the task categories where adaptive/high-effort thinking is expected to help.

Model migrations should remove obsolete thinking parameters

A request using manual thinking.type: enabled can return a 400 error on models that only support adaptive thinking.

Migration code should select request fields from the target model capability instead of carrying every old parameter forward.

Keep model configuration centralized so one application update can change thinking behavior consistently across services.

Extended thinking succeeds when reasoning is a measured workload setting

The mature Claude application treats manual extended thinking as a model-specific legacy/compatibility mode, prefers adaptive thinking on current models, configures effort by task, parses thinking/tool blocks correctly, and validates quality/cost/latency through regression tests.

The goal is not to maximize thinking tokens. It is to spend deeper reasoning only where it improves the outcome enough to justify the extra latency and cost.

Per-message effort is useful when the conversation contains both easy and hard turns. A support assistant might use ordinary effort for greetings and account lookups, then raise effort for a complex root-cause analysis or policy interpretation. Keep the routing rule outside Claude so the application can explain why a given request consumed more reasoning budget.

Thinking behavior should also be included in model fallback logic. If the primary model fails or is unavailable, the fallback may use a different thinking mode or effort surface. A fallback that silently drops from adaptive thinking to a legacy model with manual extended thinking needs an explicit request adapter and its own evaluation evidence.

Streaming clients should test how thinking changes UI behavior. Current newer-model behavior can place intermediate text into thinking blocks rather than ordinary text blocks between tool calls. If the product previously streamed every assistant fragment to users, a model migration can create apparent silence or expose content the UI was never designed to display. Treat content-block rendering as a versioned client contract.

Reasoning depth should be constrained for sensitive side effects. A more capable thinking mode can improve planning, but it should not grant new tool permissions or bypass approval. The authorization layer should remain deterministic. High effort may help decide which approved tool to use; it should never make an unapproved destructive tool suddenly available.

Prompt caching interacts with thinking configuration. Changing key request configuration or conversation structure can invalidate useful cached prefixes, and some current model migrations change preserved-thinking semantics. Track cache-hit rate before and after a reasoning-policy change so a quality improvement does not create an unexpected input-cost spike.

Regulated workflows should record the model ID, thinking mode or effort level, prompt/tool version, and evaluation evidence associated with important decisions. Auditors usually need the reproducible configuration that produced an outcome, not a vague statement that ‘Claude used more reasoning.’ This metadata belongs in release and decision logs rather than in the human-readable answer.

Failure analysis should separate ‘not enough reasoning’ from ‘bad context.’ Raising effort cannot fix missing documents, wrong tool results, stale conversation state, or a poor system instruction. When a hard case fails, first inspect evidence and task framing, then test whether deeper thinking improves the result. This avoids spending more tokens on a fundamentally mis-specified task.

Extended/adaptive thinking should also be stress-tested for adversarial prompts. A model given more reasoning time can still encounter prompt injection, misleading retrieved content, or tool-manipulation attempts. Include these cases in the evaluation set and verify that higher effort preserves security boundaries rather than merely producing a more articulate unsafe plan.

Operational metrics should break out effort tier by task category, success rate, latency, tokens, tool calls, and human escalation. If high effort is used on 80% of traffic but improves only a narrow slice, the routing policy is too broad. The goal is a small number of clearly justified expensive turns.

Thinking policy should be revisited whenever Anthropic changes the active model lineup. Current Sonnet 5.5 and newer families favor adaptive thinking, while older models still differ. Centralize capability discovery and regression tests so the application can migrate without searching every code path for legacy thinking parameters.

Effort routing should consider confidence of the cheap path. One practical pattern is to try a lower-effort model/configuration, run deterministic checks, and escalate only if validation fails or confidence is insufficient. This keeps difficult cases expensive without making every user pay the latency cost of the hardest reasoning tier.

Benchmarks should include tasks where more reasoning is actively harmful, such as concise deterministic extraction or short customer replies. Higher effort can increase verbosity or over-analysis without improving correctness. A balanced evaluation set prevents the team from assuming ‘more thinking’ is always a monotonic quality improvement.

Tool latency can dominate reasoning latency. An agent that spends two seconds thinking and twenty seconds waiting for external systems should be optimized at the tool layer first. Break p95 latency into model-thinking, generation, network, tool and orchestration components before reducing effort simply because the overall task feels slow.

Thinking configuration belongs in feature flags or model policy, not hard-coded inside prompt templates. This lets the platform team change effort by task class, customer tier or incident state without editing business prompts. It also makes rollback simple if a new effort setting produces unexpected cost or behavior.

Release notes should be reviewed for thinking-related breaking changes before model upgrades. Anthropic’s recent model generations have changed thinking support, forced tool-use behavior and response-block placement. A migration checklist that includes response parsing and thinking configuration is safer than treating the change as a model-ID replacement only.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!