Extended thinking

Extended thinking is the reasoning a model generates as tokens before it answers, billed as output tokens whether or not that reasoning is shown to you.

The model works through the problem first, restating it, trying approaches and checking intermediate results, then writes its answer. That scratch work is where the gain on hard tasks comes from, and it is also a cost: every thinking token is generated output, billed at the output rate and counted against max_tokens. Claude Code's documentation warns, as of September 2026, that the default can reach tens of thousands of tokens per request, depending on the model.

What you see is not what you pay for. On recent Claude models the thinking comes back as a summary or not at all, but billing covers the full reasoning the model generated, so the visible text understates the charge. The usage report gives the real figure in output_tokens_details.thinking_tokens.

Thinking also returns as input. On what Anthropic calls keep-all models, which as of September 2026 include Claude Opus 4.5 and later Opus models, Sonnet 4.6 and later and the Fable models, earlier turns' thinking blocks stay in the conversation and are billed as input on every later request, like the rest of the history.

Depth is now set by effort rather than by a fixed budget. In Anthropic's API the fixed-budget mode, budget_tokens, is deprecated on the 4.6 models and rejected from 4.7 onward in favour of adaptive thinking, and on Claude Fable 5, Fable 5.1 and Opus 5.5 thinking cannot be turned off at all: the effort level, from low to max, decides how deep it goes, with medium the default on Opus 5.5. In Claude Code it is set with /effort or in /model, and changing it mid-session invalidates the prompt cache on most models.

← All terms