Claude Opus vs Sonnet vs Haiku: what a coding task really costs
The price per million tokens is only half of what a coding task costs. The other half is how many tokens the model spends to finish it, and no price list shows that.
A coding agent's bill for a task is two numbers multiplied together: the price of a token on the model you picked, and the number of tokens that model spends before it is done. The price list gives you the first. The second depends on how the model works.
A model that costs twice as much per token and finishes in half the tokens costs the same. A model that is cheaper per token but reads three extra files, or answers wrongly and has to be asked again, can cost more.
The list prices, as of September 2026
Anthropic prices each model per million tokens, with separate rates for fresh input, cache writes, cache reads and output. These are its API list prices as of September 2026.
| Model | Input | Output | Cache read | Cache write 5 min | Cache write 1 h | Context |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 1.00 | 5.00 | 0.10 | 1.25 | 2.00 | 200,000 |
| Claude Sonnet 5 | 2.00 | 10.00 | 0.20 | 2.50 | 4.00 | 1,000,000 |
| Claude Sonnet 4.6 | 3.00 | 15.00 | 0.30 | 3.75 | 6.00 | 1,000,000 |
| Claude Opus 4.6 | 5.00 | 25.00 | 0.50 | 6.25 | 10.00 | 1,000,000 |
| Claude Opus 4.7 | 5.00 | 25.00 | 0.50 | 6.25 | 10.00 | 1,000,000 |
| Claude Opus 4.8 | 5.00 | 25.00 | 0.50 | 6.25 | 10.00 | 1,000,000 |
| Claude Opus 5 | 5.00 | 25.00 | 0.50 | 6.25 | 10.00 | 1,000,000 |
| Claude Opus 5.5 | 4.00 | 20.00 | 0.20 | 5.00 | 8.00 | 1,000,000 |
| Claude Fable 5 | 10.00 | 50.00 | 1.00 | 12.50 | 20.00 | 1,000,000 |
| Claude Fable 5.1 | 10.00 | 50.00 | 0.25 | 12.50 | 20.00 | 1,000,000 |
Three things in that table matter more for a coding agent than the headline input price:
- Cache reads. An agent re-sends the whole conversation on every turn and most of it comes back from the prompt cache, so this is the rate most of your input tokens pay. Most models read cache at a tenth of their input price and Fable 5.1 at a fortieth; Opus 5.5 reads it at a twentieth, $0.20 per million tokens, exactly the rate of Sonnet 5.
- Output, thinking included. Output costs five times input on every model, and thinking tokens are billed as output.
- Sonnet 5's price. Its $2 and $10 rates were announced as introductory; Anthropic has made them standard, and the rise to $3 and $15 planned for 1 September 2026 did not happen.
The other half: tokens per finished task
Tokens per task vary by model for reasons unrelated to price. A model that explores more reads more files before it answers, and every file it reads stays in the context, paid for again on each later turn. A wrong answer spends the tokens, then spends them again on the retry. A model that thinks longer spends more at the output rate.
Anthropic's own cost guide makes the same point from the other side: a more capable model often finishes with less work, meaning fewer turns, less searching and less backtracking, and that can outweigh a higher price per token. On a subset of SWE-bench Pro it reports Fable 5.1 at low effort solving 88.6% of tasks at $0.54 per solved task, against 77.4% at $0.84 for Sonnet 5 at its default, at five times the per-token price. It also shows workloads where the ranking flips, and concludes that no price list tells you which way it will go.
Measured cost per question, model by model
The table below measures both halves on a real repository: the same one-line questions, asked through Claude Code on each model and repeated three times, with the cost Claude Code itself reports at API list prices. Token figures count a cache read as a tenth of a fresh token, the rate most models bill it at. The table also shows the same questions sent through capsul on the same model; the benchmark page explains the method.
| Model | Input, bare | Input, capsul | Cost, bare | Cost, capsul | Less input |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | 55,500 | 19,300 | $0.0812 | $0.0265 | x2.88 |
| Claude Sonnet 5 | 54,200 | 18,900 | $0.153 | $0.0487 | x2.86 |
| Claude Sonnet 4.6 | 31,400 | 16,600 | $0.139 | $0.0613 | x1.89 |
| Claude Opus 4.6 | 45,700 | 16,400 | $0.332 | $0.101 | x2.78 |
| Claude Opus 4.7 | 38,700 | 20,100 | $0.28 | $0.126 | x1.92 |
| Claude Opus 4.8 | 65,800 | 13,800 | $0.498 | $0.0825 | x4.75 |
| Claude Opus 5 | 45,300 | 16,200 | $0.343 | $0.101 | x2.80 |
| Claude Opus 5.5 | 42,300 | 15,900 | $0.222 | $0.0696 | x2.66 |
| Claude Fable 5 | 41,400 | 17,600 | $0.633 | $0.219 | x2.35 |
| Claude Fable 5.1 | 49,700 | 23,400 | $0.636 | $0.279 | x2.12 |
How to read it:
- Start with Claude Code on its own: what each model spent answering, at its own list prices.
- Compare rows that share a price list. Opus 4.6, 4.7, 4.8 and Opus 5 are all priced at $5 and $25, so any gap between them is tokens, not price, and at the time of writing the gap is wide. Anthropic reports the same pattern on its own coding benchmark.
- Compare dollars across generations, not tokens. Models from 4.7 on count the same text as more tokens (see the next section), so token figures flatter Haiku 4.5, Sonnet 4.6 and Opus 4.6.
- Read cost next to the answer checks. A row that is cheap per question but finds fewer of the expected answers is not cheap per correct answer.
- Do not assume the lower price wins. Sonnet 5 lists below Sonnet 4.6, yet at the time of writing it is not the cheaper of the two on Claude Code alone, because it spent more tokens on these questions. On Anthropic's coding benchmark the price cut wins instead, although Sonnet 5 also uses more tokens per task there. Same two models, different work, different winner.
The tokenizer changed with Claude 4.7
Anthropic's pricing page states that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, the exact increase depending on content. Of the models measured above, that means Opus 4.7, 4.8, 5 and 5.5, Sonnet 5 and both Fables; Opus 4.6, Sonnet 4.6 and Haiku 4.5 use the previous one.
So a per-token comparison across that line is not like for like. The same prompt is already more tokens on the newer models before either one does any work, which is why Anthropic recommends comparing cost per solved task instead.
Effort is the second dial
Within one model, the effort level decides how many tokens it spends: how much it thinks, how many tool calls it makes and how much it writes. The levels run from low through medium, high and xhigh to max, and lower effort means fewer, terser tool calls and less thinking, which is billed at the output rate.
As of September 2026, Claude Code runs every model that supports effort at high by default, except Opus 5.5 at medium and Opus 4.7 at xhigh, so a default-to-default comparison of Sonnet 5 and Opus 5.5 is not like for like. Haiku 4.5 has no effort setting. Opus 5.5 and the Fable models always think and cannot have thinking turned off, so effort is the main control over how much they reason.
Anthropic describes Sonnet 5 at medium as comparable to Sonnet 4.6 at high. It also reports that long coding work is where effort genuinely buys accuracy, while on tasks well within a model's ability the highest levels pay for depth the task never uses.
/effort mediumSets the effort level in Claude Code; /effort on its own opens a slider.
Which model for which job
- Haiku 4.5 for mechanical edits: a rename, a formatting pass, a single-file change you will check, a quick lookup. It is the cheapest per token, half the price of Sonnet 5, and Anthropic describes it as fitting high-volume work with checkable outputs, not long agentic loops. Its window is 200,000 tokens rather than a million, and Anthropic commits to keeping it available until at least 15 October 2026, so check the deprecations page before building on it.
- Sonnet 5 for everyday work. Anthropic's Claude Code documentation says Sonnet handles most coding tasks well and costs less than Opus; per token, Sonnet 5 is half the price of Opus 5.5 on everything except cache reads, where they are equal.
- An Opus for hard problems: large cross-cutting refactors, difficult debugging, architectural decisions. Opus 5.5 is also Claude Code's default on Pro, Max, Team, Enterprise and the API as of September 2026, so if you never picked a model, it is the one you are paying for.
- Fable for the hardest and longest-running tasks. It is the most expensive per token, though Anthropic's cost guide suggests starting agent work on Fable 5.1 at
loweffort and raising it where it misses. Depending on your plan and seat tier, Claude Code may bill Fable to usage credits instead of your plan's included limits; the/modelpicker says so on the Fable row.
Switching models in Claude Code
In a session, /model followed by an alias switches at once, and /model alone opens a picker where the left and right arrow keys also set effort. claude --model sets the model at launch, and the model field in your settings file or the ANTHROPIC_MODEL environment variable makes a choice stick. The main aliases are haiku, sonnet, opus and fable, plus opusplan: Opus in plan mode, Sonnet for execution.
/model sonnetSwitches the current session to the latest Sonnet.
claude --model haikuStarts a new session on Haiku.
Switch at the start of a task, not halfway through. Each model has its own prompt cache, so after a switch the next request re-reads the whole conversation with no cache hits, which is expensive in a long session. Claude Code asks you to confirm a switch while the cache is still warm. Every plan-mode toggle under opusplan is such a switch, and on most models a mid-session effort change costs the same re-read. For subagents doing simple work, Anthropic suggests model: haiku in their configuration.
If what you want to shrink is what each question carries rather than the price of a token, capsul takes the same model names and sends the model short excerpts, what the question needs rather than the repository. Its column in the table above is that, measured.
capsul ask 'rename userId to accountId in the auth module' --model haikuDrives the Claude CLI you are already signed into, on the model you name.
Questions
Is Opus worth it for Claude Code?
For hard problems, often yes: a more capable model can finish in fewer turns with less searching and backtracking, which can outweigh a higher price per token. For routine edits it mostly pays for depth the task does not use, and Anthropic's own documentation says unexpectedly high spend on the API usually traces back to sessions that were never cleared or to Opus left as the default model. As of September 2026, Opus 5.5 is Claude Code's default on Pro, Max, Team, Enterprise and the API, so check /model before assuming you are on something cheaper.
What is the cheapest Claude model for coding?
Per token, Claude Haiku 4.5: as of September 2026 it is the cheapest model on Anthropic's API, at $1 per million input tokens and $5 per million output tokens. Per task, it depends on how many tokens a model spends and whether it gets the answer right, because a wrong answer is paid for twice. Anthropic describes Haiku as suited to high-volume work with outputs you can check, not long agentic loops, so it is the cheapest choice for mechanical edits rather than for everything.
Is Claude Sonnet 5 cheaper than Opus 5.5?
Per token, Sonnet 5 costs half as much as Opus 5.5 on fresh input, cache writes and output as of September 2026, and exactly the same on cache reads, at $0.20 per million tokens. Anthropic notes that cached input is the largest part of an agent's bill, so the real per-token gap is smaller than the headline prices suggest. Which one costs less per task then depends on how many tokens each spends: Sonnet 5 usually wins on everyday coding, while on hard problems the model that finishes in fewer turns can close the gap.
Why does the same prompt use more tokens on newer Claude models?
Claude 4.7 and later models use a newer tokenizer that, according to Anthropic, produces approximately 30% more tokens for the same text, with the exact increase depending on the content. Opus 4.6, Sonnet 4.6 and Haiku 4.5 use the previous tokenizer. Compare models on cost per task rather than on token counts, because token counts make the newer models look more expensive before either model has done any work.
Does the effort level change what a task costs?
Yes. Effort sets how much a model thinks, how many tool calls it makes and how much it writes, and thinking is billed at the output rate, so lower effort spends fewer tokens on the same task. In Claude Code you set it with /effort or with the arrow keys in /model; as of September 2026 most models default to high, Opus 5.5 to medium, and Haiku 4.5 has no effort setting. Anthropic's help centre also lists effort, alongside the model you use, among the factors that affect a subscription's usage.
$ npm i -g @penra/capsul