Prompt caching for coding agents: what it saves and what breaks it
A coding agent re-sends its whole context on every request, and the prompt cache is what stops that from being billed at full price each time. It is also fragile: change one byte near the top and the next turn pays for everything again.
Every step of an agent session is a new API request: each message you send, and each batch of tool results the agent sends back. The model keeps nothing between requests, so the agent re-sends everything: system prompt, tool definitions, project instructions, the whole conversation, then the new part. Nearly all of it matches the request before, and prompt caching is how the provider avoids processing it twice.
Claude Code manages the cache for you; on the API, one top-level cache_control field turns on automatic caching. OpenAI's models, which Codex CLI runs on, follow the same prefix rule with their own prices and lifetimes. The figures below are Claude's.
An exact prefix, in a fixed order
The cache compares the start of each request, the prefix, with what it processed recently. The match is exact: one changed character invalidates everything after it, and there is no per-file or per-section caching.
The prefix is read in a fixed order: tool definitions, then the system prompt, then the messages. A change at one level invalidates that level and every level after it. Edit a message and the tools and system prompt stay cached; change a tool definition and nothing does.
Claude Code puts what rarely changes first: its instructions and tool definitions, then your project context (CLAUDE.md, memory, rules), then the conversation. On a normal turn the entire previous request is the prefix, and only the latest exchange is new.
What a cache write and a cache read cost
Caching changes the price of input, never the answer. As of September 2026, Anthropic prices it in multiples of each model's base input price: a cache write costs 1.25 times base input with the 5-minute lifetime and 2 times with the 1-hour lifetime, and a cache read costs a tenth, except on Claude Fable 5.1 (0.025 times) and Claude Opus 5.5 (0.05 times). Anything after the last cache point is ordinary input.
| Model | Input | Output | Cache read | Cache write 5 min | Cache write 1 h | Context |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 1.00 | 5.00 | 0.10 | 1.25 | 2.00 | 200,000 |
| Claude Sonnet 5 | 2.00 | 10.00 | 0.20 | 2.50 | 4.00 | 1,000,000 |
| Claude Sonnet 4.6 | 3.00 | 15.00 | 0.30 | 3.75 | 6.00 | 1,000,000 |
| Claude Opus 4.6 | 5.00 | 25.00 | 0.50 | 6.25 | 10.00 | 1,000,000 |
| Claude Opus 4.7 | 5.00 | 25.00 | 0.50 | 6.25 | 10.00 | 1,000,000 |
| Claude Opus 4.8 | 5.00 | 25.00 | 0.50 | 6.25 | 10.00 | 1,000,000 |
| Claude Opus 5 | 5.00 | 25.00 | 0.50 | 6.25 | 10.00 | 1,000,000 |
| Claude Opus 5.5 | 4.00 | 20.00 | 0.20 | 5.00 | 8.00 | 1,000,000 |
| Claude Fable 5 | 10.00 | 50.00 | 1.00 | 12.50 | 20.00 | 1,000,000 |
| Claude Fable 5.1 | 10.00 | 50.00 | 0.25 | 12.50 | 20.00 | 1,000,000 |
Take a 100,000-token context on Claude Sonnet 5, at its September 2026 list price of $2 per million input tokens. Read from the cache, it costs $0.02 per request. Without caching, $0.20. Written again after a miss, $0.25 on the 5-minute lifetime or $0.40 on the 1-hour one: a miss costs more than no cache at all, because a write is priced above plain input.
At the 5-minute rate, a rewrite costs 12.5 times a read on most models, 25 times on Opus 5.5 and 50 times on Fable 5.1: the cheaper the read, the wider the gap. A 5-minute write pays for itself on the first read; a 1-hour write needs two.
Five minutes or one hour
The lifetime, or TTL, is how long a cached prefix survives unused, and every read resets it at no extra cost. The clock starts when a request begins, not when its answer ends: if a response takes four minutes to stream, the next request has about a minute left to reuse a 5-minute cache.
As of September 2026, Claude Code gives the main conversation one hour on a Claude subscription within its included usage, and five minutes with an API key, a cloud provider, or once a subscription draws on usage credits. Subagents get five minutes by default either way.
To choose for the main conversation, set promptCacheTtl in your settings, or the CLAUDE_CODE_PROMPT_CACHE_TTL environment variable, to 5m or 1h (Claude Code v2.1.242 or later). The hour pays when gaps between requests run from five to sixty minutes: a meeting, a review, a long build the agent waits on. On bursts that never sit idle for five minutes, it only makes every write cost more.
claude -p "hello" --output-format jsonShows which lifetime your writes used: usage.cache_creation lists one-hour writes under ephemeral_1h_input_tokens and five-minute writes under ephemeral_5m_input_tokens.
What breaks the cache without telling you
A miss raises no error: the turn is slower, and the usage shows a large write where a read should be. The usual causes in a coding agent:
- Switching model. Each model has its own cache, so
/modelmid-task rereads the whole history uncached. So does a skill that names another model, and plan mode underopusplan, which swaps Opus and Sonnet. - Changing effort. On most models each effort level has its own cache. As of September 2026, Opus 5.5 and Fable 5.1 keep it, with an API key or a Claude subscription.
- Changing the tool set. Tool definitions come first, so adding or removing one invalidates everything. Claude Code defers MCP tools by default, loading them on demand through tool search, so server changes stay out of the prefix. When tools load up front instead, as behind a custom
ANTHROPIC_BASE_URLgateway, an MCP server that crashes or reconnects on its own breaks the cache. - Anything that changes early in the prompt. A timestamp in the system prompt means everything after it is rewritten on every request and never read. A Claude Code upgrade usually changes the system prompt too, so the first session after one starts cold.
- Pausing past the lifetime: by default five minutes on an API key, one hour on a subscription.
Why the first turn back costs more
A miss is paid once, on the next request: the whole prefix is processed again and written back at the write rate, and the requests after it read from the cache again. That is why the first message after a break is slow and expensive, even when it is one line long.
Compaction is a miss by design, and a cheap one while the cache is warm, because the summarising request reads the history from the cache. After a long break it rereads the whole history uncached, so a session you have just returned to is the costly moment to compact.
The habits that follow: set model and effort at the start, finish a task before stepping away, and /clear when the subject changes instead of reviving a cold session. To drop a dead end, /rewind beats compacting: it truncates back to a prefix already cached. On Pro and Max plans, Claude Code offers to resume a large session from a summary after a long break.
How to read your cache hit rate
The API reports three input counts per response: cache_read_input_tokens (served from the cache), cache_creation_input_tokens (written to it) and input_tokens, which counts only what came after the last cache point and so looks small in a well-cached session. Total input is the sum of the three.
The hit rate is cache reads divided by that total. In a long session that stays warm, most of each request should be reads. A low rate is normal in short sessions, where the first write is a large share of all input; in a long one, a large write turn after turn points at a cause from the list above.
/usageIn Claude Code v2.1.251 or later, the Prompt cache (main) line shows the share of input served from cache, the misses, and whether the cache is still warm.
From v2.1.260 that line also names the likely cause of the last miss when it can, such as changed tool definitions. It covers the main conversation only: a subagent has its own prompt and tools, so its first request never reads the parent's cache. On the API, the Console's Usage page charts the share of input read from cache, and cache diagnostics (in beta as of September 2026) reports where two consecutive requests diverged.
Subscription or API: who pays for a miss
With an API key, a miss costs money, at the rates in the table, and rate-limit headroom: for most Claude models, cache reads do not count toward the input-tokens-per-minute limit, while writes and uncached input do.
On a Pro or Max plan there is no per-token bill, and the dollar figure in /usage is an estimate meant for API users. A miss is paid from your plan limits instead: Claude Code's documentation lists cache misses among the reasons a long session uses more of them than your activity suggests. On Pro, Max, Team and Enterprise plans, the /usage breakdown flags misses once they reach a tenth of recent usage.
A cached token is still a token
Even at a tenth of the price, the whole context is re-read on every request. In the words of Claude Code's documentation, a one-line question in a session that has been open all day still draws usage for the whole conversation. So the habits stack: keep the front of the prompt stable so it stays cached, and the context small so each read, and each miss, costs little.
If you would rather not keep it small by hand, capsul drives the agent CLI you are already signed into and sends what the task asks for, not the repository, under a token budget you set. It stops at the budget and says what it left out; the benchmark page publishes the measured effect per model, with answer checks beside it.
Questions
Does Claude Code use prompt caching automatically?
Yes. Claude Code manages prompt caching for you unless you turn it off with the DISABLE_PROMPT_CACHING environment variable, and it orders each request so the parts that rarely change come first. As of September 2026, the main conversation gets a one-hour cache on a Claude subscription within plan usage, and a five-minute cache with an API key or a cloud provider.
Is the Claude prompt cache 5 minutes or 1 hour?
Both exist. The API default is five minutes, each read resets the timer at no extra cost, and the one-hour lifetime makes every cache write cost 2 times base input instead of 1.25 times. As of September 2026, Claude Code uses the hour for the main conversation on a Claude subscription and five minutes with an API key, unless you set promptCacheTtl. The hour pays off when gaps between requests run past five minutes.
How much do cache read tokens cost on Claude?
As of September 2026, a cache read costs a tenth of the model's base input price, except on Claude Fable 5.1 (0.025 times) and Claude Opus 5.5 (0.05 times). On Claude Sonnet 5 that is $0.20 per million tokens, against $2 for uncached input. Cache writes cost more than plain input, so a cache only saves money once it is read.
Why is my prompt cache hit rate low?
Usually because something at the front of the prompt keeps changing, or because requests are spaced further apart than the cache lifetime. In a coding agent the common causes are switching model or effort mid-session, tool definitions changing, a timestamp in the system prompt, and pauses longer than five minutes on an API key. Short sessions also score low by construction, because the first write is a large share of all their input. In Claude Code, /usage counts the misses and, from v2.1.260, names the likely cause of the last one.
Does editing CLAUDE.md break the prompt cache?
Not mid-session in Claude Code. The project-root and user-level CLAUDE.md files are read once when the session starts, so an edit neither invalidates the cache nor takes effect until the next /clear, /compact or restart. Nested CLAUDE.md files in subdirectories load later, when Claude first reads a file in that directory.
$ npm i -g @penra/capsul