Cache TTL

A cache TTL (time to live) is how long a cached prompt prefix survives without being used, after which the next request has to process it and write it to the cache again.

On Anthropic's API the default lifetime is five minutes, and a one-hour lifetime costs more to write: twice the base input price instead of 1.25 times, as of September 2026. The timer is not a fixed expiry. Every request that reads the prefix refreshes it at no extra cost, so a cache stays warm for as long as requests keep arriving inside the window. The clock starts when a request begins, not when its response ends, so a long answer uses up part of the window.

Choosing a lifetime is a bet on pauses. A one-hour cache pays off when work stops and resumes: the first request after a twenty-minute break reads the prefix instead of rebuilding it. On bursts of work that never idle past five minutes, it only raises the write price.

Coding agents usually choose for you. As of September 2026, Claude Code requests the one-hour lifetime for the main conversation on a Claude subscription within the plan's included usage, and five minutes with an API key, a cloud provider or usage credits, while subagents get five minutes by default. The promptCacheTtl setting overrides the choice for the main conversation.

Whatever the lifetime, the first message after it lapses is the expensive one, because it misses the cache and reprocesses the entire context. With reads at a tenth of the input price and writes at 1.25 times or more, rebuilding a large context once costs more than reading it from the cache ten times.

← All terms