Cache write

A cache write is the processing and storing of a prompt prefix for reuse, billed above the normal input price in exchange for much cheaper reads on the requests that follow.

Caching is paid for on the way in. On Anthropic's API, as of September 2026, writing a prefix costs 1.25 times the base input price for the default five-minute lifetime and twice the base price for the one-hour lifetime, while reading it back costs a tenth of the base price on most models. Responses report the written part separately, as cache_creation_input_tokens.

The arithmetic decides when caching pays. Anthropic's pricing puts the break-even at one read for a five-minute write and two reads for a one-hour write. A prefix that is written and never read costs more than sending it uncached, and a prefix shorter than the model's minimum (512 to 4,096 tokens depending on the model, as of September 2026) is not cached at all.

In an agent session, writes happen at predictable moments: on the first request, on the newly added part of each turn so that the next turn can read it, and across the whole context after anything that breaks the prefix, such as a model switch, a changed tool list, an edited system prompt or a pause longer than the cache lifetime.

A healthy session therefore reads far more than it writes. When cache_creation_input_tokens stays high turn after turn, something early in the request keeps changing, and every turn pays the write premium without collecting the discount.

← All terms