Cache-weighted input tokens

Cache-weighted input tokens measure a request's input with each token counted in proportion to its price relative to fresh input: fresh tokens in full, cache reads at a tenth and cache writes at their premium.

Raw token counts treat every input token alike, and billing does not. On Anthropic's API, as of September 2026, a token read from the prompt cache costs a tenth of a fresh one on most models, and a token written to the cache costs 1.25 or 2 times as much, depending on the cache lifetime. Two runs with identical raw input can therefore differ severalfold in cost, depending only on how much of that input was cached.

A raw count misleads in both directions. Adding cached tokens at full weight overstates a session that caches well. Counting only the uncached field understates it: in Anthropic's usage report, input_tokens covers only what was neither read from nor written to the cache, and the total is that figure plus cache_read_input_tokens and cache_creation_input_tokens. A comparison of two ways of running the same task has to say which count it used; otherwise it flatters one side without saying so.

Weighting converts all input into fresh-input equivalents: fresh tokens times one, cache reads times 0.1, cache writes times their write multiplier. The result follows what input costs without being tied to any one model's price, which keeps it comparable across models. The fixed weights are a convention, not a price list: some newer models charge less than a tenth for a cache read, 0.05 on Claude Opus 5.5 and 0.025 on Claude Fable 5.1 as of September 2026.

It measures input only. Output, thinking tokens included, is priced separately and sits outside the figure. Nor is it a subscription's usage meter, which the provider computes in its own way.

← All terms