Claude Code vs Codex CLI: where each one spends its tokens

Which agent uses fewer tokens depends more on the model, the setup and the task than on the agent. Here is where each one spends input, what each shows you, how each bills, and what measuring both actually says.

Claude Code and Codex CLI are CLI coding agents built the same way: a model in a loop that reads files, runs commands and edits code in the directory you started it in. So they spend tokens the same way. Every request carries far more than your question, and the question is the smallest part of it.

The short answer to which one uses fewer tokens is that neither does reliably. The model you pick, the instructions and tools you have switched on, and how much the agent explores move the total more than the choice of agent.

Where the tokens go, on both sides

Each request carries three layers of input, in both tools:

  • Ambient input: the agent's own instructions, the definition of every tool the model may call, your project instructions (CLAUDE.md for Claude Code, AGENTS.md for Codex) and whatever skills, plugins or MCP servers are switched on. It is there before you type anything.
  • Exploration: every file read, search and command output joins the conversation and stays there for the rest of the session.
  • Transcript: the whole conversation is sent again with every request, including each tool result so far. One question usually takes several requests, one per step the agent takes.

Prompt caching softens the transcript on both sides: a prefix the provider has already processed is billed as a cache read, at a tenth of the fresh rate for most models as of September 2026, on Anthropic's price list and on OpenAI's alike. Both agents also compact the history into a summary when the window fills, and /compact does it on demand.

What each one loads before you type

Claude Code orders every request the same way: its core instructions and tool definitions first, then your project context (CLAUDE.md, auto memory, rules), then the conversation. MCP tools are deferred by default on supported models, so an idle server costs its tool names and instructions rather than its full schemas until Claude needs one. Skills cost a one-line description each until invoked.

Codex CLI builds an instruction chain from AGENTS.md files when it starts, one from your Codex home and then one per directory from the project root down to where you are, and stops adding files at 32 KiB by default. It adds a catalogue of available skills, whose token budget defaults to two percent of the model's context window. MCP servers in config.toml contribute their tools, deferred behind a search step when the model supports it and sent in full otherwise; enabled_tools and disabled_tools trim what each server exposes.

Neither list is fixed: it depends on what you have installed. That is the first reason a comparison run on someone else's machine does not transfer to yours.

What each one shows you

Claude Code reports usage in tokens and in estimated dollars:

/usage

Claude Code's /usage screen: tokens by model, cache reads and writes, a list-price estimate and, on a subscription, your plan usage.

/context

Claude Code: what is filling the context window right now, by category.

Codex CLI reports tokens and, when you are signed in with ChatGPT, the share of your limits you have used:

/status

Codex: the model, the context window used and left, and how much of the five-hour and weekly limits remains.

codex debug prompt-input 'where is the retry limit set'

Prints, as JSON, the input list Codex would send for that prompt, without calling the model.

/usage in Codex shows your account's token activity by day, by week or in total, and quitting prints a token line for the session. The debug prompt-input output leaves out the base instructions and the tool schemas, which travel beside that list, so read it as a floor rather than the whole request.

How each one bills

Both agents run either on a subscription or on pay-per-token billing, and the two meter differently:

  • Claude Code on Claude Pro or Max: usage counts against a session limit that resets every five hours and a weekly limit, shared with your use of Claude itself. Team and Enterprise seats work the same way. The /usage dollar figure is a list-price estimate that Anthropic says is not relevant for billing on Pro and Max.
  • Claude Code on an API key or the Claude Console: billed per token at API rates, cache reads and writes included. An ANTHROPIC_API_KEY left in your environment takes over from your subscription once you approve it, and claude -p uses it without asking.
  • Codex CLI signed in with ChatGPT: usage counts against your plan's allowance, which OpenAI's pricing page estimates in local messages per five-hour window, with weekly limits that may also apply. Past the included usage, credits let you continue. On Plus and Pro, Codex shows you tokens and percentages of your limits, not dollars.
  • Codex CLI with an API key: billed through your OpenAI Platform account at standard API rates.

What our benchmark measured

We ran both agents on their own, the way anyone runs them (claude -p and codex exec with their usual configuration, not Claude Code's --bare mode), on the same five one-line questions about one repository. The unit is cache-weighted input per question: everything the agent sent to answer it, across all its requests, with fresh tokens at full weight and cache reads at a tenth. Each row is one model.

For each model, the tables give the agent on its own, which is the figure to read when comparing agents, and the same question sent through capsul on the same model, which is what the benchmark was built to measure. The answer checks count, out of five, the questions whose answer contained the expected detail, checked automatically.

ModelInput, bareInput, capsulCost, bareCost, capsulLess input
Claude Haiku 4.555,50019,300$0.0812$0.0265x2.88
Claude Sonnet 554,20018,900$0.153$0.0487x2.86
Claude Sonnet 4.631,40016,600$0.139$0.0613x1.89
Claude Opus 4.645,70016,400$0.332$0.101x2.78
Claude Opus 4.738,70020,100$0.28$0.126x1.92
Claude Opus 4.865,80013,800$0.498$0.0825x4.75
Claude Opus 545,30016,200$0.343$0.101x2.80
Claude Opus 5.542,30015,900$0.222$0.0696x2.66
Claude Fable 541,40017,600$0.633$0.219x2.35
Claude Fable 5.149,70023,400$0.636$0.279x2.12
Claude models in Claude Code: cache-weighted input and API-equivalent cost per one-line coding question, bare agent vs with capsul. Computed from the published datasets: Claude campaign of September 22, 2026, Codex campaign of September 4, 2026.

The Claude table adds a cost column: what Claude Code itself reports those tokens would cost per question at API list prices. On a subscription you do not pay that figure.

ModelInput, bareInput, capsulLess input
gpt-5.6-sol33,10014,700x2.26
gpt-5.6-terra24,00011,100x2.16
gpt-5.6-luna32,10013,600x2.36
gpt-5.547,2008,130x5.81
gpt-5.431,8008,220x3.86
gpt-5.4-mini42,30012,100x3.49
GPT models in Codex CLI: cache-weighted input per one-line coding question, bare agent vs with capsul. Codex reports tokens, not dollars. Computed from the published datasets: Claude campaign of September 22, 2026, Codex campaign of September 4, 2026.

The Codex table has no cost column, because on a ChatGPT plan Codex reports tokens, not dollars, and that is the plan these runs used.

Why the two tables are not a ranking

Read side by side, the two tables are indicative at best, for five reasons:

  • Different dates and code. The Claude rows were measured on 22 September 2026 and the main Codex campaign on 4 September 2026, each on the agent versions and capsul code of its day.
  • Different models. No model appears in both tables, so every cross-agent comparison is also a comparison of models.
  • Different tokens. Each vendor cuts text into tokens its own way, and Anthropic notes that its newer models count the same text as roughly 30 percent more tokens than its older ones. A token is not a fixed amount of text.
  • Different ledgers. Claude reports cache writes, which the benchmark weights above fresh input, as Anthropic bills them. The Codex runs reported none, and Codex's credit billing on ChatGPT plans has no separate cache-write charge.
  • Different repetitions. Each Claude question ran three times per arm and each Codex question twice, and the first run writes the cache that the later ones read.

What survives those caveats is still useful. On a one-line question, both agents on their own send tens of thousands of cache-weighted tokens, nearly all of it context rather than the question. And within each table the spread between models is wide: the model you pick moves the number at least as much as the agent you pick.

What to do with this

If you already pay for one of the two plans, switching agents to save tokens is rarely the lever that matters. The ones that do are the same in both tools:

  • Start each task in a fresh session. /clear does it in both, and the old transcript stops being re-sent.
  • Name the file you mean. Exploration is the layer that varies most from one run to the next.
  • Keep the ambient layer small: a short CLAUDE.md or AGENTS.md, and no MCP servers or plugins you are not using.
  • Match the model and the reasoning effort to the task: /model and /effort in Claude Code, /model in Codex, which sets both where the model allows it.

To compare the two on your own code, measure a real task, not anyone's table, ours included. Run it three times per agent, each in a fresh session in the same repository: the first run writes the cache and the next ones read it. With claude -p and --output-format json, add up the per-model input, cache-read and cache-write counts, which include subagents (the plain usage field does not). With codex exec --json, read the usage on the final turn.completed event and subtract the cached part from the input. Weight fresh input fully, cache reads at a tenth and any reported cache writes at their billed premium, and you have a cost per task, in tokens, for each agent on your own work.

If you use both agents, capsul drives either one from the same command, under a token budget you set, and keeps a local record of what each one spends.

capsul burn

Local spend, broken down by agent, model and project.

Questions

Does Claude Code or Codex CLI use fewer tokens?

Neither does reliably. The model, the instructions and tools you have switched on, and how much the agent explores move the total more than the choice of agent, and each vendor cuts text into tokens differently. In our benchmark of both agents run on their own, the spread between models on one agent was at least as wide as the gap between the two agents.

Is Codex CLI cheaper than Claude Code?

On a subscription the price is the plan, and what differs is how long it lasts before its five-hour or weekly limit. On an API key, both bill per token at their vendor's list prices, with cached input at about a tenth of the fresh rate for most models. Either way, the cheaper agent for you is the one that answers your tasks with a smaller model and less context, which only a measurement on your own code can show.

Why does Claude Code show a dollar cost if I pay for Pro or Max?

/usage computes the session's cost from its token counts at API list prices, which is what the same work would cost on an API key. Anthropic's documentation says that for Pro and Max subscribers this figure is not relevant for billing. What counts against your plan appears on the same screen as usage bars.

How do I see Codex CLI token usage on a ChatGPT plan?

Type /status for the context window used and left and the share of your five-hour and weekly limits remaining, or /usage for your account's token activity by day, by week or in total. When you quit, Codex prints the session's token counts: its total adds fresh input and output, and cached input is listed separately.

Can I compare the token counts the two agents report?

Not as printed. Claude reports input, cache reads and cache writes as separate counts, while the input_tokens in Codex's JSON output already includes cached input. Convert both to cache-weighted input, fresh tokens at full weight and cache reads at a tenth, and compare the same task run several times on each.

$ npm i -g @penra/capsul

← All guides