gpt-5.4 in Codex CLI: tokens per question

This earlier GPT model is another reference point for Codex CLI coding sessions.

The figures

Bare Codex CLI
31,800

cache-weighted input tokens per question

With capsul
8,220

cache-weighted input tokens per question

less input sent
x3.86

90% interval: x2.88 to x5.47

answer checks passed out of five, bare then with capsul
4 / 4
output tokens per question, bare then with capsul
1,660 / 467

Codex CLI reports tokens, not dollars, on a ChatGPT plan: no cost is shown for GPT models.

A familiar model can still use very different amounts of input depending on how much the agent reads before answering. The measurements below hold the coding questions constant while changing how context is supplied. That makes the row useful for judging token traffic without guessing from the model name.

No per-question dollar figure is reported for this subscription run. Read the interval with the answer checks and the individual tasks. If you use this model for larger edits or longer conversations, the input accumulated in those sessions deserves its own measurement.

Question by question

Five one-line questions about a real TypeScript codebase, each asked 2 times per arm. Weighted input per question, mean of the repetitions.

QuestionBarecapsulLess inputCorrect runs, bare / capsul
Q136,80010,400x3.530 / 0
Q230,6005,970x5.122 / 2
Q317,0003,310x5.121 / 2
Q456,00015,900x3.532 / 2
Q518,5005,520x3.352 / 2

How this was measured

Two arms on the same model: Codex CLI as anyone runs it, and the same question through capsul. The unit is cache-weighted input: fresh tokens at full weight, cache reads at a tenth. Campaign of September 4, 2026, 2 repetitions per question and arm.

20 / 20 cells · /benchmarks/2026-09-04-codex.json

Questions

How many tokens does gpt-5.4 use per question in Codex CLI?

In the September 4, 2026 benchmark, a bare Codex CLI session on gpt-5.4 sent 31,800 cache-weighted input tokens per one-line coding question on average. With capsul on the same model, 8,220.

Is the saving on gpt-5.4 established?

Yes. The 90% interval runs from x2.88 to x5.47, entirely above one, on 20 measured cells.

Does capsul change gpt-5.4's answers?

Not on this protocol: gpt-5.4 passed 4 of five answer checks bare and 4 of five with capsul.