gpt-5.4-mini in Codex CLI: tokens per question

Mini is the smaller member of this earlier GPT family, suited to well-scoped coding work.

The figures

Bare Codex CLI
42,300

cache-weighted input tokens per question

With capsul
12,100

cache-weighted input tokens per question

less input sent
x3.49

90% interval: x3.05 to x4.03

answer checks passed out of five, bare then with capsul
3 / 4
output tokens per question, bare then with capsul
2,390 / 904

Codex CLI reports tokens, not dollars, on a ChatGPT plan: no cost is shown for GPT models.

It is a sensible candidate for locating code, making a mechanical change or checking a narrow test failure. A smaller model is not automatically the cheapest route to a correct answer if it needs repeated attempts or extra exploration. The answer-check column belongs beside the token count for that reason.

Codex CLI used a ChatGPT subscription in this run, so the page does not turn tokens into dollars. The individual questions reveal where this model handled the task and where it did not. Use them with the interval before extending the measured result to your own workflow.

Question by question

Five one-line questions about a real TypeScript codebase, each asked 2 times per arm. Weighted input per question, mean of the repetitions.

QuestionBarecapsulLess inputCorrect runs, bare / capsul
Q145,70011,500x3.971 / 0
Q235,60021,800x1.632 / 2
Q37,9207,110x1.110 / 2
Q481,20010,400x7.780 / 2
Q541,2009,770x4.222 / 2

How this was measured

Two arms on the same model: Codex CLI as anyone runs it, and the same question through capsul. The unit is cache-weighted input: fresh tokens at full weight, cache reads at a tenth. Campaign of September 4, 2026, 2 repetitions per question and arm.

20 / 20 cells · /benchmarks/2026-09-04-codex.json

Questions

How many tokens does gpt-5.4-mini use per question in Codex CLI?

In the September 4, 2026 benchmark, a bare Codex CLI session on gpt-5.4-mini sent 42,300 cache-weighted input tokens per one-line coding question on average. With capsul on the same model, 12,100.

Is the saving on gpt-5.4-mini established?

Yes. The 90% interval runs from x3.05 to x4.03, entirely above one, on 20 measured cells.

Does capsul change gpt-5.4-mini's answers?

It passed more of them: 3 of five answer checks bare, 4 of five with capsul. Five checks is a small sample: read it as no loss rather than a gain.