gpt-5.6-luna in Codex CLI: tokens per question
Luna is the efficient member of this GPT family for focused, repeatable coding tasks.
The figures
- Bare Codex CLI
- 32,100
- With capsul
- 13,600
- less input sent
- x2.36
- answer checks passed out of five, bare then with capsul
- 4 / 3
- output tokens per question, bare then with capsul
- 1,180 / 295
cache-weighted input tokens per question
cache-weighted input tokens per question
90% interval: x1.91 to x3.01
Codex CLI reports tokens, not dollars, on a ChatGPT plan: no cost is shown for GPT models.
A precise file reference and a narrow request give this tier a fair chance on lookups, small edits and straightforward tests. A broad investigation may need a stronger model or more exploration, either of which changes the cost of the session. The benchmark therefore keeps token use beside answer checks rather than treating low input as a complete result.
This Codex CLI run was covered by a ChatGPT subscription and reports no per-call dollar cost. Compare its interval and question-level rows with the work you expect to repeat. A difference on one question does not establish the same difference across your entire codebase.
Question by question
Five one-line questions about a real TypeScript codebase, each asked 2 times per arm. Weighted input per question, mean of the repetitions.
| Question | Bare | capsul | Less input | Correct runs, bare / capsul |
|---|---|---|---|---|
| Q1 | 39,200 | 13,300 | x2.96 | 2 / 0 |
| Q2 | 30,600 | 14,500 | x2.11 | 2 / 2 |
| Q3 | 21,200 | 12,700 | x1.66 | 2 / 2 |
| Q4 | 48,000 | 19,800 | x2.43 | 0 / 0 |
| Q5 | 21,800 | 7,800 | x2.80 | 2 / 2 |
How this was measured
Two arms on the same model: Codex CLI as anyone runs it, and the same question through capsul. The unit is cache-weighted input: fresh tokens at full weight, cache reads at a tenth. Campaign of September 4, 2026, 2 repetitions per question and arm.
20 / 20 cells · /benchmarks/2026-09-04-codex.json
Questions
How many tokens does gpt-5.6-luna use per question in Codex CLI?
In the September 4, 2026 benchmark, a bare Codex CLI session on gpt-5.6-luna sent 32,100 cache-weighted input tokens per one-line coding question on average. With capsul on the same model, 13,600.
Is the saving on gpt-5.6-luna established?
Yes. The 90% interval runs from x1.91 to x3.01, entirely above one, on 20 measured cells.
Does capsul change gpt-5.6-luna's answers?
It passed fewer: 4 of five answer checks bare, 3 of five with capsul. The benchmark publishes this rather than leaving it out.