gpt-5.6-luna in Codex CLI: tokens per question

Luna is the efficient member of this GPT family for focused, repeatable coding tasks.

The figures

Bare Codex CLI
32,100

cache-weighted input tokens per question

With capsul
13,600

cache-weighted input tokens per question

less input sent
x2.36

90% interval: x1.91 to x3.01

answer checks passed out of five, bare then with capsul
4 / 3
output tokens per question, bare then with capsul
1,180 / 295

Codex CLI reports tokens, not dollars, on a ChatGPT plan: no cost is shown for GPT models.

A precise file reference and a narrow request give this tier a fair chance on lookups, small edits and straightforward tests. A broad investigation may need a stronger model or more exploration, either of which changes the cost of the session. The benchmark therefore keeps token use beside answer checks rather than treating low input as a complete result.

This Codex CLI run was covered by a ChatGPT subscription and reports no per-call dollar cost. Compare its interval and question-level rows with the work you expect to repeat. A difference on one question does not establish the same difference across your entire codebase.

Question by question

Five one-line questions about a real TypeScript codebase, each asked 2 times per arm. Weighted input per question, mean of the repetitions.

QuestionBarecapsulLess inputCorrect runs, bare / capsul
Q139,20013,300x2.962 / 0
Q230,60014,500x2.112 / 2
Q321,20012,700x1.662 / 2
Q448,00019,800x2.430 / 0
Q521,8007,800x2.802 / 2

How this was measured

Two arms on the same model: Codex CLI as anyone runs it, and the same question through capsul. The unit is cache-weighted input: fresh tokens at full weight, cache reads at a tenth. Campaign of September 4, 2026, 2 repetitions per question and arm.

20 / 20 cells · /benchmarks/2026-09-04-codex.json

Questions

How many tokens does gpt-5.6-luna use per question in Codex CLI?

In the September 4, 2026 benchmark, a bare Codex CLI session on gpt-5.6-luna sent 32,100 cache-weighted input tokens per one-line coding question on average. With capsul on the same model, 13,600.

Is the saving on gpt-5.6-luna established?

Yes. The 90% interval runs from x1.91 to x3.01, entirely above one, on 20 measured cells.

Does capsul change gpt-5.6-luna's answers?

It passed fewer: 4 of five answer checks bare, 3 of five with capsul. The benchmark publishes this rather than leaving it out.