gpt-5.6-sol in Codex CLI: tokens per question
Sol is the flagship member of this GPT family in Codex CLI, intended for demanding coding work.
The figures
- Bare Codex CLI
- 33,100
- With capsul
- 14,700
- less input sent
- x2.26
- answer checks passed out of five, bare then with capsul
- 3 / 3
- output tokens per question, bare then with capsul
- 1,020 / 225
cache-weighted input tokens per question
cache-weighted input tokens per question
90% interval: x1.81 to x3.01
Codex CLI reports tokens, not dollars, on a ChatGPT plan: no cost is shown for GPT models.
It is the tier to consider for an ambiguous bug, a complex implementation or a review that needs sustained judgement. Those tasks can also invite broad file exploration, which is why the input the agent sends matters alongside the model choice. The figures below compare the same short questions on bare Codex CLI and with capsul.
This run used a ChatGPT subscription, so the page reports tokens rather than a per-question dollar amount. Check the interval and answer checks before using the measured difference as a reason to change your workflow. A long project may accumulate context differently from these brief tasks.
Question by question
Five one-line questions about a real TypeScript codebase, each asked 2 times per arm. Weighted input per question, mean of the repetitions.
| Question | Bare | capsul | Less input | Correct runs, bare / capsul |
|---|---|---|---|---|
| Q1 | 36,000 | 19,200 | x1.88 | 0 / 0 |
| Q2 | 23,900 | 9,850 | x2.43 | 1 / 2 |
| Q3 | 24,600 | 7,830 | x3.14 | 2 / 2 |
| Q4 | 61,400 | 21,900 | x2.80 | 0 / 0 |
| Q5 | 19,400 | 14,500 | x1.34 | 2 / 2 |
How this was measured
Two arms on the same model: Codex CLI as anyone runs it, and the same question through capsul. The unit is cache-weighted input: fresh tokens at full weight, cache reads at a tenth. Campaign of September 4, 2026, 2 repetitions per question and arm.
20 / 20 cells · /benchmarks/2026-09-04-codex.json
Questions
How many tokens does gpt-5.6-sol use per question in Codex CLI?
In the September 4, 2026 benchmark, a bare Codex CLI session on gpt-5.6-sol sent 33,100 cache-weighted input tokens per one-line coding question on average. With capsul on the same model, 14,700.
Is the saving on gpt-5.6-sol established?
Yes. The 90% interval runs from x1.81 to x3.01, entirely above one, on 20 measured cells.
Does capsul change gpt-5.6-sol's answers?
Not on this protocol: gpt-5.6-sol passed 3 of five answer checks bare and 3 of five with capsul.