gpt-5.5 in Codex CLI: tokens per question

This earlier GPT release gives Codex CLI users a baseline against the newer family variants.

The figures

Bare Codex CLI
47,200

cache-weighted input tokens per question

With capsul
8,130

cache-weighted input tokens per question

less input sent
x5.81

90% interval: x4.29 to x7.89

answer checks passed out of five, bare then with capsul
4 / 4
output tokens per question, bare then with capsul
1,500 / 81.2

Codex CLI reports tokens, not dollars, on a ChatGPT plan: no cost is shown for GPT models.

Model selection affects how an agent approaches a coding question, but the surrounding context is still part of every request. Use this row to inspect the input sent by bare Codex CLI and by the same task with capsul. It is especially useful if this is the model already available in your everyday setup.

The benchmark records tokens and answer checks, not a dollar bill for a ChatGPT subscription. Compare the task rows before attributing a difference to the model's age or tier. The observations describe short questions under a fixed protocol, not every longer session you might run.

Question by question

Five one-line questions about a real TypeScript codebase, each asked 2 times per arm. Weighted input per question, mean of the repetitions.

QuestionBarecapsulLess inputCorrect runs, bare / capsul
Q163,8008,500x7.500 / 0
Q240,90012,300x3.312 / 2
Q321,5003,700x5.792 / 2
Q483,00010,200x8.132 / 2
Q527,0005,910x4.562 / 2

How this was measured

Two arms on the same model: Codex CLI as anyone runs it, and the same question through capsul. The unit is cache-weighted input: fresh tokens at full weight, cache reads at a tenth. Campaign of September 4, 2026, 2 repetitions per question and arm.

20 / 20 cells · /benchmarks/2026-09-04-codex.json

Questions

How many tokens does gpt-5.5 use per question in Codex CLI?

In the September 4, 2026 benchmark, a bare Codex CLI session on gpt-5.5 sent 47,200 cache-weighted input tokens per one-line coding question on average. With capsul on the same model, 8,130.

Is the saving on gpt-5.5 established?

Yes. The 90% interval runs from x4.29 to x7.89, entirely above one, on 20 measured cells.

Does capsul change gpt-5.5's answers?

Not on this protocol: gpt-5.5 passed 4 of five answer checks bare and 4 of five with capsul.