gpt-5.6-terra in Codex CLI: tokens per question

Terra is the balanced member of this GPT family, positioned between flagship capability and lower-cost throughput.

The figures

Bare Codex CLI
24,000

cache-weighted input tokens per question

With capsul
11,100

cache-weighted input tokens per question

less input sent
x2.16

90% interval: x1.57 to x3.20

answer checks passed out of five, bare then with capsul
3 / 4
output tokens per question, bare then with capsul
582 / 142

Codex CLI reports tokens, not dollars, on a ChatGPT plan: no cost is shown for GPT models.

It is a plausible default for routine implementation, test repair and code review when the task needs more judgement than a mechanical edit. In an agent session, the context gathered to understand a repository can dominate the user's prompt. The per-question rows show how that traffic changed under the measured protocol.

Codex CLI ran under a ChatGPT subscription here, so no dollar estimate is attached to this model. The useful evidence is the token difference together with the answer checks. If your work is longer or more exploratory, measure that workflow before assuming this short-question result will carry across.

Question by question

Five one-line questions about a real TypeScript codebase, each asked 2 times per arm. Weighted input per question, mean of the repetitions.

QuestionBarecapsulLess inputCorrect runs, bare / capsul
Q126,6006,100x4.370 / 0
Q221,8009,960x2.192 / 2
Q315,7009,620x1.632 / 2
Q433,30024,400x1.370 / 2
Q522,4005,480x4.082 / 2

How this was measured

Two arms on the same model: Codex CLI as anyone runs it, and the same question through capsul. The unit is cache-weighted input: fresh tokens at full weight, cache reads at a tenth. Campaign of September 4, 2026, 2 repetitions per question and arm.

20 / 20 cells · /benchmarks/2026-09-04-codex.json

Questions

How many tokens does gpt-5.6-terra use per question in Codex CLI?

In the September 4, 2026 benchmark, a bare Codex CLI session on gpt-5.6-terra sent 24,000 cache-weighted input tokens per one-line coding question on average. With capsul on the same model, 11,100.

Is the saving on gpt-5.6-terra established?

Yes. The 90% interval runs from x1.57 to x3.20, entirely above one, on 20 measured cells.

Does capsul change gpt-5.6-terra's answers?

It passed more of them: 3 of five answer checks bare, 4 of five with capsul. Five checks is a small sample: read it as no loss rather than a gain.