Context compaction
Context compaction replaces an agent's conversation history with a summary written by the model, so that a long session fits back inside the context window at the price of detail.
It happens in two ways: automatically, when the context approaches the window's limit, and on request. In Claude Code the command is /compact, which accepts a focus such as /compact focus on the API changes. Codex CLI has its own /compact, which summarizes the visible chat to free tokens.
The summary is not free. It is written in a separate request that carries the whole conversation, so compacting a large context is itself a large request. Claude Code's documentation notes that while the prompt cache is warm, that request reads most of its input from the cache and costs a fraction of what its size suggests, but after a break longer than the cache lifetime it reprocesses the full history as uncached input. The next request then starts a new, shorter history, which has to be cached afresh.
What survives depends on where it came from. In Claude Code, the system prompt still applies, and the project-root CLAUDE.md and auto memory are reloaded from disk. Full tool outputs and intermediate reasoning are gone, replaced by the summary, and an instruction given only in conversation may be lost with them.
Compaction keeps a session going; it does not make it cheap. When the next task is unrelated, a fresh session carries nothing forward, and /clear in Claude Code costs nothing. When continuity matters, compacting at a natural break, with a focus, beats letting the automatic pass choose its moment in the middle of a task.