Claude Code subagents: what they save, and what they cost
A subagent keeps its reading out of your conversation, not off your bill. Whether it saves tokens depends on what it would otherwise have left in your context, and for how many turns.
A subagent is a separate Claude worker that Claude Code starts inside your session. It has its own context window, its own system prompt and its own tools. It makes its tool calls in that window, reading, searching and running commands, then hands back only its final answer. Claude also delegates on its own, for example to the built-in Explore subagent when it needs to search code.
So subagents save context, not necessarily tokens. Your main conversation stays small because the reading never lands in it. The total processed across all agents can be higher, because every subagent pays for its own setup and its own exploration, and several at once multiply both.
What a subagent pays for
A subagent starts without your conversation history, the skills already invoked or the files Claude has already read, working from a message Claude writes to hand over the task. Every request it makes carries:
- Its own system prompt, shorter than your session's, and the tools it may use, MCP tools included.
- Your CLAUDE.md files and a git status snapshot. The built-in Explore and Plan subagents skip both to start cheaper.
- The task message, plus the full text of any skills its definition preloads.
Then it runs its own agentic loop. The result of each tool call, every file read and every command output, joins its context, and each next step re-sends that growing context. The longer the search, the more this loop outweighs the setup.
Caching helps less than in your main thread. A subagent's prefix differs from your conversation's, so its first request cannot read your prompt cache and its setup is processed fresh. The cache it then builds lasts five minutes by default, even on a subscription where your main conversation's lasts an hour. The subagentPromptCacheTtl setting can raise it to an hour, at a higher cache-write price.
Finally the result comes back: the subagent's last message, plus a short trailer with token counts and duration. It joins your conversation and is re-sent with every later turn, so a subagent asked for a two-line verdict costs your main thread less than one asked for a full report.
What it saves
The saving is everything the subagent read and ran that never reached your conversation: the dozen files opened to find one function, the thousands of lines of test output behind three failures. Done in your main thread, all of it would stay in context and be re-sent on every turn until you cleared or compacted. Prompt caching bills those re-reads at a reduced rate, but they still count against your limits and still crowd the window the model reasons over.
Claude Code's documentation walks through the shape of it: a research subagent reads about 6,100 tokens of files, and the main conversation receives a 420-token result. Your context grew by the result. The reading was still paid for, in the subagent's window.
So a subagent wins when what it keeps out is large and your session has many turns left, and loses when the task is small or it must re-read what your conversation already holds.
When subagents pay off
- Exploring code you do not know yet, where most of what gets read turns out to be irrelevant.
- Verbose operations you only need a verdict from: a test run reduced to its failures, a log to its errors, a page of documentation to one answer.
- Long tasks. The more turns your main session has ahead, the more each token kept out of it is worth.
- Independent investigations in parallel. They finish in the time of the slowest one rather than the sum of all, which saves time, not tokens: each one pays its own setup.
- Work a smaller model does well, such as searching, summarising and running checks.
When they waste tokens
- Small, targeted changes. When you already know the file, the setup outweighs the exploration it replaces, and a fresh subagent needs time to gather context before it can act.
- Work that depends on what your conversation already knows. A fresh subagent re-reads files Claude has already read. Frequent back-and-forth, or phases that share context such as planning, implementing and testing, belong in the main conversation.
- Overlapping parallel agents. Three subagents that each need the same core module each read it, and each result lands in your context.
- Many agents at once on a subscription. Every subagent draws on the same plan limits as your main conversation, and Claude Code's documentation says plainly that running several at once multiplies token usage.
Delegation can also nest. As of September 2026 a subagent can start its own, up to three layers below your conversation by default, and up to 20 can run at once. Setting CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH to 1 turns nesting off, and CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS changes the ceiling.
Two other tools cover part of the same ground. /btw answers a side question from what your conversation already holds, with no tools, and keeps the answer out of history. /subtask starts a fork, a subagent that inherits your conversation and reads its cache. The documentation calls that cheaper than a fresh subagent for tasks that need the same context, though a fork starts as large as your conversation.
Put a subagent on a cheaper model
A custom subagent is a Markdown file in .claude/agents/ for one project or ~/.claude/agents/ for all your projects, with YAML frontmatter above its system prompt. The model field takes haiku, sonnet, opus, fable, a full model ID such as claude-opus-5-5, or inherit. For a subagent that searches, summarises or runs checks, add model: haiku.
Claude Code picks a subagent's model in this order: a model Claude passes for that one call, the definition's model field, the CLAUDE_CODE_SUBAGENT_MODEL environment variable, then your main conversation's model. That last step has a cost: switch to Opus with /model and every subagent that inherits your model switches with you.
The built-in Explore subagent no longer defaults to Haiku. Since Claude Code v2.1.198 it inherits your session's model, capped at Opus on the Claude API. To keep exploration cheap, create your own subagent named Explore with model: haiku, which overrides the built-in one; unlike the built-in, it loads your CLAUDE.md files unless you add omitClaudeMd: true. CLAUDE_CODE_SUBAGENT_MODEL alone does not move Explore or Plan; adding CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1 (v2.1.257 or later) applies it to them too.
As of September 2026, Anthropic's API list prices put Claude Haiku 4.5 at $1 per million input tokens and $5 per million output, against $2 and $10 for Sonnet 5 and $4 and $20 for Opus 5.5. For subscribers, Anthropic's help centre puts it as Opus costing several times more per turn than Sonnet, and Sonnet more than Haiku.
One limit applies: a subagent's context window is sized by its own model, and Haiku 4.5 has a 200K-token window where Sonnet 5 and Opus 5.5 have 1M. Two more fields bound the spend: maxTurns caps how many agentic turns a subagent may take, and on Sonnet or Opus, effort: low lowers its reasoning effort without changing your session's.
/tasksLists running and just-finished subagents, with the model each one actually runs on.
See what subagents consumed
For totals, use the tools that add everything up:
/usageon a Pro, Max, Team or Enterprise plan attributes recent usage to subagents, skills, plugins and MCP servers, each as a share of the total.dandwswitch between the last 24 hours and seven days. It reads local session history, so other machines and claude.ai are not included, and itsPrompt cache (main)line covers your main conversation only.- In scripts,
claude -p --output-format jsonreturnstotal_cost_usdand a per-model breakdown. The Agent SDK documentation is explicit that theusagefield counts only the top-level loop, whiletotal_cost_usdandmodelUsageinclude subagent requests. - For a team, the OpenTelemetry token and cost counters tag each request with
query_source, set tosubagentfor subagent requests, and withagent.name. - Each subagent's transcript, showing what it read and ran, is kept under
~/.claude/projects/{project}/{sessionId}/subagents/for 30 days by default.
/usageOn subscription plans, the breakdown shows what share of recent usage went to subagents.
Before you delegate
The cheapest exploration, delegated or not, is the one that does not happen. Name the file when you know it, and ask the narrow question rather than the broad one. Keep each custom subagent's description short: descriptions take up context in your main conversation, while a subagent's system prompt loads only when it runs.
If the exploration is the part you want to cap, capsul drives the Claude Code you are already signed into, sends what the task asks for rather than the repository, and stops at the budget you set, telling you what it left out.
capsul ask 'why does the session expire early' --budget 3000Nothing goes over the ceiling, and what was left out is reported.
Questions
Do subagents save tokens in Claude Code?
They save space in your main conversation, not necessarily tokens overall. Each subagent pays for its own setup and its own reads, and its first request cannot use your conversation's prompt cache. It comes out ahead when it keeps bulky output, such as a wide search or a long test log, out of a session with many turns left.
Which model does a Claude Code subagent use?
A custom subagent runs on the model in its model field (haiku, sonnet, opus, fable, a full model ID, or inherit); without one, on the CLAUDE_CODE_SUBAGENT_MODEL default if you set it, otherwise on your main conversation's model. A model Claude passes for a single call overrides all of these. As of September 2026 the built-in Explore subagent inherits your session's model, capped at Opus on the Claude API, so a custom subagent named Explore with model: haiku is how to keep searching cheap.
How do I see how many tokens my subagents used?
On a Pro, Max, Team or Enterprise plan, /usage shows the share of recent usage that went to subagents, over the last 24 hours or seven days. In scripts, read total_cost_usd or the per-model breakdown rather than usage, which the Agent SDK documentation says leaves subagent requests out. The token count shown beside a subagent reflects its context size, not the total it processed.
Do parallel subagents use up my Claude plan faster?
Yes. Each subagent makes its own requests against the same plan limits as your main conversation, and Claude Code's documentation states that running several subagents at once multiplies token usage. Running them in parallel saves wall-clock time, because independent tasks finish in the time of the slowest one, but it does not save tokens.
Are subagents the same as agent teams?
No. Subagents run inside one session and report back to it. Agent teams, experimental and off by default, run separate Claude Code instances that message each other, and Anthropic's cost page estimates they use about seven times the tokens of a standard session when teammates run in plan mode.
$ npm i -g @penra/capsul