What MCP servers cost in tokens, and how to cut it
An MCP server costs tokens in three places: the tool list the agent carries, the definitions it loads, and the results it keeps. Here is how to see each one, and the switches that bring it down.
A model can only call a tool that has been described to it, and the description travels inside the request: a name, a few sentences of text and a JSON schema for the inputs. Anthropic's tool use pricing counts all of it as input tokens, like any other part of the prompt. So MCP servers do use tokens, and in a coding agent they use them on every turn, because every turn re-sends the whole context.
How much depends on whether the agent defers those definitions. As of September 2026, Claude Code and current Codex CLI releases both hold most MCP tool definitions back until the model asks for them, on the models that support it. An idle server is now cheap but not free, and a few configurations quietly put the full cost back.
What an MCP server adds to a request
- Tool definitions: the name, description and JSON input schema of each tool. Anthropic's tool search documentation gives the example of five common servers (GitHub, Slack, Sentry, Grafana and Splunk), which can put about 55,000 tokens of definitions in front of the model before it does any work.
- Server instructions: a short text the server returns when it connects, telling the model what it is for. Claude Code loads it at session start and cuts it at 2,048 characters by default.
- Calls and results: each call the model makes and each result the server returns become part of the transcript. A result holding an issue list or a database schema can run to thousands of tokens, and it is re-sent on every later turn, at the cached rate while the cache holds, until compaction or clearing removes it.
- The tool-use system prompt: whenever a request carries any tool, the Claude API adds its own instructions, 286 tokens on Opus 5.5, 354 on Sonnet 5 and 496 on Haiku 4.5 in Anthropic's published table as of September 2026. A coding agent pays this anyway, because its built-in tools are tools too; MCP servers do not multiply it.
Deferred or upfront: know which one you have
Claude Code turns tool search on by default. At session start only tool names and server instructions enter the context, and the full schema of a tool loads when Claude searches for it. Anthropic says tool search typically cuts definition overhead by more than 85 percent, loading the three to five tools a request needs. A loaded definition then stays in the conversation and is re-sent with the rest of the history.
MCP definitions go back into every request, in full, in these cases:
ANTHROPIC_BASE_URLpoints to a proxy or gateway that is not Anthropic's. SetENABLE_TOOL_SEARCH=trueif your gateway forwards the blocks tool search relies on.ENABLE_TOOL_SEARCH=falseis set, orCLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS, which keeps tool search off even whenENABLE_TOOL_SEARCHis also set.ENABLE_TOOL_SEARCH=autois set: definitions then load upfront as long as they total less than 10 percent of the context window.- A server is marked
alwaysLoad: true, for that server's tools. - The deployment rejects tool search: Claude models older than the 4.5 generation on Google Cloud's Agent Platform, and Microsoft Foundry deployments hosted on Azure.
Seeing the overhead in Claude Code
/context allBreaks the context window down by category and lists the tokens each loaded MCP tool uses.
/context shows everything in the window by category: system prompt, system tools, MCP tools, memory files, skills and messages, and all expands the per-item breakdown. Run it at the start of a session to see what your setup costs before you type anything, and again after a few tool calls to see what the results added.
/mcpLists each configured server with its status and tool count, and turns servers on or off for the current project.
On a Pro, Max, Team or Enterprise plan, /usage also attributes recent usage to individual MCP servers, counting only the requests that carried one of that server's results. It tells you which server is expensive to use rather than which is expensive to keep, which is usually the more useful answer.
Claude Code also warns when a single MCP result exceeds 10,000 tokens, and caps results at 25,000 tokens by default through MAX_MCP_OUTPUT_TOKENS. A larger result is saved to a file, and the conversation receives the path instead of the content.
Cutting it in Claude Code
- Turn off what the project does not use. Toggle a server off in
/mcp, or type/mcp disable <server>. Claude Code records the choice per project and keeps the configuration. - Scope servers to the projects that need them.
claude mcp adddefaults to local scope, this project only.--scope userloads a server in every project on the machine, and a checked-in.mcp.jsonloads it for everyone who approves it in that repository. - Check your claude.ai connectors. Signed in with a claude.ai subscription, Claude Code automatically loads the connectors you added on claude.ai. Set
"disableClaudeAiConnectors": truein a project's.claude/settings.json, or start Claude Code withENABLE_CLAUDEAI_MCP_SERVERS=false, to leave them out. - Use the CLI when one exists. Anthropic's cost guidance notes that
gh,aws,gcloudandsentry-cliadd no per-tool listing at all, and Claude can run them directly. - Keep deferral on. Reserve
alwaysLoadfor the few tools needed on every turn. - Remove single tools you never want. A permission deny rule naming one tool, in the form
mcp__<server>__<tool>, takes it out of Claude's context.
A long tool list costs accuracy as well as tokens
A large catalogue has a second price. Anthropic's documentation places the point where Claude's tool selection starts to degrade at 30 to 50 available tools. Deferral helps, because Claude only sees the handful it searched for, but a smaller catalogue helps more.
If you write or choose servers, Anthropic's tool design guidance points the same way: consolidate related operations into one tool with an action parameter instead of one tool per action, prefix names by service (github_, slack_) so one search finds a whole group, and return only the fields the next step needs. Fewer, broader tools shrink the definitions. Lean responses shrink the part that stays in the transcript.
When a tool change breaks the cache
Prompt caching is a prefix match. On the Claude API the prefix is built in a fixed order, tools, then system, then messages, and changing any tool's name, description or parameters invalidates the entire cache. The next request processes the whole conversation again as new input instead of reading it from the cache, where it would have cost a tenth of the input price on most models as of September 2026.
In Claude Code this depends on the loading mode. With deferred tools, a server that connects, disconnects or changes its tool list only appends to the conversation, and the cached prefix survives. With tools loaded upfront, any such change invalidates it, and it can happen without you touching anything: a local server's process exits, a remote session expires, a server reconnects after a transient failure. Edits to the MCP configuration take effect at the next start, so when tools load upfront, the cheap moment to add or remove a server is between sessions.
Codex CLI
Codex CLI keeps MCP servers in ~/.codex/config.toml, or in .codex/config.toml inside a trusted project, one [mcp_servers.<name>] table per server. The CLI, the IDE extension and the ChatGPT desktop app share that configuration, so a server added for one is available in all three.
The bill has the same shape: every tool Codex exposes is described to the model, and every result joins the transcript. Current releases (0.156 as of September 2026) defer MCP definitions behind a search tool on the models that support it, listing the servers upfront. That behaviour lives in the open-source code rather than the documentation, so treat it as current, not permanent.
/mcp verboseIn the Codex TUI, lists the MCP tools this session can call, with server diagnostics.
/status shows the session's token usage and remaining context. The switches live in the same configuration file:
enabled = falseturns a server off without deleting its entry.enabled_toolsexposes only the tools you list, anddisabled_toolsremoves some, applied after the allow list.tools.<tool>.output_token_limit, inside a server's table, sets a token budget for one tool's output, overriding the model's default truncation.- Servers that one repository needs belong in its
.codex/config.toml, not in the global file.
The rest of the bill
With definitions deferred and idle servers off, what remains is the context itself: the files the agent reads, the command output, and the history that carries both. That is the part capsul works on. It drives the Claude Code or Codex CLI you are already signed into, sends what the task asks for rather than the repository, and stops at the budget you set and tells you what it left out.
capsul ask 'why does the webhook handler retry twice' --budget 3000Questions
Do MCP servers use tokens when you do not call them?
Yes. With deferred loading, the Claude Code default on supported models as of September 2026, an idle server costs its tool names and its instructions on every request. With deferral off, it costs the full name, description and JSON schema of every tool on every request. Turning the server off for the project is what brings it to zero.
How do I see how many tokens my MCP tools use in Claude Code?
Run /context all, which breaks the context window down by category and lists the tokens each loaded MCP tool uses. /mcp shows each server's status and tool count. On Pro, Max, Team and Enterprise plans, /usage also shows each MCP server's share of recent usage, based on the requests that carried its results.
How many MCP tools is too many for Claude Code?
Anthropic's documentation says Claude's tool selection starts to degrade past 30 to 50 available tools, and that five common servers can add about 55,000 tokens of definitions when everything loads upfront. Tool search, on by default in Claude Code, addresses both by loading only the few tools a request needs. The simplest rule still holds: connect in each project only the servers that project uses.
Does adding an MCP server mid-session break the prompt cache?
It does when tool definitions are loaded upfront, because tools sit at the start of the cached prefix and any change to them invalidates everything after. With deferred tools, the Claude Code default on supported models, a server connecting or disconnecting only appends content and the cache survives. Configuration edits apply at the next start, so changing servers between sessions costs nothing extra.
Where does Codex CLI configure MCP servers?
In ~/.codex/config.toml, one [mcp_servers.<name>] table per server, or in .codex/config.toml inside a trusted project. Set enabled = false to turn a server off without deleting it, and use enabled_tools or disabled_tools to expose only some of its tools. The Codex CLI, IDE extension and ChatGPT desktop app share this configuration.
$ npm i -g @penra/capsul