What a CLAUDE.md costs in tokens, and how long it should be

Your CLAUDE.md is read before your first prompt and sent again with every request after it. Caching makes that cheaper, not free, so each line has to earn its place.

A CLAUDE.md is a Markdown file of standing instructions that Claude Code loads into the context window when a session starts. It is not part of the system prompt: Claude Code delivers it as a user message placed right after it, and Claude treats it as guidance, not enforced configuration.

The model keeps nothing between requests, so Claude Code re-sends the whole context on every call: system prompt, project context, every earlier message and tool result. Each round of tool use is a new call, so one prompt that makes Claude read five files and run the tests is several requests, each carrying your CLAUDE.md in full.

In short: every line is paid on every request, at a discount while the cache holds and above the normal input price each time the cache is rebuilt. Keep the file short and move what is not needed every time into files that load on demand.

Which files load, and when

Claude Code concatenates every memory file it finds instead of letting one replace another. As of September 2026, the Claude Code documentation lists these sources:

  • Managed policy: an organisation-wide CLAUDE.md deployed by IT, such as /etc/claude-code/CLAUDE.md on Linux. It cannot be excluded.
  • User: ~/.claude/CLAUDE.md and the rules in ~/.claude/rules/, applied to every project on your machine.
  • Project: ./CLAUDE.md or ./.claude/CLAUDE.md, plus every .claude/rules/ file without a paths field, shared through version control.
  • Local: ./CLAUDE.local.md, your own notes for one project, added to .gitignore.
  • Auto memory: the first 200 lines or 25KB of the MEMORY.md index Claude keeps for the repository. The topic files it points to are read only when needed.

Files in the working directory and the directories above it load at launch. A subdirectory's CLAUDE.md loads only when Claude reads a file in that directory, and a rule with paths only when Claude reads a file it covers. Imports are the trap: an @path/to/file line pulls that file in at launch, up to four hops deep, so splitting a long file into imports tidies it without saving a single token.

Why caching does not make it free

With prompt caching, the API compares the start of each request with what it processed recently and bills the identical part at a reduced rate. Claude Code orders every request so that what rarely changes comes first: the system prompt and tool definitions, then the project context (your memory files), then the conversation. On an ordinary turn, your CLAUDE.md is a cache read.

At Anthropic's published API prices (September 2026), a cache read costs a tenth of the base input price on most Claude models, and less on a few of the newest. A cache write costs 1.25 times the base price with a five-minute lifetime, 2 times with a one-hour lifetime. Whenever the prefix has to be written again, your CLAUDE.md is billed above the price of plain input. That happens more often than it looks:

  • After a break longer than the cache lifetime: one hour for the main conversation on a Claude subscription within its included usage, five minutes by default with an API key.
  • When you switch model (each model has its own cache) or, on most models, change the effort level.
  • When the tool definitions ahead of it change, for instance an MCP server whose tools load upfront connecting mid-session.
  • In every subagent except the built-in Explore and Plan: each loads your CLAUDE.md files into its own context and builds its own cache, with a five-minute lifetime by default even on a subscription.
  • In each worktree or directory you start a session from, because the cache is effectively scoped to one directory.

Caching changes the price, not the size. The file occupies the window on every request, which leaves less room for the conversation before auto-compaction, and longer files reduce adherence, because each rule gets less attention. On a subscription there is no invoice, but the same context draws on your plan's usage limits.

How to see what yours costs

Open a fresh session and, before typing anything, look at what is already in the window:

/context

Breaks the window down by category. The Memory files list shows every CLAUDE.md, rules and auto memory file that loaded.

/memory lists your memory files and opens any of them in your editor. From Claude Code v2.1.251, /usage adds a Prompt cache (main) line with the share of input served from cache and the number of misses; a growing miss count means the prefix, memory files included, keeps being rewritten.

What belongs in it

Anthropic's test for each line: would removing it cause Claude to make mistakes? If not, cut it. What survives is usually short:

  • Build, test and lint commands Claude cannot guess.
  • Code style rules that differ from the defaults, not the ones your formatter already enforces.
  • Repository etiquette: branch names, commit and pull request conventions.
  • Architectural decisions specific to the project.
  • Environment quirks, such as a required variable or a service that must be running.
  • Pitfalls: the non-obvious behaviour that has already cost someone an afternoon.

Write each rule so it can be checked: "Run npm test before committing" works, "test your changes" does not. The documentation's target is under 200 lines per file.

What to move out, and where

Most of a bloated CLAUDE.md is not wrong, it is in the wrong place. Each kind of content has a cheaper home:

  • What Claude can read from the code (directory layouts, dependency lists, file-by-file descriptions): delete it.
  • Long reference material, such as an API reference or a style guide: keep it in its own file and mention the path without the @. Claude then opens it only when a task calls for it.
  • Multi-step procedures, such as a release checklist: make them skills. Only a skill's short description loads at session start; the rest loads when it is used.
  • Rules for one part of the codebase: a .claude/rules/ file with a paths pattern, or a CLAUDE.md in that subdirectory. Both load only when Claude reads a file they apply to.
  • Anything that must happen every time, such as formatting after an edit: a hook. A hook costs no context unless it returns output, and unlike an instruction it is enforced.

How to trim it

Start with the checkup that ships with Claude Code. From v2.1.206, /doctor proposes cuts for a checked-in CLAUDE.md: it removes content Claude can derive from the codebase, keeps pitfalls, rationale and conventions that differ from tool defaults, and moves the always-loaded guidance that remains into skills and nested files that load on demand.

/doctor

Reports its findings first and asks before changing any file.

Then a few passes by hand:

  • Delete each rule Claude follows unprompted: remove the line and check whether behaviour changes.
  • Turn notes meant for maintainers into block-level HTML comments. Claude Code strips <!-- ... --> blocks before the content reaches the model.
  • Resolve contradictions. Two rules that disagree are paid for twice, and Claude may follow either one.
  • In a monorepo, skip other teams' files with the claudeMdExcludes setting.
  • Give custom subagents that need nothing from it omitClaudeMd: true in their definition (v2.1.271 or later).

AGENTS.md and Codex CLI

Codex CLI reads AGENTS.md instead. According to its documentation (September 2026), it builds an instruction chain once per run, usually once per session in the interactive interface:

  • Global: in ~/.codex, or CODEX_HOME if set, AGENTS.override.md if it exists, otherwise AGENTS.md.
  • Project: from the project root, usually the Git root, down to the directory you launched from, at most one file per directory: AGENTS.override.md, then AGENTS.md, then any name listed in project_doc_fallback_filenames.
  • It joins them from the root down, skips empty files, and stops adding text once the combined size reaches project_doc_max_bytes, 32 KiB by default.

The result goes into the first turn of the session, so every later request carries it. Codex's own tips for making usage limits last include shrinking AGENTS.md and nesting files in the directories they govern. As of September 2026, its credit rate card bills cached input at a tenth of fresh input, with no separate cache-write charge. To see the instructions as the model receives them:

codex debug prompt-input

Prints the model-visible input as JSON instead of sending it, instruction files included.

One file can serve both agents. From v2.1.277, Claude Code reads AGENTS.md as project instructions when there is no CLAUDE.md or CLAUDE.local.md in the working directory or above it; otherwise, put @AGENTS.md at the top of your CLAUDE.md. Claude Code does not read AGENTS.override.md.

The fixed cost and the variable one

Trimming the instruction file shrinks the fixed part of the bill. The rest is what each task reads on top of it: files, command output, history. Naming the file you mean and starting a new session per task keep that part down.

For a hard ceiling on it, capsul drives the Claude Code or Codex CLI you are already signed into, sends what the task asks for under a token budget you set, and tells you what it left out. Measured results are on the benchmark page.

Questions

How long should a CLAUDE.md be?

Anthropic's Claude Code documentation recommends under 200 lines per file (as of September 2026), because longer files take more context and reduce how reliably Claude follows them. There is no hard line limit: Claude Code loads a file of up to 4 MiB in full and skips a larger one. The practical test is whether removing a line would cause Claude to make mistakes; if it would not, cut it.

Does CLAUDE.md count against my Claude usage limits?

Yes. It is part of the input of every request in a session, so it draws on a subscription's usage limits, or on an API bill, like the rest of the context. Prompt caching makes each read cheaper while the cache is warm, but every time the cache is rebuilt the file is written to it again, which on the API costs more than plain input.

Do @imports in CLAUDE.md reduce token usage?

No. Imported files are expanded and loaded at launch together with the CLAUDE.md that references them, up to four hops deep, so they cost the same as pasting the text in. To keep a document out of context until a task needs it, mention its path without the @, or move it into a skill or a path-scoped rule.

Does Codex CLI read CLAUDE.md, and does Claude Code read AGENTS.md?

Codex reads AGENTS.md or AGENTS.override.md, from ~/.codex and from the project root down to your current directory, and reads other names only if you list them in project_doc_fallback_filenames. Claude Code v2.1.277 and later reads AGENTS.md when there is no CLAUDE.md or CLAUDE.local.md in your working directory or above it, and otherwise through an @AGENTS.md import in your CLAUDE.md. One short AGENTS.md can therefore serve both agents.

Do edits to CLAUDE.md apply to the current session?

Not for the project and user files. Claude Code reads them once at session start and keeps that version, so an edit takes effect after /clear, /compact or a restart, and it does not break the cache in the meantime. A nested CLAUDE.md or a path-scoped rule that has not loaded yet does pick up the edit when Claude first reads a file it applies to.

$ npm i -g @penra/capsul

← All guides