veasy

Claude Code · AI agents · prompt caching

How Claude Code is built around prompt caching

Thariq from the Claude Code team wrote about designing the whole harness around prompt caching. His main points, rewritten: how to order the prompt, why the model and the tools stay fixed for a session, how plan mode and tool search avoid breaking the cache, and how compaction reuses it.

Diagram: the four layers of a Claude Code request, from top to bottom: system prompt and tools, CLAUDE.md and memory, session state, messages. Changing a layer means everything below it is billed again; three changes that broke the order are listed beside it.

The Claude Code team at Anthropic has an alert on its prompt cache hit rate, and when the rate falls too low they declare a SEV, an incident. Thariq (@trq212), who works on Claude Code, explained why in a post on X in February 2026, "Lessons from Building Claude Code: Prompt Caching Is Everything". He covers how they order the prompt, the times they broke the cache themselves, and the features they designed so the cache survives.

What prompt caching does for an agent

On every turn an agent sends the model almost the whole conversation again: the system prompt, the tool definitions, every message so far. With prompt caching the API reuses the work it already did on the earlier part of the request instead of processing it from scratch, so each turn is cheaper and faster. According to a diagram in his post, reading tokens from the cache costs about a tenth of the normal input price.

The match works on the prefix. The API caches the request from its first token up to each cache_control marker, and a later request only gets a hit if its beginning is identical. One difference near the top and everything after it is computed again.

A high hit rate lowers the cost of running Claude Code and, Thariq says, lets the team offer more generous rate limits on the subscription plans. That is why it has an alert.

Order the prompt from stable to changing

Since the cache matches a prefix, the order of the parts decides how many requests can share it. Claude Code puts them in this order:

  1. The base system prompt and the tool definitions, the same for every user and cached globally.
  2. CLAUDE.md and memory, the same within one project.
  3. Session context, such as the environment and the MCP servers, the same within one session.
  4. The messages, which grow every turn.

Thariq says this order is more fragile than it looks, and lists ways his team broke it: a detailed timestamp in the static system prompt, tool definitions sorted differently from one call to the next, and tool parameters that changed during a session, for example the list of agents the AgentTool can call. None of these looks like a big change, yet each one turned the rest of the request into a cache miss.

Send updates as messages

Information in the prompt goes stale. The date changes, or the user edits a file. Fixing the system prompt feels like the right move, but it means a cache miss on the whole conversation and a bigger bill for the user.

Claude Code puts the update in the conversation instead: the next user message or tool result gets a <system-reminder> tag with the new fact, for example that it is now Wednesday. The old prefix stays the same, so the cache still hits.

Keep the same model for the whole session

Each model has its own cache. His example: you are 100,000 tokens into a conversation with Opus and have an easy question. Switching to Haiku for it costs more than letting Opus answer, because Haiku first has to build a cache for all 100,000 tokens.

When a task should go to another model, Claude Code uses a subagent. The main model writes a short handoff message with only what the other model needs. The Explore agents in Claude Code run on Haiku this way.

Keep the same tools for the whole session

Thariq counts changing the tool set mid-conversation among the most common ways people break the cache. Giving the model only the tools it needs right now sounds sensible. The tool definitions sit in the prefix, though, so adding or removing one throws away the cache for the entire conversation.

Two Claude Code features are built around that limit.

Plan mode. The obvious design would switch to read-only tools while plan mode is on. Claude Code keeps every tool in every request and makes entering and leaving plan mode tools of their own, EnterPlanMode and ExitPlanMode. When the user turns plan mode on, the agent gets a system message saying it is in plan mode and what that means: read the code, don't edit files, call ExitPlanMode when the plan is ready. The tool definitions stay exactly as they were. Because EnterPlanMode is a tool, the model can also call it on its own when it runs into a hard problem, and that doesn't break the cache either.

Tool search. A session can have dozens of MCP tools, and sending all their definitions with every request is expensive. Instead of dropping some, Claude Code sends a short stub for each one, only the name with defer_loading: true. When the model needs a tool, it finds it with the ToolSearch tool, and the full schema is loaded then. The stubs are always the same and in the same order, so the prefix doesn't move. Tool search is also available in the API for anyone building their own agent.

Compaction that reuses the cache

When a conversation is about to fill the context window, Claude Code compacts it: it summarizes everything so far and carries on from the summary.

The easy implementation is a separate API call: a different system prompt, no tools, and the whole conversation passed in to be summarized. That call shares no prefix with the main conversation, so every input token is billed at full price, right when the conversation is at its longest.

Claude Code forks the conversation instead. The compaction request uses the same system prompt, user context, system context and tool definitions as the parent conversation, keeps all of its messages, and adds the summary request as one new user message at the end. To the API it looks almost the same as the parent's last request, so the cache hits and the only new tokens are the summary instruction.

For this to work, Claude Code keeps part of the context window free, a compaction buffer, for the summary request and the summary it produces. Anthropic built compaction into the API based on this work, so you don't have to rebuild it yourself.

His lessons, shortened

Thariq closes with five lessons:

  • The cache is a prefix match. A change near the top invalidates everything after it, so get the order right first.
  • Put changing facts, such as the date or the current mode, in messages, not in the system prompt.
  • Keep the model and the tool set fixed for the session. Model state changes as tools, like plan mode, and defer loading tools rather than removing them.
  • Watch the cache hit rate the way you watch uptime. The Claude Code team treats a cache break as an incident, because a few more points of cache misses change cost and latency a lot.
  • Side jobs such as compaction, summaries or running a skill should use the same parameters as the parent conversation, so they hit its cache.

Source

ShareFacebookLinkedInX

Comments