Follow Discord
Sweep 02 Oct 2026 · 18:55Z Build v2.1.288 509 read Stable v2.1.285 Latest v2.1.288 Next v2.1.288 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · claude-code

How Claude Code uses prompt caching changedprompt-caching

Nearest release: v2.1.284, published 6 hours before upstream edited the page. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Upstream edited this page at 28 Sep 2026 23:26 UTC, give or take a minute or two: the time comes from Anthropic’s own sitemap rather than from a commit. This site recorded the change at 28 Sep 2026 23:37 UTC.

Upstream edited
Recorded here
Lines+33added
Lines−36removed
From line 18 where the diff opens
First seen 14 Aug 2026 this site's first read of the page
Recorded edits38to this page, all time

The whole hunk

from line 18, old and new numbered
/
lines
from line 18
1818 
1919To get the most out of prefix matching, Claude Code orders each request so content that rarely changes between turns comes first:
2020 
21| Layer | Content | Changes when |
22| --------------- | ----------------------------------------------- | ----------------------------------------------- |
23| System prompt | Core instructions, tool definitions | The set of loaded tool definitions changes |
24| Project context | CLAUDE.md, auto memory, unscoped rules | Session starts, or after `/clear` or `/compact` |
25| Conversation | Your messages, Claude's responses, tool results | Every turn |
21| Layer | Content | Changes when |
22| - | - | - |
23| System prompt | Core instructions, tool definitions | The set of loaded tool definitions changes |
24| Project context | CLAUDE.md, auto memory, unscoped rules | Session starts, or after `/clear` or `/compact` |
25| Conversation | Your messages, Claude's responses, tool results | Every turn |
2626 
2727A change to the conversation layer leaves the system prompt and project context cached. A change to the system prompt invalidates everything, because all later content now sits behind a different prefix. The third column gives common triggers rather than an exhaustive list, and the sections below cover the full set.
2828 
from line 104
104104 
105105### Connecting or removing an MCP server
106106 
107Tool definitions sit in the system prompt layer, so the cache invalidates when the set of tool definitions in the request changes between turns. Toggling the [advisor tool](/docs/en/advisor) is an exception: its definition sits after the cache breakpoint, so enabling or disabling `/advisor` keeps the cached prefix intact. Whether an [MCP server](/docs/en/mcp) change does this depends on whether its tools are deferred by [tool search](/docs/en/mcp#scale-with-mcp-tool-search) or loaded into the prefix:
107Tool definitions sit in the system prompt layer, so the cache invalidates when the set of tool definitions in the request changes between turns. Toggling the [advisor tool](/docs/en/advisor) is an exception: its definition sits after the cache breakpoint, so enabling or disabling `/advisor` keeps the cached prefix intact. Whether an [MCP server](/docs/en/mcp) change does this depends on whether [tool search](/docs/en/mcp#scale-with-mcp-tool-search) defers the session's MCP tools, the default on supported models:
108108 
109* **Deferred tools**, the default on supported models: a server connecting, disconnecting, or changing its tool list only appends new content and doesn't disturb anything already cached.
110* **Tools loaded into the prefix**: adding a definition invalidates the cache, and so does removing one on purpose. This is the case when [tool search is unavailable or disabled](/docs/en/mcp#configure-tool-search), such as on Google Cloud's Agent Platform models earlier than the Claude 4.5 generation, with a custom `ANTHROPIC_BASE_URL` gateway, or on a Microsoft Foundry [deployment hosted on Azure](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry#hosting-options) once Claude Code detects that the deployment rejects tool search.
109* **Tools deferred**: Claude Code keeps the tool list from the conversation's first request for the whole conversation, so a server connecting or disconnecting mid-session doesn't disturb anything already cached. A server that finishes connecting after the first request supplies its tools as deferred definitions that Claude loads on demand.
110* **Tools loaded upfront**: adding a definition invalidates the cache, and so does removing one on purpose. This applies when tool search is [below its `auto` threshold, disabled, or unavailable](/docs/en/mcp#configure-tool-search), such as on Google Cloud's Agent Platform models earlier than the Claude 4.5 generation, with a custom `ANTHROPIC_BASE_URL` gateway, or on a Microsoft Foundry [deployment hosted on Azure](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry#hosting-options) once Claude Code detects that the deployment rejects tool search.
111111 
112112Without tool search, whether a mid-session server change invalidates the cache depends on what changed. For each change, this table gives whether the cache is kept and what happens to the tool definitions in the next request.
113113 
114| Mid-session change | Cache | Tool definitions in the next request |
115| ---------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
116| A server connects, or a [dynamic tool update](/docs/en/mcp#dynamic-tool-updates) adds tools | Invalidated | The new definitions are added |
117| A server drops out with no action on your part, such as a stdio server's process exiting | Kept | The server's definitions stay unchanged. A call to one of its tools returns an error instead of running |
118| A remote server [reconnects automatically](/docs/en/mcp#automatic-reconnection) after its connection drops | Kept, unless a request sent while the server reconnects adds the `WaitForMcpServers` tool, which invalidates the cache once | The server's definitions stay unchanged. A request sent while the server reconnects can add `WaitForMcpServers` when the conversation hasn't listed it yet, and the tool then stays listed for the rest of the conversation |
119| You remove a tool on purpose, such as with a [deny rule](#denying-an-entire-tool) or by disabling its server in `/mcp` | Invalidated | The definition is removed |
114| Mid-session change | Cache | Tool definitions in the next request |
115| - | - | - |
116| A server connects, or a [dynamic tool update](/docs/en/mcp#dynamic-tool-updates) adds tools | Invalidated | The new definitions are added |
117| A server drops out with no action on your part, such as a stdio server's process exiting | Kept | The server's definitions stay unchanged. A call to one of its tools returns an error instead of running |
118| A remote server [reconnects automatically](/docs/en/mcp#automatic-reconnection) after its connection drops | Kept, unless a request sent while the server reconnects adds the `WaitForMcpServers` tool, which invalidates the cache once | The server's definitions stay unchanged. A request sent while the server reconnects can add `WaitForMcpServers` when the conversation hasn't listed it yet, and the tool then stays listed for the rest of the conversation |
119| You remove a tool on purpose, such as with a [deny rule](#denying-an-entire-tool) or by disabling its server in `/mcp` | Invalidated | The definition is removed |
120120 
121121When you resume a conversation whose tools load into the prefix, one of its MCP servers can still be connecting as the first request goes out. If the transcript recorded that server's tool definitions, that request includes them as recorded, so it doesn't change when the server finishes connecting with the same tools.
122122 
from line 132
132132 
133133#### Plugins that provide MCP servers
134134 
135When you enable or disable a plugin that provides [MCP servers](/docs/en/plugins/components#mcp-servers), Claude Code follows the same rules as when you [connect or remove an MCP server](#connecting-or-removing-an-mcp-server):
135When you enable or disable a plugin that provides [MCP servers](/docs/en/plugins/components#mcp-servers), Claude Code follows the same rules as when you [connect or remove an MCP server](#connecting-or-removing-an-mcp-server).
136136 
137* If Claude Code defers the server's tools, it keeps the cache.
138* If Claude Code loads them into the prefix, the next request re-reads the entire conversation.
139 
140137#### Code intelligence plugins
141138 
142139When you enable a [code intelligence plugin](/docs/en/plugins/code-intelligence), Claude gets the [LSP tool](/docs/en/tools-reference#lsp-tool-behavior).
from line 263
266263 
267264Unless you choose a TTL yourself, Claude Code requests the one-hour TTL only on a Claude subscription within your plan's included usage. There it requests the hour for the main conversation, plus a small set of helper requests that Anthropic controls server-side. This table gives each bucket's default TTL under both kinds of billing.
268265 
269| Request bucket | Claude subscription, within plan usage | Usage credits, API key, or cloud provider |
270| ----------------- | ------------------------------------------------------------------------------ | ----------------------------------------- |
271| Main conversation | One hour | Five minutes |
272| Everything else | Five minutes, except the server-controlled helper requests, which get one hour | Five minutes |
266| Request bucket | Claude subscription, within plan usage | Usage credits, API key, or cloud provider |
267| - | - | - |
268| Main conversation | One hour | Five minutes |
269| Everything else | Five minutes, except the server-controlled helper requests, which get one hour | Five minutes |
273270 
274271Once you go over your plan's usage limit and Claude Code draws on [usage credits](https://support.claude.com/en/articles/12429409-extra-usage-for-paid-claude-plans), you are billed for that usage, so Claude Code drops the main conversation to the cheaper five-minute TTL. To keep the one-hour TTL there, [choose the TTL yourself](#choose-the-ttl-yourself).
275272 
from line 296
299296 
300297## Cache scope
301298 
302In Claude Code, the cache is effectively scoped to one machine and directory. Each conversation carries the working directory, platform, shell, and OS version, and the system prompt names your auto memory paths, so two sessions in different directories build different prefixes and miss each other's cache. That includes worktrees of the same repository, since each worktree has its own working directory.
299In Claude Code, the cache is effectively scoped to one machine and directory. The system prompt embeds your auto memory paths, and the conversation opens with an announcement of the working directory, platform, shell, and OS version. Two sessions in different directories therefore build different prefixes and miss each other's cache.
303300 
304301Sessions you run in parallel in the same directory build matching prefixes and read each other's cache. Sequential sessions share the prefix only when the git status snapshot taken at startup matches, since each conversation also carries the branch and recent commits from that snapshot.
305302 
306The underlying API cache is broader. Caches are isolated between organizations, and on some providers, [between workspaces within an organization](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#cache-storage-and-sharing). Within those boundaries, any two requests with the same model and prefix read the same cache. For Agent SDK callers running fleets of automated processes, see [improve prompt caching across users and machines](/docs/en/agent-sdk/modifying-system-prompts#improve-prompt-caching-across-users-and-machines) to suppress the per-machine sections of the system prompt and share the cache across machines.
303The underlying API cache is broader. Caches are isolated between organizations, and on some providers, [between workspaces within an organization](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#cache-storage-and-sharing). Within those boundaries, any two requests with the same model and prefix read the same cache. For Agent SDK callers running fleets of automated processes, see [improve prompt caching across users and machines](/docs/en/agent-sdk/modifying-system-prompts#improve-prompt-caching-across-users-and-machines) to move the auto memory location out of the system prompt and share the system prompt's cache entry across users and machines.
307304 
308305## Check cache performance
309306 
310307Cache performance shows up as two token counts the API reports on every response. The most direct way to watch them live is a [statusline script](/docs/en/statusline) that reads the `current_usage` object:
311308 
312| Field | Meaning |
313| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
314| `cache_creation_input_tokens` | Tokens written to the cache on this turn, billed at the cache write rate |
315| `cache_read_input_tokens` | Tokens served from cache on this turn, billed at the model's [cached token rate](https://platform.claude.com/docs/en/about-claude/pricing), below the standard input rate |
309| Field | Meaning |
310| - | - |
311| `cache_creation_input_tokens` | Tokens written to the cache on this turn, billed at the cache write rate |
312| `cache_read_input_tokens` | Tokens served from cache on this turn, billed at the model's [cached token rate](https://platform.claude.com/docs/en/about-claude/pricing), below the standard input rate |
316313 
317314A high read-to-creation ratio means caching is working well. If creation stays high turn after turn, something is changing in your prefix. The [actions that invalidate the cache](#actions-that-invalidate-the-cache) section lists the usual causes.
318315 
from line 338
341338 
342339Disabling caching is occasionally useful when debugging caching behavior with a specific model or provider. To turn it off, set one of these environment variables to `1`:
343340 
344| Variable | Effect |
345| ------------------------------- | ----------------------------------- |
346| `DISABLE_PROMPT_CACHING` | Disable for all models |
347| `DISABLE_PROMPT_CACHING_HAIKU` | Disable for the default Haiku model |
348| `DISABLE_PROMPT_CACHING_SONNET` | Disable for Sonnet only |
349| `DISABLE_PROMPT_CACHING_OPUS` | Disable for Opus only |
350| `DISABLE_PROMPT_CACHING_FABLE` | Disable for Fable only |
341| Variable | Effect |
342| - | - |
343| `DISABLE_PROMPT_CACHING` | Disable for all models |
344| `DISABLE_PROMPT_CACHING_HAIKU` | Disable for the default Haiku model |
345| `DISABLE_PROMPT_CACHING_SONNET` | Disable for Sonnet only |
346| `DISABLE_PROMPT_CACHING_OPUS` | Disable for Opus only |
347| `DISABLE_PROMPT_CACHING_FABLE` | Disable for Fable only |
351348 
352349`DISABLE_PROMPT_CACHING_HAIKU` applies to the default Haiku model, the model the `haiku` alias resolves to. It disables caching wherever that model runs, including the main conversation when it is your main model. Covering the main conversation requires Claude Code v2.1.283 or later.
353350 
Feedback