The whole hunk
from line 6, old and new numbered
/
lines
from line 6
66
77Preserved thinking is a property of newer Claude models that guards against distillation. It decides whether the model can use a thinking block that you send back from an earlier turn. Starting with Claude Fable 5.1, when a `thinking` or `redacted_thinking` block comes back in a request, the API checks the block's `signature` for two things:
88
9* **The model is the one that produced the block, or a newer one.** A model reads its own thinking blocks and those of earlier models. Claude Fable 5.1 reads blocks from Claude Opus 5, but Claude Opus 5 can't read blocks from Claude Fable 5.1. If the current model can't read a block, the API drops it from that request without an error. See [Switching models mid-conversation](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models).
9* **The model can read the block.** Each model reads its own thinking blocks and those of a fixed set of other models. Claude Fable 5.1 reads blocks from Claude Opus 5 and, on the Claude API, from Claude Opus 5.5; neither Claude Opus 5 nor Claude Opus 5.5 reads blocks from Claude Fable 5.1. If the current model can't read a block, the API drops it from that request without an error. See [Switching models mid-conversation](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models).
1010* **Nothing before the thinking block has changed.** The top-level `system` prompt, `tools`, and `messages` before the block are its prefix. If the prefix differs from what you sent when the block was produced, that block and every later thinking block are invalid, and the API rejects the request with a 400 error or drops the invalid blocks, whichever you choose. See [Keeping the prefix unchanged](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#prefix-check).
1111
12The model check applies to every account. The API enforces the prefix check by default for accounts created on or after August 31, 2026, 00:00 UTC. On older accounts, it enforces the prefix check only on requests that set `thinking.block_binding.prefix_mismatch_behavior`. **Later models will enforce the prefix check for all accounts**, so make your integration append-only now.
12The model check applies to every account. The API enforces the prefix check by default for accounts created on or after August 31, 2026, 00:00 UTC. On older accounts, it enforces the prefix check only on requests that set `thinking.block_binding.prefix_mismatch_behavior`. **Make your integration append-only regardless of your account's age**, so the same code works on every account, including newer accounts enforced by default.
1313
1414## Who needs to change anything
1515
from line 32
3232
3333Claude Fable 5.1 and Claude Mythos 5.1 read thinking blocks produced by each other and by earlier Claude models. No earlier model reads thinking blocks from Claude Fable 5.1 or Claude Mythos 5.1.
3434
35* **A conversation that moves up to Claude Fable 5.1 keeps its reasoning.** The earlier model's thinking blocks stay readable, so the model thinks as usual from the first turn after the switch.
35Claude Opus 5.5 reads thinking blocks from Claude Opus 5 and earlier Opus, Sonnet, and Haiku models, but not from Claude Fable or Claude Mythos models. On the Claude API, Claude Fable 5.1 and Claude Mythos 5.1 read thinking blocks from Claude Opus 5.5; no other model does. So a conversation that moves from Claude Opus 5 onto Claude Opus 5.5 keeps its reasoning, and so does one that moves from Claude Opus 5.5 up to Claude Fable 5.1 or Claude Mythos 5.1 on the Claude API. One that moves from Claude Fable 5.1 or Claude Mythos 5.1 to Claude Opus 5.5, or from Claude Opus 5.5 to any model other than those two, runs the turns after the switch without the previous model's reasoning. The blocks are dropped, not rejected, as described below.
36
37* **A conversation that moves to Claude Fable 5.1 from an earlier model, or from Claude Opus 5.5 on the Claude API, keeps its reasoning.** The earlier model's thinking blocks stay readable, so the model thinks as usual from the first turn after the switch.
3638* **A conversation that moves down to an earlier model loses Claude Fable 5.1's reasoning for that request.** This happens when a router sends a turn to a cheaper model, after a [classifier refusal fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback), or during a [server-side fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#server-side-fallback). The API removes the unreadable blocks before the prompt reaches the model. They aren't billed and don't count toward `input_tokens`.
3739
3840Keep sending the full history on every request, thinking blocks included, and let the API drop what the current model can't read. The API never edits your `messages` array, so the dropped blocks stay in your history. When the same history goes back to Claude Fable 5.1, its blocks are readable again, along with the earlier model's thinking. The reasoning is lost for good only if your client removes the blocks itself, for example a harness that strips thinking on a model switch or rebuilds the history from what each model used.
from line 59
5759
5860## Keeping the prefix unchanged
5961
60On Claude Fable 5.1, a thinking block stays valid only while everything you sent before it is unchanged on later requests. The checked prefix has three parts:
62On Claude Fable 5.1 and Claude Opus 5.5, a thinking block stays valid only while everything you sent before it is unchanged on later requests. The checked prefix has three parts:
6163
6264* The top-level `system` prompt
6365* The set of `tools`
from line 378
376378
377379### When the API enforces the check
378380
379The API enforces the prefix check on Claude Fable 5.1 for new accounts.
381The API enforces the prefix check on Claude Fable 5.1 and Claude Opus 5.5 for new accounts.
380382
381* **Accounts created on or after August 31, 2026, 00:00 UTC:** the API checks Claude Fable 5.1 requests and applies `"error"` unless you set `"drop_block"`. The same definition of a new account applies to the Claude API and to cloud platforms.
383* **Accounts created on or after August 31, 2026, 00:00 UTC:** the API checks Claude Fable 5.1 and Claude Opus 5.5 requests and applies `"error"` unless you set `"drop_block"`. The same definition of a new account applies to the Claude API and to cloud platforms.
382384* **Older accounts:** the API enforces the check only on requests that set `prefix_mismatch_behavior`. Setting the field opts a request in, so you can see what a new account sees without creating one. On requests that leave it unset, the API still runs the check but lets failing blocks through to the model. With the beta header, the response lists each one in `input_transformations` as `thinking_mismatch_allowed`, so you can find prefix edits without changing what the model receives.
383* **Later models:** every account, on every request.
384385
385386To find out which group your account is in, take a Claude Fable 5.1 conversation that contains a thinking block, change something before that block, and send it to Claude Fable 5.1 without the beta header or the `block_binding` field. A 400 response that names the header means your account is enforced by default. A 200 response means it isn't. To confirm, send the same request again with the beta header, still without `block_binding`: the response lists every thinking block after your edit in `input_transformations` as `thinking_mismatch_allowed`.
386387
from line 1205
12041205
12051206Each common prefix edit has a replacement that gives the model the same information and leaves earlier bytes unchanged, so later thinking stays valid. Find the edit your code makes today in the first column:
12061207
1207| Instead of | Use | Beta header |
1208| --------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------- |
1209| Rebuilding the top-level `system` prompt | A [mid-conversation system message](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#new-instructions) | None |
1210| Re-rendering the context in your first user message (environment, date, memory, project instructions) on each request | Render it once and resend it unchanged. When something changes, [put the new version in the newest turn](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#changing-context) | None |
1211| Clearing or shortening old `tool_result` content, or re-encoding old images, in place | Shorten a tool result or downscale an image before the first time you send it, not after. To clear old results later, [trim context on the server](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#server-side-trimming) with `clear_tool_uses_20250919` | `context-management-2025-06-27` |
1212| Injecting a reminder and deleting it on the next request | A [turn-scoped system message](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#per-turn-reminders) (`clear_at: "next_user_message"`) | `mid-conversation-system-clear-at-2026-08-21` |
1213| Adding or removing entries in `tools` | [`tool_addition` and `tool_removal` blocks](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#tool-changes) | `mid-conversation-tool-changes-2026-07-01` |
1214| Changing top-level `output_config.effort` (restarts the cache, doesn't affect thinking) | A [per-message `output_config`](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#effort-changes) | `mid-conversation-output-config-2026-07-01` |
1215| Dropping or summarizing old turns on the client | [On-demand compaction](https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand) to keep the recent turns with their thinking, other server-side [compaction or context editing](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#server-side-trimming), or [client-side compaction](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#custom-compaction-on-the-client) that keeps no stale thinking | `compact-2026-09-04` |
1216| An image or document URL whose bytes change between requests | A [`file_id` from the Files API](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#files-by-id), or base64 | None |
1208| Instead of | Use | Beta header |
1209| --------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
1210| Rebuilding the top-level `system` prompt | A [mid-conversation system message](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#new-instructions) | None |
1211| Re-rendering the context in your first user message (environment, date, memory, project instructions) on each request | Render it once and resend it unchanged. When something changes, [put the new version in the newest turn](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#changing-context) | None |
1212| Clearing or shortening old `tool_result` content, or re-encoding old images, in place | Shorten a tool result or downscale an image before the first time you send it, not after. To clear old results later, [trim context on the server](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#server-side-trimming) with `clear_tool_uses_20250919` | `context-management-2025-06-27` |
1213| Injecting a reminder and deleting it on the next request | A [turn-scoped system message](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#per-turn-reminders) (`clear_at: "next_user_message"`) | `mid-conversation-system-clear-at-2026-08-21` |
1214| Adding or removing entries in `tools` | [`tool_addition` and `tool_removal` blocks](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#tool-changes) | `inline-tools-2026-09-15` (add `mcp-client-2026-09-15` when the tool comes from an MCP server connected through the MCP connector), or the older `mid-conversation-tool-changes-2026-07-01`, which works on the Claude API, Amazon Bedrock, and Google Cloud |
1215| Changing top-level `output_config.effort` (restarts the cache, doesn't affect thinking) | A [per-message `output_config`](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#effort-changes) | `mid-conversation-output-config-2026-07-01` |
1216| Dropping or summarizing old turns on the client | [On-demand compaction](https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand) to keep the recent turns with their thinking, other server-side [compaction or context editing](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#server-side-trimming), or [client-side compaction](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#custom-compaction-on-the-client) that keeps no stale thinking | `compact-2026-09-04` |
1217| An image or document URL whose bytes change between requests | A [`file_id` from the Files API](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#files-by-id), or base64 | None |
12171218
12181219All of these assume you [send assistant turns back exactly as returned](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#append-assistant-turns-exactly-as-returned). Mid-conversation system messages, turn-scoped system messages, and tool changes aren't available on every model: [Mid-conversation system messages and tool changes](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages) lists the models that accept them. If your code serves several models, keep editing the top-level `system` prompt for the models that don't accept them.
12191220
from line 1221
12201221To use several betas in one request, combine the values in one `anthropic-beta` header. Beta names are the same on Amazon Bedrock and Google Cloud wherever the beta is available there (see [Beta headers](https://platform.claude.com/docs/en/api/beta-headers)):
12211222
12221223```text wrap
1223anthropic-beta: thinking-binding-controls-2026-08-01,mid-conversation-system-clear-at-2026-08-21,mid-conversation-tool-changes-2026-07-01
1224anthropic-beta: thinking-binding-controls-2026-08-01,mid-conversation-system-clear-at-2026-08-21,inline-tools-2026-09-15
12241225```
12251226
12261227### Send assistant turns back exactly as returned
from line 1318
13171318
13181319### Add or remove tools with `tool_addition` and `tool_removal`
13191320
1320Editing the `tools` array mid-session invalidates preserved thinking blocks. Instead, declare every tool the session might need in `tools` on the first request and never change the array. To change which tools the model can use from some point on, append a `role: "system"` message that carries a `tool_removal` or `tool_addition` block. These are [mid-conversation tool changes](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#mid-conversation-tool-changes) and need the beta header `mid-conversation-tool-changes-2026-07-01`. For example, to withdraw a dangerous tool after a mode switch:
1321Editing the `tools` array mid-session invalidates preserved thinking blocks. Leave the array as you first sent it, and change which tools the model can use by appending a `role: "system"` message that carries `tool_addition` or `tool_removal` blocks. These are [mid-conversation tool changes](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#mid-conversation-tool-changes) and need the beta header `inline-tools-2026-09-15`, which is available on the Claude API. The older `mid-conversation-tool-changes-2026-07-01` header still works for changes that name a tool by reference, on the Claude API, Amazon Bedrock, and Google Cloud.
13211322
1323You have two ways to use these blocks:
1324
1325* **Declare every tool up front.** Put every tool the session might need in `tools` on the first request, with `defer_loading: true` on any the model shouldn't see yet. Then turn tools on and off with `tool_addition` and `tool_removal` blocks that name them.
1326* **Start with a snapshot and add tools as you go.** Put the tools you know about in `tools` on the first request. When a new tool comes along, define it inside a `tool_addition` block instead of editing `tools`.
1327
1328Either way, `tools` never changes, so earlier thinking stays valid and the prompt cache still hits, with the one exception noted below.
1329
1330For example, to withdraw a dangerous tool after a mode switch:
1331
13221332```json
13231333{
13241334 "role": "system",
from line 1339
13291339}
13301340```
13311341
1332To offer a tool later instead, declare it in `tools` with `defer_loading: true` so the model doesn't see it at first. When it becomes available, append a `tool_addition` block:
1342To turn on a tool you declared with `defer_loading: true`, append a `tool_addition` block that names it:
13331343
13341344```json
13351345{
from line 1351
13411351}
13421352```
13431353
1344Sometimes you can't declare a tool up front because you don't know its schema yet. An MCP server discovered at runtime is the common case. Append that tool to `tools` with `defer_loading: true`, then offer it with a `tool_addition` block. Adding a deferred tool is safe: the prefix check ignores a deferred tool until a `tool_addition` block references it, so earlier thinking stays valid. Adding a tool without `defer_loading: true` changes the prefix and invalidates earlier thinking.
1354Sometimes you can't declare a tool up front because you don't know its schema yet: a tool your application discovers at runtime, or an MCP server that connects after the first turn. Define it inside the `tool_addition` block instead of touching `tools`. With `inline-tools-2026-09-15`, the block's `tool` can be `{"type": "tool_definition", "definition": {...}}`, carrying the same entry you would have put in `tools`:
13451355
1356```json
1357{
1358 "role": "system",
1359 "content": [
1360 {
1361 "type": "tool_addition",
1362 "tool": {
1363 "type": "tool_definition",
1364 "definition": {
1365 "name": "db_query",
1366 "description": "Run a read-only SQL query against the analytics database.",
1367 "input_schema": {
1368 "type": "object",
1369 "properties": { "sql": { "type": "string" } },
1370 "required": ["sql"]
1371 }
1372 }
1373 }
1374 }
1375 ]
1376}
1377```
1378
1379The new tool arrives in `messages`, `tools` never changes, and earlier thinking stays valid. Keep at least one tool without `defer_loading: true` in `tools`: if every tool there is deferred, the first tool you define this way costs one full prompt cache miss.
1380
1381If the API connects to the MCP server for you through the [MCP connector](https://platform.claude.com/docs/en/agents-and-tools/mcp-connector), also send `mcp-client-2026-09-15`. It covers everything `mcp-client-2025-11-20` does, so send it instead of that one. The block's `definition` can then be an `mcp_toolset` for a server listed in `mcp_servers`. When the API has to fetch a server's tool list, the response starts with an `mcp_tool_listing` block for that server. Send it back unchanged with the rest of the assistant turn, and keep sending `mcp-client-2026-09-15` on every later request that carries it. The block pins the toolset to that list, so the API doesn't contact the server again for it. These MCP connector features are available on the Claude API.
1382
1383With only the older header, you can still append a tool you learn about mid-session to `tools` with `defer_loading: true`, then offer it with a `tool_addition` block. That's safe because the prefix check ignores a deferred tool until a `tool_addition` block references it. Adding a tool without `defer_loading: true` changes the prefix and invalidates earlier thinking.
1384
13461385The `role: "system"` messages that carry these blocks join the prefix for later thinking. Leave them in place on later requests.
13471386
13481387### Change effort with a per-message `output_config`
from line 1686
16471686 </Accordion>
16481687
16491688 <Accordion title="My tool list changes mid-session. How do I avoid invalidating the conversation?">
1650 Don't edit `tools`. Declare the full set at session start, mark tools that aren't available yet with `defer_loading: true`, and offer or withdraw them with `tool_addition` and `tool_removal` blocks. If you learn a tool's schema only mid-session, such as from an MCP server discovered at runtime, you can still append it to `tools` with `defer_loading: true` and offer it the same way. That's safe because an unreferenced deferred tool isn't part of the prefix. The `role: "system"` messages that carry these blocks join the prefix for later thinking, so don't move, reword, or delete them afterward. See [Add or remove tools with `tool_addition` and `tool_removal`](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#tool-changes).
1689 Don't edit `tools`. Declare the full set at session start, mark tools that aren't available yet with `defer_loading: true`, and offer or withdraw them with `tool_addition` and `tool_removal` blocks. If you learn a tool's schema only mid-session, define it inside the `tool_addition` block (`inline-tools-2026-09-15`, plus `mcp-client-2026-09-15` for a server the API's MCP connector reaches) and leave `tools` unchanged. The `role: "system"` messages that carry these blocks join the prefix for later thinking, so don't move, reword, or delete them afterward. See [Add or remove tools with `tool_addition` and `tool_removal`](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#tool-changes).
16511690 </Accordion>
16521691
16531692 <Accordion title="I compact by summarizing older turns and keeping recent turns verbatim. Does that still work?">