Follow Discord
Sweep 03 Oct 2026 · 20:28Z Build v2.1.289 510 read Stable v2.1.285 Latest v2.1.289 Next v2.1.289 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · api

preserved-thinking changedbuild-with-claude/preserved-thinking

Nearest release: v2.1.286, published an hour after this site recorded the change. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Recorded here
Lines+10added
Lines−0removed
From line 430 where the diff opens
First seen 1 Sep 2026 this site's first read of the page
Recorded edits17to this page, all time

The whole hunk

from line 430, old and new numbered
/
lines
from line 430
430430 
431431### Check whether your code edits the prefix
432432 
433<Tip>
434 **Automate this check with the Claude API skill.** In the latest version of [Claude Code](https://code.claude.com/docs/en/overview), open the repository that builds your requests and run the bundled [Claude API skill](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/claude-api-skill#checking-an-integration-for-preserved-thinking):
435 
436 ```text wrap
437 /claude-api preserved-thinking-migration
438 ```
439 
440 The skill captures a few of your own multi-turn sessions and diffs consecutive requests to find where your code edits the prefix. It then replays the sessions with `"drop_block"` and counts the thinking blocks the API drops. After that, it fixes one cause at a time and measures again after each fix.
441</Tip>
442 
433443First, diff what you send. Capture the request bodies your integration sends over a few normal turns, including a compaction or a tool change. For each pair of consecutive requests, compare `system`, `tools`, and the `messages` they share. They should be identical up to the newly appended turns.
434444 
435445Then confirm against the API. Add the `thinking-binding-controls-2026-08-01` beta header, set `prefix_mismatch_behavior` to `"drop_block"`, and run a normal multi-turn session through your integration on claude-fable-5-1. The following example runs two turns the way your integration should: `messages` only grows, each assistant turn goes back exactly as the API returned it, `thinking` blocks included, and `block_binding` is set on every request. After each turn it prints the number of `thinking` blocks in the response and the number of dropped blocks:
Feedback