Follow Discord
Sweep 09 Oct 2026 · 17:27Z Build v2.1.296 517 read Stable v2.1.287 Latest v2.1.296 Next v2.1.296 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One capture · api

One read of Claude Developer Platformapi-20261009T153709Z

536 pages moved out of 760 read.

Pages moved 536 significant first
Pages read 760 in this capture
Captured 15:37 UTC
Corpus hash 9680386eb15e corpus-hash

What this read moved

376-400 of 536, page 16 of 22

This capture is too large to show at once. Changes 376-400 of 536 are below, significant first; the rest are on the following screens.

api/skills Changed · +48 / -0 lines

from line 13
1313 
1414### Headers
1515 
16- `"anthropic-version": optional string`
17 
18 The version of the Claude API you want to use.
19 
20 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
21 
1622- `"anthropic-workspace-id": optional string`
1723 
1824 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 168
162168 
163169### Headers
164170 
171- `"anthropic-version": optional string`
172 
173 The version of the Claude API you want to use.
174 
175 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
176 
165177- `"anthropic-workspace-id": optional string`
166178 
167179 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 303
291303 
292304### Headers
293305 
306- `"anthropic-version": optional string`
307 
308 The version of the Claude API you want to use.
309 
310 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
311 
294312- `"anthropic-workspace-id": optional string`
295313 
296314 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 425
407425 
408426### Headers
409427 
428- `"anthropic-version": optional string`
429 
430 The version of the Claude API you want to use.
431 
432 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
433 
410434- `"anthropic-workspace-id": optional string`
411435 
412436 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 604
580604 
581605#### Headers
582606 
607- `"anthropic-version": optional string`
608 
609 The version of the Claude API you want to use.
610 
611 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
612 
583613- `"anthropic-workspace-id": optional string`
584614 
585615 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 719
689719 
690720#### Headers
691721 
722- `"anthropic-version": optional string`
723 
724 The version of the Claude API you want to use.
725 
726 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
727 
692728- `"anthropic-workspace-id": optional string`
693729 
694730 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 829
793829 
794830#### Headers
795831 
832- `"anthropic-version": optional string`
833 
834 The version of the Claude API you want to use.
835 
836 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
837 
796838- `"anthropic-workspace-id": optional string`
797839 
798840 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 925
883925 Requests carrying the `skills-2025-10-02` beta header address versions by their Unix epoch timestamp instead (e.g., "1759178010641129").
884926 
885927#### Headers
928 
929- `"anthropic-version": optional string`
930 
931 The version of the Claude API you want to use.
932 
933 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
886934 
887935- `"anthropic-workspace-id": optional string`
888936 

api/skills/create Changed · +6 / -0 lines

from line 11
1111 
1212## Headers
1313 
14- `"anthropic-version": optional string`
15 
16 The version of the Claude API you want to use.
17 
18 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
19 
1420- `"anthropic-workspace-id": optional string`
1521 
1622 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).

api/skills/delete Changed · +6 / -0 lines

from line 19
1919 
2020## Headers
2121 
22- `"anthropic-version": optional string`
23 
24 The version of the Claude API you want to use.
25 
26 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
27 
2228- `"anthropic-workspace-id": optional string`
2329 
2430 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).

api/skills/list Changed · +6 / -0 lines

from line 36
3636 
3737## Headers
3838 
39- `"anthropic-version": optional string`
40 
41 The version of the Claude API you want to use.
42 
43 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
44 
3945- `"anthropic-workspace-id": optional string`
4046 
4147 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).

api/skills/retrieve Changed · +6 / -0 lines

from line 19
1919 
2020## Headers
2121 
22- `"anthropic-version": optional string`
23 
24 The version of the Claude API you want to use.
25 
26 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
27 
2228- `"anthropic-workspace-id": optional string`
2329 
2430 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).

api/skills/versions Changed · +24 / -0 lines

from line 21
2121 
2222### Headers
2323 
24- `"anthropic-version": optional string`
25 
26 The version of the Claude API you want to use.
27 
28 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
29 
2430- `"anthropic-workspace-id": optional string`
2531 
2632 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 136
130136 
131137### Headers
132138 
139- `"anthropic-version": optional string`
140 
141 The version of the Claude API you want to use.
142 
143 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
144 
133145- `"anthropic-workspace-id": optional string`
134146 
135147 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 246
234246 
235247### Headers
236248 
249- `"anthropic-version": optional string`
250 
251 The version of the Claude API you want to use.
252 
253 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
254 
237255- `"anthropic-workspace-id": optional string`
238256 
239257 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).
from line 342
324342 Requests carrying the `skills-2025-10-02` beta header address versions by their Unix epoch timestamp instead (e.g., "1759178010641129").
325343 
326344### Headers
345 
346- `"anthropic-version": optional string`
347 
348 The version of the Claude API you want to use.
349 
350 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
327351 
328352- `"anthropic-workspace-id": optional string`
329353 

api/skills/versions/create Changed · +6 / -0 lines

from line 19
1919 
2020## Headers
2121 
22- `"anthropic-version": optional string`
23 
24 The version of the Claude API you want to use.
25 
26 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
27 
2228- `"anthropic-workspace-id": optional string`
2329 
2430 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).

api/skills/versions/delete Changed · +6 / -0 lines

from line 25
2525 
2626## Headers
2727 
28- `"anthropic-version": optional string`
29 
30 The version of the Claude API you want to use.
31 
32 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
33 
2834- `"anthropic-workspace-id": optional string`
2935 
3036 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).

api/skills/versions/list Changed · +6 / -0 lines

from line 33
3333 
3434## Headers
3535 
36- `"anthropic-version": optional string`
37 
38 The version of the Claude API you want to use.
39 
40 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
41 
3642- `"anthropic-workspace-id": optional string`
3743 
3844 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).

api/skills/versions/retrieve Changed · +6 / -0 lines

from line 25
2525 
2626## Headers
2727 
28- `"anthropic-version": optional string`
29 
30 The version of the Claude API you want to use.
31 
32 Read more about versioning and our version history [here](https://platform.claude.com/docs/en/api/versioning).
33 
2834- `"anthropic-workspace-id": optional string`
2935 
3036 Optional header to select the Workspace for this request. The value is a Workspace ID (for example, `wrkspc_011CZkZaBF1tNoB5wlCeusgy`).

build-with-claude/claude-on-vertex-ai Changed · +20 / -9 lines

from line 69
6969 import com.anthropic.models.messages.MessageCreateParams;
7070 import com.anthropic.models.messages.Model;
7171 import com.anthropic.vertex.backends.VertexBackend;
72 import com.google.auth.oauth2.GoogleCredentials;
7273 
73 void main() {
74 void main() throws Exception {
7475 AnthropicClient client = AnthropicOkHttpClient.builder()
75 .backend(VertexBackend.fromEnv())
76 .backend(
77 VertexBackend.builder()
78 .googleCredentials(GoogleCredentials.getApplicationDefault())
79 .region("global")
80 .project("MY_PROJECT_ID")
81 .build()
82 )
7683 .build();
7784 
7885 MessageCreateParams params = MessageCreateParams.builder()
from line 274
267274 import com.anthropic.models.messages.MessageCreateParams;
268275 import com.anthropic.models.messages.Model;
269276 import com.anthropic.vertex.backends.VertexBackend;
277 import com.google.auth.oauth2.GoogleCredentials;
270278 
271 void main() {
272 // Uses default Google Cloud credentials
279 void main() throws Exception {
273280 AnthropicClient client = AnthropicOkHttpClient.builder()
274 .backend(VertexBackend.fromEnv())
281 .backend(
282 VertexBackend.builder()
283 .googleCredentials(GoogleCredentials.getApplicationDefault())
284 .region("global")
285 .project("MY_PROJECT_ID")
286 .build()
287 )
275288 .build();
276289 
277290 Message message = client
from line 551
538551 import com.google.auth.oauth2.GoogleCredentials;
539552 
540553 void main() throws Exception {
541 // Uses default Google Cloud credentials
542554 AnthropicClient client = AnthropicOkHttpClient.builder()
543555 .backend(
544556 VertexBackend.builder()
from line 934
922934 import com.google.auth.oauth2.GoogleCredentials;
923935 
924936 void main() throws Exception {
925 // Uses default Google Cloud credentials with specific region
926937 AnthropicClient client = AnthropicOkHttpClient.builder()
927938 .backend(
928939 VertexBackend.builder()

build-with-claude/claude-platform-on-aws Changed · +28 / -28 lines

from line 387
387387 -H "anthropic-version: 2023-06-01" \
388388 -H "anthropic-workspace-id: $ANTHROPIC_AWS_WORKSPACE_ID" \
389389 -d '{
390 "model": "claude-sonnet-5",
390 "model": "claude-sonnet-5-5",
391391 "max_tokens": 1024,
392392 "messages": [
393393 {"role": "user", "content": "Hello!"}
from line 404
404404 ant messages create \
405405 --base-url https://aws-external-anthropic.us-west-2.api.aws \
406406 --workspace-id "$ANTHROPIC_AWS_WORKSPACE_ID" \
407 --model claude-sonnet-5 \
407 --model claude-sonnet-5-5 \
408408 --max-tokens 1024 \
409409 --message '{role: user, content: "Hello!"}' \
410410 --transform content
from line 416
416416 client = AnthropicAWS()
417417 
418418 message = client.messages.create(
419 model="claude-sonnet-5",
419 model="claude-sonnet-5-5",
420420 max_tokens=1024,
421421 messages=[{"role": "user", "content": "Hello!"}],
422422 )
from line 429
429429 const client = new AnthropicAws();
430430 
431431 const message = await client.messages.create({
432 model: "claude-sonnet-5",
432 model: "claude-sonnet-5-5",
433433 max_tokens: 1024,
434434 messages: [{ role: "user", content: "Hello!" }]
435435 });
from line 444
444444 
445445 var message = await client.Messages.Create(new()
446446 {
447 Model = Model.ClaudeSonnet5,
447 Model = Model.ClaudeSonnet5_5,
448448 MaxTokens = 1024,
449449 Messages = [new() { Role = Role.User, Content = "Hello!" }]
450450 });
from line 459
459459 }
460460 
461461 message, err := client.Messages.New(context.Background(), anthropic.MessageNewParams{
462 Model: anthropic.ModelClaudeSonnet5,
462 Model: anthropic.ModelClaudeSonnet5_5,
463463 MaxTokens: 1024,
464464 Messages: []anthropic.MessageParam{
465465 anthropic.NewUserMessage(anthropic.NewTextBlock("Hello!")),
from line 487
487487 
488488 Message message = client.messages().create(
489489 MessageCreateParams.builder()
490 .model(Model.CLAUDE_SONNET_5)
490 .model(Model.CLAUDE_SONNET_5_5)
491491 .maxTokens(1024)
492492 .addUserMessage("Hello!")
493493 .build()
from line 503
503503 $client = new Client();
504504 
505505 $message = $client->messages->create(
506 model: 'claude-sonnet-5',
506 model: 'claude-sonnet-5-5',
507507 maxTokens: 1024,
508508 messages: [['role' => 'user', 'content' => 'Hello!']],
509509 );
from line 517
517517 client = Anthropic::AWSClient.new
518518 
519519 message = client.messages.create(
520 model: "claude-sonnet-5",
520 model: "claude-sonnet-5-5",
521521 max_tokens: 1024,
522522 messages: [{ role: "user", content: "Hello!" }]
523523 )
from line 546
546546* **Beta features:** Pass the standard `anthropic-beta` header to access beta features, just as you would with the Claude API.
547547* **Agent Skills:** Use pre-built and custom [Agent Skills](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview) with the same `container.skills` parameter as the Claude API. All pre-built Skills (PowerPoint, Excel, Word, PDF) work out of the box.
548548* **Code execution:** Run code in Anthropic's managed sandbox using the [code execution tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool).
549* **Tool use:** Computer use and all other [tool use capabilities](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview) are available.
550* **Extended thinking:** Enable extended thinking with the same parameters as the Claude API.
549* **Tool use:** All [tool use capabilities](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview) are available except the computer use and browser use toolsets; see [Features not supported](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws#features-not-supported).
550* **Thinking:** [Adaptive thinking](https://platform.claude.com/docs/en/build-with-claude/thinking), the [effort parameter](https://platform.claude.com/docs/en/build-with-claude/effort), and, on the models that support it, [extended thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) work with the same parameters as the Claude API.
551551* **Streaming:** Full SSE streaming support for real-time responses.
552552* **Batch processing:** Submit batch requests for high-throughput workloads.
553553* **Prompt caching:** Cache tools, system prompts, and message history to reduce latency and cost. All prompt caching capabilities (5-minute TTL, 1-hour TTL, and automatic caching) are available.
from line 571
571571The following capabilities are not currently available on Claude Platform on AWS:
572572 
573573* **HIPAA readiness:** Anthropic's HIPAA-ready program is not available. See [API and data retention](https://platform.claude.com/docs/en/manage-claude/api-and-data-retention).
574* **Computer use and browser use toolsets:** `computer_toolset_20260801` and `browser_toolset_20260801` are not currently available on Claude Platform on AWS. The beta [computer use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool#earlier-tool-versions) tool versions remain available.
574* **Computer use and browser use toolsets:** `computer_toolset_20260801` and `browser_toolset_20260801` are not currently available on Claude Platform on AWS. The beta computer use tool versions remain available for the models listed under [Earlier tool versions](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool#earlier-tool-versions).
575575 
576576- **Admin API:** Workspace endpoints (create, get, list, update, and archive on `/v1/organizations/workspaces`) and external key endpoints (register, get, list, update, and delete on `/v1/organizations/external_keys`, for [CMEK](https://platform.claude.com/docs/en/manage-claude/cmek); keys are validated when attached to a workspace rather than through a validate endpoint) are available. Other Admin API endpoints (organization members, workspace members, invites, API keys, usage reports, cost reports, and rate limit reports) are not currently available. View usage and cost data in the [Claude Console](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws#using-the-claude-console) instead. AWS IAM manages organization membership.
577577- **Workspace member management:** Adding or removing users from individual workspaces is not available. AWS IAM policies on workspace ARNs control access.
from line 610
610610 -H "anthropic-version: 2023-06-01" \
611611 -H "anthropic-workspace-id: $ANTHROPIC_AWS_WORKSPACE_ID" \
612612 -d '{
613 "model": "claude-sonnet-5",
613 "model": "claude-sonnet-5-5",
614614 "max_tokens": 1024,
615615 "inference_geo": "us",
616616 "messages": [
from line 628
628628 ant messages create \
629629 --base-url https://aws-external-anthropic.us-west-2.api.aws \
630630 --workspace-id "$ANTHROPIC_AWS_WORKSPACE_ID" \
631 --model claude-sonnet-5 \
631 --model claude-sonnet-5-5 \
632632 --max-tokens 1024 \
633633 --inference-geo us \
634634 --message '{role: user, content: "Hello!"}' \
from line 640
640640 
641641 client = AnthropicAWS()
642642 message = client.messages.create(
643 model="claude-sonnet-5",
643 model="claude-sonnet-5-5",
644644 max_tokens=1024,
645645 inference_geo="us",
646646 messages=[{"role": "user", "content": "Hello!"}],
from line 652
652652 import AnthropicAws from "@anthropic-ai/aws-sdk";
653653 const client = new AnthropicAws();
654654 const message = await client.messages.create({
655 model: "claude-sonnet-5",
655 model: "claude-sonnet-5-5",
656656 max_tokens: 1024,
657657 inference_geo: "us",
658658 messages: [{ role: "user", content: "Hello!" }]
from line 668
668668 
669669 var message = await client.Messages.Create(new()
670670 {
671 Model = Model.ClaudeSonnet5,
671 Model = Model.ClaudeSonnet5_5,
672672 MaxTokens = 1024,
673673 InferenceGeo = "us",
674674 Messages = [new() { Role = Role.User, Content = "Hello!" }]
from line 684
684684 }
685685 
686686 message, err := client.Messages.New(context.Background(), anthropic.MessageNewParams{
687 Model: anthropic.ModelClaudeSonnet5,
687 Model: anthropic.ModelClaudeSonnet5_5,
688688 MaxTokens: 1024,
689689 InferenceGeo: anthropic.String("us"),
690690 Messages: []anthropic.MessageParam{
from line 713
713713 
714714 Message message = client.messages().create(
715715 MessageCreateParams.builder()
716 .model(Model.CLAUDE_SONNET_5)
716 .model(Model.CLAUDE_SONNET_5_5)
717717 .maxTokens(1024)
718718 .inferenceGeo("us")
719719 .addUserMessage("Hello!")
from line 730
730730 $client = new Client();
731731 
732732 $message = $client->messages->create(
733 model: 'claude-sonnet-5',
733 model: 'claude-sonnet-5-5',
734734 maxTokens: 1024,
735735 inferenceGeo: 'us',
736736 messages: [['role' => 'user', 'content' => 'Hello!']],
from line 745
745745 client = Anthropic::AWSClient.new
746746 
747747 message = client.messages.create(
748 model: "claude-sonnet-5",
748 model: "claude-sonnet-5-5",
749749 max_tokens: 1024,
750750 inference_geo: "us",
751751 messages: [{ role: "user", content: "Hello!" }]
from line 886
886886 -H "anthropic-version: 2023-06-01" \
887887 -H "anthropic-workspace-id: $ANTHROPIC_AWS_WORKSPACE_ID" \
888888 -d '{
889 "model": "claude-sonnet-5",
889 "model": "claude-sonnet-5-5",
890890 "max_tokens": 1024,
891891 "messages": [
892892 {"role": "user", "content": "Hello!"}
from line 905
905905 client = AnthropicAWS()
906906 
907907 response = client.messages.with_raw_response.create(
908 model="claude-sonnet-5",
908 model="claude-sonnet-5-5",
909909 max_tokens=1024,
910910 messages=[{"role": "user", "content": "Hello!"}],
911911 )
from line 924
924924 
925925 const { data: message, response } = await client.messages
926926 .create({
927 model: "claude-sonnet-5",
927 model: "claude-sonnet-5-5",
928928 max_tokens: 1024,
929929 messages: [{ role: "user", content: "Hello!" }]
930930 })
from line 943
943943 
944944 var response = await client.WithRawResponse.Messages.Create(new()
945945 {
946 Model = Model.ClaudeSonnet5,
946 Model = Model.ClaudeSonnet5_5,
947947 MaxTokens = 1024,
948948 Messages = [new() { Role = Role.User, Content = "Hello!" }]
949949 });
from line 963
963963 message, err := client.Messages.New(
964964 context.Background(),
965965 anthropic.MessageNewParams{
966 Model: anthropic.ModelClaudeSonnet5,
966 Model: anthropic.ModelClaudeSonnet5_5,
967967 MaxTokens: 1024,
968968 Messages: []anthropic.MessageParam{
969969 anthropic.NewUserMessage(anthropic.NewTextBlock("Hello!")),
from line 996
996996 
997997 HttpResponseFor<Message> response = client.messages().withRawResponse().create(
998998 MessageCreateParams.builder()
999 .model(Model.CLAUDE_SONNET_5)
999 .model(Model.CLAUDE_SONNET_5_5)
10001000 .maxTokens(1024)
10011001 .addUserMessage("Hello!")
10021002 .build()
from line 1014
10141014 $client = new Client();
10151015 
10161016 $response = $client->messages->raw->create(
1017 model: 'claude-sonnet-5',
1017 model: 'claude-sonnet-5-5',
10181018 maxTokens: 1024,
10191019 messages: [['role' => 'user', 'content' => 'Hello!']],
10201020 );

build-with-claude/compaction-thinking-blocks Changed · +2 / -3 lines

from line 23
2323 supportedPlatforms:
2424 Claude API: beta
2525 Claude Platform on AWS: beta
26 Amazon Bedrock: not available
2726 Google Cloud: beta
2827 Microsoft Foundry: beta
2928---
from line 57
5857 
5958To add an instruction or change the available tools without touching `system` or `tools`, append the change to `messages`, as described in [Make changes without editing the prefix](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#replace-prefix-edits).
6059 
61Mid-conversation system messages inside the summarized turns are summarized too, so their text instructions stop applying after the swap. To keep one in force, state it again in a `role: "system"` message directly after the first new `user` turn that follows the kept turns. Tool changes inside those turns carry over on their own when the compaction request also carries `inline-tools-2026-09-15`: the returned block records their net effect in its `tool_changes` field, so send the block back unmodified. If the block has no `tool_changes` field, restate those tool changes the same way. A system message placed between the block and the kept turns breaks their thinking.
60Mid-conversation system messages inside the summarized turns are summarized too, so their text instructions stop applying after the swap. A [per-message effort](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) level they set isn't carried over either: until a later message sets a level, turns run at the request's top-level `output_config.effort`, or the model's default if you don't set one. To keep an instruction or effort level in force, state it again in a `role: "system"` message directly after the first new `user` turn that follows the kept turns. Tool changes inside those turns carry over on their own when the compaction request also carries `inline-tools-2026-09-15`: the returned block records their net effect in its `tool_changes` field, so send the block back unmodified. If the block has no `tool_changes` field, restate those tool changes the same way. A system message placed between the block and the kept turns breaks their thinking.
6261 
6362## Check that the kept thinking held
6463 
from line 702
703702Dropped thinking blocks: 0
704703```
705704 
706In production, `"drop_block"` keeps requests succeeding when a condition doesn't hold, and reports each dropped block in `input_transformations` with `reason: "prefix_binding_mismatch"`. An entry whose `path` falls in a kept turn means that turn's thinking didn't hold. [What the API does with an invalid block](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#mismatch-behavior) describes what is dropped and says how to alert on it.
705In production, `"drop_block"` keeps requests succeeding when a condition doesn't hold, and reports each dropped block in `input_transformations` with `reason: "prefix_binding_mismatch"`. An entry whose `path` falls in a kept turn means that turn's thinking didn't hold. [What the API does with an invalid block](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#mismatch-behavior) describes what is dropped and says how to alert on it. On Claude Sonnet 5.5 and Claude Haiku 5.5, `block_binding` works only with `thinking: {"type": "adaptive"}`. A request that sets it with `thinking: {"type": "between_tools"}` on Claude Sonnet 5.5, or with `thinking: {"type": "disabled"}` on Claude Haiku 5.5, returns a 400 error. On those requests, make sure every condition holds, or remove the `thinking` and `redacted_thinking` blocks from the kept turns.
707706 

build-with-claude/compaction-threshold Changed · +95 / -20 lines

from line 2900
29002900 
29012901On Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, and Claude Haiku 5.5, remove the `thinking` and `redacted_thinking` blocks from any assistant turn you re-insert after the compaction block, or send `thinking.block_binding.prefix_mismatch_behavior: "drop_block"` with the `thinking-binding-controls-2026-08-01` [beta header](https://platform.claude.com/docs/en/api/beta-headers). Those blocks were produced when the full history was present, so they no longer pass the [conversation check](https://platform.claude.com/docs/en/build-with-claude/thinking#preserved-in-conversation). Where the check is enforced, the continuation request is rejected with a 400 error. The preserved text and tool blocks can stay as they are. Letting the API summarize everything, without re-inserting earlier turns, avoids this. On Claude Sonnet 5.5, `block_binding` works only with `thinking: {"type": "adaptive"}`. With `between_tools`, remove the blocks instead. On Claude Haiku 5.5, `block_binding` works only with `thinking: {"type": "adaptive"}`, so with `thinking: {"type": "disabled"}`, remove the blocks instead.
29022902 
2903Here's an example that uses `pause_after_compaction` to preserve the prior exchange and the current user message (three messages total) verbatim instead of summarizing them:
2903Here's an example that uses `pause_after_compaction` to preserve the prior exchange and the current user message (three messages total) instead of summarizing them. It keeps their text and tool blocks unchanged, removes the `thinking` and `redacted_thinking` blocks from the preserved assistant turn, and drops that turn if nothing else is left:
29042904 
29052905<CodeGroup>
29062906 ```bash cURL
from line 2992
29922992 compaction_block = response.content[0]
29932993 
29942994 # Preserve the prior exchange + current user message (3 messages)
2995 # by including them after the compaction block
2996 preserved_messages = messages[-3:] if len(messages) >= 3 else messages
2995 # by including them after the compaction block, without thinking
2996 # blocks; an assistant turn left with no content is dropped
2997 preserved_messages = [
2998 {
2999 **message,
3000 "content": [
3001 block
3002 for block in message["content"]
3003 if block.type not in ("thinking", "redacted_thinking")
3004 ],
3005 }
3006 if message["role"] == "assistant"
3007 else message
3008 for message in messages[-3:]
3009 ]
3010 preserved_messages = [m for m in preserved_messages if m["content"]]
29973011 
29983012 # Build new message list: compaction + preserved messages
29993013 new_assistant_content = [compaction_block]
from line 3072
30583072 const compactionBlock = response.content[0];
30593073 
30603074 // Preserve the prior exchange + current user message (3 messages)
3061 // by including them after the compaction block
3062 const preservedMessages = messages.length >= 3 ? messages.slice(-3) : [...messages];
3075 // by including them after the compaction block, without thinking
3076 // blocks; an assistant turn left with no content is dropped
3077 const preservedMessages = messages
3078 .slice(-3)
3079 .map((message) =>
3080 message.role === "assistant" && Array.isArray(message.content)
3081 ? {
3082 ...message,
3083 content: message.content.filter(
3084 (block) => block.type !== "thinking" && block.type !== "redacted_thinking"
3085 )
3086 }
3087 : message
3088 )
3089 .filter((message) => message.content.length > 0);
30633090 
30643091 // Build new message list: compaction + preserved messages
30653092 const messagesAfterCompaction: Anthropic.Beta.Messages.BetaMessageParam[] = [
from line 3157
31303157 if (!response.Content[0].TryPickCompaction(out _))
31313158 throw new InvalidOperationException("Expected compaction block");
31323159 
3133 var preserved = messages.Count >= 3
3134 ? messages.Skip(messages.Count - 3).ToList()
3135 : new List<BetaMessageParam>(messages);
3160 // Preserve the prior exchange + current user message (3 messages),
3161 // without thinking blocks; an assistant turn left empty is dropped
3162 var preserved = messages
3163 .Skip(Math.Max(0, messages.Count - 3))
3164 .Select(message => message.Content.TryPickBetaContentBlockParams(out var blocks)
3165 ? new BetaMessageParam
3166 {
3167 Role = message.Role,
3168 Content = blocks
3169 .Where(block => block.Type.GetString() is not ("thinking" or "redacted_thinking"))
3170 .ToList()
3171 }
3172 : message)
3173 .Where(message => !message.Content.TryPickBetaContentBlockParams(out var blocks) || blocks.Count > 0)
3174 .ToList();
31363175 
31373176 var messagesAfterCompaction = new List<BetaMessageParam>
31383177 {
from line 3254
32153254 if response.StopReason == "compaction" {
32163255 compactionParam := response.Content[0].ToParam()
32173256 
3257 // Preserve the prior exchange + current user message (3 messages),
3258 // without thinking blocks; an assistant turn left empty is dropped
32183259 var preserved []anthropic.BetaMessageParam
3219 if len(messages) >= 3 {
3220 preserved = messages[len(messages)-3:]
3221 } else {
3222 preserved = messages
3260 for _, message := range messages[max(0, len(messages)-3):] {
3261 var content []anthropic.BetaContentBlockParamUnion
3262 for _, block := range message.Content {
3263 if block.OfThinking == nil && block.OfRedactedThinking == nil {
3264 content = append(content, block)
3265 }
3266 }
3267 if len(content) == 0 {
3268 continue
3269 }
3270 message.Content = content
3271 preserved = append(preserved, message)
32233272 }
32243273 
32253274 messagesAfterCompaction := []anthropic.BetaMessageParam{
from line 3346
32973346 // Check if compaction occurred and paused
32983347 if (response.stopReason().isPresent()
32993348 && response.stopReason().get().equals(BetaStopReason.COMPACTION)) {
3300 // Preserve the prior exchange + current user message (3 messages)
3301 List<BetaMessageParam> preservedMessages = messages.size() >= 3
3302 ? new ArrayList<>(messages.subList(messages.size() - 3, messages.size()))
3303 : new ArrayList<>(messages);
3349 // Preserve the prior exchange + current user message (3 messages),
3350 // without thinking blocks; an assistant turn left empty is dropped
3351 List<BetaMessageParam> preservedMessages = new ArrayList<>();
3352 for (BetaMessageParam message : messages.subList(Math.max(0, messages.size() - 3), messages.size())) {
3353 if (message.content().isBetaContentBlockParams()) {
3354 message = message.toBuilder()
3355 .contentOfBetaContentBlockParams(message.content().asBetaContentBlockParams().stream()
3356 .filter(block -> !block.isThinking() && !block.isRedactedThinking())
3357 .toList())
3358 .build();
3359 if (message.content().asBetaContentBlockParams().isEmpty()) {
3360 continue;
3361 }
3362 }
3363 preservedMessages.add(message);
3364 }
33043365 
33053366 // Build new message list: compaction + preserved messages
33063367 List<BetaMessageParam> messagesAfterCompaction = new ArrayList<>();
from line 3429
33683429 if ($response->stopReason === 'compaction') {
33693430 $compactionBlock = $response->content[0];
33703431 
3371 $preserved = count($messages) >= 3
3372 ? array_slice($messages, -3)
3373 : $messages;
3432 // Preserve the prior exchange + current user message (3 messages),
3433 // without thinking blocks; an assistant turn left empty is dropped
3434 $preserved = array_values(array_filter(array_map(
3435 fn($message) => $message['role'] === 'assistant'
3436 ? array_merge($message, ['content' => array_values(array_filter(
3437 $message['content'],
3438 fn($block) => !in_array($block->type, ['thinking', 'redacted_thinking'], true)
3439 ))])
3440 : $message,
3441 array_slice($messages, -3)
3442 ), fn($message) => $message['content'] !== []));
33743443 
33753444 $messagesAfterCompaction = array_merge(
33763445 [['role' => 'assistant', 'content' => [$compactionBlock]]],
from line 3500
34313500 if response.stop_reason == :compaction
34323501 compaction_block = response.content[0]
34333502 
3434 preserved = messages.length >= 3 ? messages[-3..-1] : messages.dup
3503 # Preserve the prior exchange + current user message (3 messages),
3504 # without thinking blocks; an assistant turn left empty is dropped
3505 preserved = messages.last(3).map do |message|
3506 next message unless message[:role] == "assistant"
3507 
3508 message.merge(content: message[:content].reject { |block| %i[thinking redacted_thinking].include?(block.type) })
3509 end.reject { |message| message[:content].empty? }
34353510 
34363511 messages_after_compaction = [
34373512 { role: "assistant", content: [compaction_block] }

build-with-claude/context-editing Changed · +3 / -3 lines

from line 17
1717Context editing allows you to selectively clear specific content from conversation history as it grows. Beyond optimizing costs and staying within limits, this is about actively curating what Claude sees: context is a finite resource with diminishing returns, and irrelevant content degrades model focus. Context editing gives you fine-grained runtime control over that curation. For the broader principles behind context management, see [Effective context engineering](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents). This page covers:
1818 
1919* **Tool result clearing** - Best for agentic workflows with heavy tool use where old tool results are no longer needed
20* **Thinking block clearing** - For managing thinking blocks when using extended thinking, with options to preserve recent thinking for context continuity
20* **Thinking block clearing** - For managing thinking blocks from earlier turns, with options to preserve recent thinking for context continuity
2121* **Client-side SDK compaction** - An SDK-based alternative for summary-based context management (server-side compaction is generally preferred)
2222 
2323| Approach | Where it runs | Strategies | How it works |
from line 41
4141 
4242### Thinking block clearing
4343 
44The `clear_thinking_20251015` strategy manages `thinking` blocks in conversations when extended thinking is enabled. This strategy gives you control over thinking preservation: you can choose to keep more thinking blocks to maintain reasoning continuity, or clear them more aggressively to save context space.
44The `clear_thinking_20251015` strategy manages `thinking` blocks from earlier assistant turns. This strategy gives you control over thinking preservation: you can choose to keep more thinking blocks to maintain reasoning continuity, or clear them more aggressively to save context space.
4545 
4646<Tip>
4747 **Default behavior:** The default varies by model class.
from line 696
696696 
697697## Thinking block clearing usage
698698 
699Enable thinking block clearing to manage context and prompt caching effectively when extended thinking is enabled:
699Use thinking block clearing to manage context and prompt caching in conversations that carry thinking blocks from earlier turns:
700700 
701701<CodeGroup>
702702 ```bash cURL

build-with-claude/effort Changed · +3 / -3 lines

from line 339
339339 
340340### Recommended effort levels for Claude Haiku 5.5
341341 
342Claude Haiku 5.5 supports all five effort levels, and `medium` is the default on the Claude API and in Claude Code. Effort is the main control for how much the model thinks, and with it quality, latency, and cost. **Start with `medium`** for most work, including agentic coding. Use `low`, the cheapest and fastest level, for chat, short tool tasks, and simple, high-volume requests. In long agent prompts, the model is more likely to skip a search, stop early, or skip a check at `low`. Use `high` for knowledge work, longer agent tasks, and strict instruction following. Use `xhigh` or `max` only where your evals show a quality gain, and compare them with Claude Sonnet 5.5 on performance, cost, and speed. Thinking is on by default and counts toward `max_tokens`, so leave room for it. See [Prompting Claude Haiku 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#use-effort-to-control-thinking).
342Claude Haiku 5.5 supports all five effort levels, and `medium` is the default. Effort is the main control for how much the model thinks, and with it quality, latency, and cost. **Start with `medium`** for most work, including agentic coding. Use `low`, the cheapest and fastest level, for chat, short tool tasks, and simple, high-volume requests. In long agent prompts, the model is more likely to skip a search, stop early, or skip a check at `low`. Use `high` for knowledge work, longer agent tasks, and strict instruction following. Use `xhigh` or `max` only where your evals show a quality gain, and compare them with Claude Sonnet 5.5 on performance, cost, and speed. Thinking is on by default and counts toward `max_tokens`, so leave room for it. See [Prompting Claude Haiku 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#use-effort-to-control-thinking).
343343 
344344To get less thinking, lower the effort level. You can also send `thinking: {"type": "disabled"}` at `high` effort or below. At `xhigh` or `max`, it returns a 400 error, so use adaptive thinking there: omit the `thinking` field or send `thinking: {"type": "adaptive"}`.
345345 
from line 381
381381 
382382Without the beta value, a per-message `output_config` returns a 400 error: `messages.N.output_config: Extra inputs are not permitted`, where `N` is the index of the `system` message in `messages`. With the beta value, models without per-message effort, including Claude Fable 5, return a 400 error: `output_config.effort requires a model that supports per-turn effort; this model does not`. On Amazon Bedrock, those models and Claude Opus 5 return the `Extra inputs are not permitted` error instead. On Claude Sonnet 5.5 with `thinking: {"type": "between_tools"}` and on Claude Haiku 5.5 with `thinking: {"type": "disabled"}`, effort can't change mid-conversation: a per-message `output_config.effort` that differs from the level in effect returns a 400 error. To vary effort per turn, use adaptive thinking.
383383 
384Add a `role: "system"` message with empty `content` and the new level in `output_config.effort`. The new level takes effect from the next `user` turn and holds until a later message changes it. Everything before that message is unchanged, so the cached prefix still matches.
384Add a `role: "system"` message with empty `content` and the new level in `output_config.effort`. Placed directly after a `user` turn with new input, it sets the level for Claude's reply to that turn. Placed anywhere else, such as between an `assistant` turn and the next `user` turn as in the following example, it takes effect from the next `user` turn with new input. A `user` turn that holds only `tool_result` blocks doesn't count as new input, so a change placed after one in a tool-use loop waits for the next `user` turn with new input. The new level then holds until a later message changes it. Everything before that message is unchanged, so the cached prefix still matches.
385385 
386386The following example starts at `high`, then drops to `low` for a routine follow-up:
387387 
from line 654
654654 
655655An effort-only system message carries no text, so the [placement rules for mid-conversation system messages](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#limitations) don't apply. It can appear anywhere in `messages`, including as the first entry or between an `assistant` turn and the next `user` turn. Values are the level names (`low`, `medium`, `high`, `xhigh`, and `max`).
656656 
657On Claude Fable 5.1, prefer this form over changing the top-level value between requests. A top-level change restarts the cache and also steers the model less reliably: its earlier replies were written at the previous level, and it tends to stay consistent with them.
657Prefer this form over changing the top-level value between requests. A top-level change restarts the cache and, on Claude Fable 5.1, also steers the model less reliably: its earlier replies were written at the previous level, and it tends to stay consistent with them.
658658 
659659### Top-level effort on the next request
660660 

build-with-claude/fallback-credit Changed · +3 / -3 lines

from line 562
562562 
563563## Where it works
564564 
565Fallback credit is in beta on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. Refusals in [Message Batches](https://platform.claude.com/docs/en/build-with-claude/batch-processing) don't mint credit tokens, and redemption applies only to direct Messages API requests: a token passed on a batch request is accepted but ignored.
565Fallback credit is in beta on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. Refusals in [Message Batches](https://platform.claude.com/docs/en/build-with-claude/batch-processing) don't mint credit tokens, and redemption applies only to direct Messages API requests: a token passed on a batch request is accepted but ignored. A Claude Haiku 5.5 refusal carries no fallback credit, so a retry after one pays the full cost of writing the fallback model's prompt cache.
566566 
567The retry model must be one of the refused model's permitted fallback targets. For Claude Fable 5.1 and Claude Fable 5, those are Claude Opus 4.8 (`claude-opus-4-8`) and Claude Opus 5 (`claude-opus-5`).
567The retry model must be one of the refused model's permitted fallback targets. For Claude Fable 5.1, Claude Fable 5, and Claude Opus 5.5, those are Claude Opus 4.8 (`claude-opus-4-8`) and Claude Opus 5 (`claude-opus-5`). For Claude Opus 5, it is Claude Opus 4.8. For Claude Sonnet 5.5, it is Claude Sonnet 5 (`claude-sonnet-5`).
568568 
569569<Accordion title="Looking up permitted fallback targets programmatically">
570570 On the Claude API and Claude Platform on AWS, the target list is published as `allowed_fallback_models` on each model's entry in the [Models API](https://platform.claude.com/docs/en/api/models/list) when the `server-side-fallback-2026-07-01` beta header is set. The list is not yet visible under the `fallback-credit-*` header alone. It is not exposed on Amazon Bedrock, Google Cloud, or Microsoft Foundry.

build-with-claude/mid-conversation-effort-example Changed · +500 / -557 lines

The two sides of this change are more than 400 edits apart, too far apart to line up, so this is the differ's own diff of it and the words inside a line are not marked.

from line 57
5757 ```typescript TypeScript
5858 import { exec } from "node:child_process";
5959 import { createHash } from "node:crypto";
60 import { rmSync } from "node:fs";
61 import { mkdtemp, readFile, rename, writeFile } from "node:fs/promises";
60 import { mkdtempDisposable, readFile, rename, writeFile } from "node:fs/promises";
6261 import { tmpdir } from "node:os";
6362 import { join } from "node:path";
64 import { promisify } from "node:util";
6563 
6664 import Anthropic from "@anthropic-ai/sdk";
6765 
from line 99
10199 const string systemPrompt = "You are a helpful general-purpose agent. Answer the user's request directly.";
102100 
103101 const int requestTimeoutSeconds = 600;
104 // The other ports stream with max_tokens 64000. This port uses non-streaming
105 // Messages.Create, and the API rejects non-streaming requests at that size.
106 // 8192 is the non-streaming ceiling for Opus 4.0 and 4.1 and a conservative
107 // choice for newer Opus models.
108 const int requestMaxTokens = 8192;
109102 const int bashTimeoutSeconds = 60;
110103 const int toolResultMaxChars = 8000;
111104 const int maxConcurrent = 10;
112105 var docTestMode = Environment.GetEnvironmentVariable("DOC_TEST_MODE") is { Length: > 0 };
113 int maxTotalSubtasks = docTestMode ? 2 : 200;
114 int maxSubagentTurns = docTestMode ? 1 : 15;
115 int maxMainTurns = docTestMode ? 1 : 30;
106 var maxTotalSubtasks = docTestMode ? 2 : 200;
107 var maxSubagentTurns = docTestMode ? 1 : 15;
108 var maxMainTurns = docTestMode ? 1 : 30;
116109 const int turnsBetweenRefreshers = 10;
117 var journalPath = Environment.GetEnvironmentVariable("ORCH_JOURNAL") is { Length: > 0 } p ? p : "orchestration_journal.json";
110 var journalPath = Environment.GetEnvironmentVariable("ORCH_JOURNAL") is { Length: > 0 } configuredPath
111 ? configuredPath
112 : "orchestration_journal.json";
118113 ```
119114 
120115 ```go Go
from line 166
171166 ```
172167 
173168 ```java Java
169 // java.base (java.util, java.nio.file, java.util.concurrent, ...) is imported
170 // implicitly in a compact source file, so only the SDK and Jackson need imports.
174171 import com.anthropic.client.AnthropicClient;
175172 import com.anthropic.client.okhttp.AnthropicOkHttpClient;
173 import com.anthropic.core.JsonArray;
174 import com.anthropic.core.JsonBoolean;
175 import com.anthropic.core.JsonObject;
176 import com.anthropic.core.JsonString;
176177 import com.anthropic.core.JsonValue;
177178 import com.anthropic.core.RequestOptions;
178179 import com.anthropic.helpers.MessageAccumulator;
179 import com.anthropic.models.messages.ContentBlock;
180180 import com.anthropic.models.messages.ContentBlockParam;
181181 import com.anthropic.models.messages.Message;
182182 import com.anthropic.models.messages.MessageCreateParams;
from line 193
193193 import com.fasterxml.jackson.core.type.TypeReference;
194194 import com.fasterxml.jackson.databind.JsonNode;
195195 import com.fasterxml.jackson.databind.ObjectMapper;
196 import java.io.IOException;
197 import java.io.UncheckedIOException;
198 import java.nio.charset.StandardCharsets;
199 import java.nio.file.Files;
200 import java.nio.file.Path;
201 import java.nio.file.StandardCopyOption;
202 import java.security.MessageDigest;
203 import java.time.Duration;
204 import java.util.ArrayList;
205 import java.util.Comparator;
206 import java.util.HashMap;
207 import java.util.HexFormat;
208 import java.util.List;
209 import java.util.Map;
210 import java.util.Objects;
211 import java.util.Optional;
212 import java.util.concurrent.Callable;
213 import java.util.concurrent.CancellationException;
214 import java.util.concurrent.CompletableFuture;
215 import java.util.concurrent.ExecutionException;
216 import java.util.concurrent.ExecutorService;
217 import java.util.concurrent.Executors;
218 import java.util.concurrent.Future;
219 import java.util.concurrent.TimeUnit;
220 import java.util.concurrent.locks.ReentrantLock;
221 import java.util.stream.Collectors;
222 import java.util.stream.IntStream;
223196 
224197 AnthropicClient client = AnthropicOkHttpClient.fromEnv();
225198 
from line 215
242215 static final int MAX_MAIN_TURNS = DOC_TEST_MODE ? 1 : 30;
243216 static final int TURNS_BETWEEN_REFRESHERS = 10;
244217 static final Path JOURNAL_PATH = Path.of(Optional.ofNullable(System.getenv("ORCH_JOURNAL"))
245 .filter(s -> !s.isEmpty()).orElse("orchestration_journal.json"));
218 .filter(path -> !path.isEmpty())
219 .orElse("orchestration_journal.json"));
246220 ```
247221 
248222 ```php PHP
from line 331
357331 ```
358332 
359333 ```java Java
360 static final String MODE_ENTER =
361 "Orchestration mode is on: optimize for the most exhaustive, correct answer rather than "
362 + "the fastest one. Use the Workflow tool on every substantive task, sized to the problem's "
363 + "natural decomposition rather than the maximum the tool allows. See the Workflow tool's "
364 + "description for standing consent, granularity guidance, and quality patterns. Work solo "
365 + "only on conversational or trivial turns.";
334 static final String MODE_ENTER = """
335 Orchestration mode is on: optimize for the most exhaustive, correct answer rather than \
336 the fastest one. Use the Workflow tool on every substantive task, sized to the problem's \
337 natural decomposition rather than the maximum the tool allows. See the Workflow tool's \
338 description for standing consent, granularity guidance, and quality patterns. Work solo \
339 only on conversational or trivial turns.""";
366340 static final String MODE_REFRESH =
367341 "Orchestration mode is still on. Use the Workflow tool; see its standing consent section.";
368342 static final String MODE_EXIT =
from line 447
473447 ```typescript TypeScript
474448 const WORKFLOW_TOOL: Anthropic.Tool = {
475449 name: "Workflow",
476 description:
450 description: [
477451 "Orchestrate a multiagent workflow: split a large task into independent subtasks " +
478 "and run them as parallel agents, then collect their results.\n\n" +
452 "and run them as parallel agents, then collect their results.",
479453 "Opt-in: only use this tool when the user explicitly asks for a workflow, or when a " +
480 "system message confirms that orchestration mode is on.\n\n" +
454 "system message confirms that orchestration mode is on.",
481455 "Quality patterns: adversarial verification (a second wave of agents checks the first " +
482 "wave's findings against the source), a completeness critic (one agent hunts for what " +
483 "the others missed), and multiphase sequencing (understand, design, implement, and " +
484 "review as separate workflow calls, reading results between phases). A useful default " +
485 "is hybrid: scout inline first to discover the work-list, then fan out over it.\n\n" +
456 "wave's findings against the source), a completeness critic (one agent hunts for what " +
457 "the others missed), and multiphase sequencing (understand, design, implement, and " +
458 "review as separate workflow calls, reading results between phases). A useful default " +
459 "is hybrid: scout inline first to discover the work-list, then fan out over it.",
486460 "Granularity: scope each subtask to a distinct concern, component, or question rather " +
487 "than per line or per file section. Scale the count to what the user asked for: a " +
488 "focused review of a module of a few hundred lines rarely needs more than about ten " +
489 "subtasks; a broad audit of a large codebase can justify more.\n\n" +
461 "than per line or per file section. Scale the count to what the user asked for: a " +
462 "focused review of a module of a few hundred lines rarely needs more than about ten " +
463 "subtasks; a broad audit of a large codebase can justify more.",
490464 "Standing consent: while a system message confirms orchestration mode is on, that " +
491 "opt-in is standing. Author and run a workflow for every substantive task by default, " +
492 "and lean toward verifying findings adversarially. Work solo only on conversational " +
493 "turns or trivial mechanical edits. When a system message says the mode is off, " +
494 "revert to the opt-in rule above.",
465 "opt-in is standing. Author and run a workflow for every substantive task by default, " +
466 "and lean toward verifying findings adversarially. Work solo only on conversational " +
467 "turns or trivial mechanical edits. When a system message says the mode is off, " +
468 "revert to the opt-in rule above.",
469 ].join("\n\n"),
495470 input_schema: {
496471 type: "object",
497472 properties: {
from line 516
541516 Tool workflowTool = new()
542517 {
543518 Name = "Workflow",
544 Description =
545 "Orchestrate a multiagent workflow: split a large task into independent subtasks "
546 + "and run them as parallel agents, then collect their results.\n\n"
547 + "Opt-in: only use this tool when the user explicitly asks for a workflow, or when a "
548 + "system message confirms that orchestration mode is on.\n\n"
549 + "Quality patterns: adversarial verification (a second wave of agents checks the first "
550 + "wave's findings against the source), a completeness critic (one agent hunts for what "
551 + "the others missed), and multiphase sequencing (understand, design, implement, and "
552 + "review as separate workflow calls, reading results between phases). A useful default "
553 + "is hybrid: scout inline first to discover the work-list, then fan out over it.\n\n"
554 + "Granularity: scope each subtask to a distinct concern, component, or question rather "
555 + "than per line or per file section. Scale the count to what the user asked for: a "
556 + "focused review of a module of a few hundred lines rarely needs more than about ten "
557 + "subtasks; a broad audit of a large codebase can justify more.\n\n"
558 + "Standing consent: while a system message confirms orchestration mode is on, that "
559 + "opt-in is standing. Author and run a workflow for every substantive task by default, "
560 + "and lean toward verifying findings adversarially. Work solo only on conversational "
561 + "turns or trivial mechanical edits. When a system message says the mode is off, "
562 + "revert to the opt-in rule above.",
563 InputSchema = new InputSchema
519 Description = """
520 Orchestrate a multiagent workflow: split a large task into independent subtasks and run them as parallel agents, then collect their results.
521 
522 Opt-in: only use this tool when the user explicitly asks for a workflow, or when a system message confirms that orchestration mode is on.
523 
524 Quality patterns: adversarial verification (a second wave of agents checks the first wave's findings against the source), a completeness critic (one agent hunts for what the others missed), and multiphase sequencing (understand, design, implement, and review as separate workflow calls, reading results between phases). A useful default is hybrid: scout inline first to discover the work-list, then fan out over it.
525 
526 Granularity: scope each subtask to a distinct concern, component, or question rather than per line or per file section. Scale the count to what the user asked for: a focused review of a module of a few hundred lines rarely needs more than about ten subtasks; a broad audit of a large codebase can justify more.
527 
528 Standing consent: while a system message confirms orchestration mode is on, that opt-in is standing. Author and run a workflow for every substantive task by default, and lean toward verifying findings adversarially. Work solo only on conversational turns or trivial mechanical edits. When a system message says the mode is off, revert to the opt-in rule above.
529 """,
530 InputSchema = new()
564531 {
565532 Properties = new Dictionary<string, JsonElement>
566533 {
from line 550
583550 Description =
584551 "Report the final findings for your subtask. Call this exactly once, when you are "
585552 + "done investigating; it ends your task.",
586 InputSchema = new InputSchema
553 InputSchema = new()
587554 {
588555 Properties = new Dictionary<string, JsonElement>
589556 {
from line 657
690657 ```java Java
691658 static final Tool WORKFLOW_TOOL = Tool.builder()
692659 .name("Workflow")
693 .description("Orchestrate a multiagent workflow: split a large task into independent subtasks "
694 + "and run them as parallel agents, then collect their results.\n\n"
695 + "Opt-in: only use this tool when the user explicitly asks for a workflow, or when a "
696 + "system message confirms that orchestration mode is on.\n\n"
697 + "Quality patterns: adversarial verification (a second wave of agents checks the first "
698 + "wave's findings against the source), a completeness critic (one agent hunts for what "
699 + "the others missed), and multiphase sequencing (understand, design, implement, and "
700 + "review as separate workflow calls, reading results between phases). A useful default "
701 + "is hybrid: scout inline first to discover the work-list, then fan out over it.\n\n"
702 + "Granularity: scope each subtask to a distinct concern, component, or question rather "
703 + "than per line or per file section. Scale the count to what the user asked for: a "
704 + "focused review of a module of a few hundred lines rarely needs more than about ten "
705 + "subtasks; a broad audit of a large codebase can justify more.\n\n"
706 + "Standing consent: while a system message confirms orchestration mode is on, that "
707 + "opt-in is standing. Author and run a workflow for every substantive task by default, "
708 + "and lean toward verifying findings adversarially. Work solo only on conversational "
709 + "turns or trivial mechanical edits. When a system message says the mode is off, "
710 + "revert to the opt-in rule above.")
660 .description("""
661 Orchestrate a multiagent workflow: split a large task into independent subtasks \
662 and run them as parallel agents, then collect their results.
663 
664 Opt-in: only use this tool when the user explicitly asks for a workflow, or when a \
665 system message confirms that orchestration mode is on.
666 
667 Quality patterns: adversarial verification (a second wave of agents checks the first \
668 wave's findings against the source), a completeness critic (one agent hunts for what \
669 the others missed), and multiphase sequencing (understand, design, implement, and \
670 review as separate workflow calls, reading results between phases). A useful default \
671 is hybrid: scout inline first to discover the work-list, then fan out over it.
672 
673 Granularity: scope each subtask to a distinct concern, component, or question rather \
674 than per line or per file section. Scale the count to what the user asked for: a \
675 focused review of a module of a few hundred lines rarely needs more than about ten \
676 subtasks; a broad audit of a large codebase can justify more.
677 
678 Standing consent: while a system message confirms orchestration mode is on, that \
679 opt-in is standing. Author and run a workflow for every substantive task by default, \
680 and lean toward verifying findings adversarially. Work solo only on conversational \
681 turns or trivial mechanical edits. When a system message says the mode is off, \
682 revert to the opt-in rule above.""")
711683 .inputSchema(Tool.InputSchema.builder()
712 .properties(JsonValue.from(Map.of(
713 "subtasks", Map.of(
684 .properties(Tool.InputSchema.Properties.builder()
685 .putAdditionalProperty("subtasks", JsonValue.from(Map.of(
714686 "type", "array",
715687 "items", Map.of("type", "string"),
716 "description", "Independent subtask prompts to run as parallel agents"))))
688 "description", "Independent subtask prompts to run as parallel agents")))
689 .build())
717690 .putAdditionalProperty("required", JsonValue.from(List.of("subtasks")))
718691 .build())
719692 .build();
from line 695
722695 
723696 static final Tool REPORT_TOOL = Tool.builder()
724697 .name("report_findings")
725 .description("Report the final findings for your subtask. Call this exactly once, when you are "
726 + "done investigating; it ends your task.")
698 .description("""
699 Report the final findings for your subtask. Call this exactly once, when you are \
700 done investigating; it ends your task.""")
727701 .inputSchema(Tool.InputSchema.builder()
728 .properties(JsonValue.from(Map.of(
729 "summary", Map.of("type", "string", "description", "Two or three sentences of synthesis"),
730 "findings", Map.of(
702 .properties(Tool.InputSchema.Properties.builder()
703 .putAdditionalProperty("summary", JsonValue.from(Map.of(
704 "type", "string",
705 "description", "Two or three sentences of synthesis")))
706 .putAdditionalProperty("findings", JsonValue.from(Map.of(
731707 "type", "array",
732708 "items", Map.of(
733709 "type", "object",
from line 717
741717 "severity", Map.of(
742718 "type", "string",
743719 "enum", List.of("high", "medium", "low", "info"))),
744 "required", List.of("claim", "evidence", "severity"))))))
720 "required", List.of("claim", "evidence", "severity")))))
721 .build())
745722 .putAdditionalProperty("required", JsonValue.from(List.of("summary", "findings")))
746723 .build())
747724 .build();
from line 912
935912 ```
936913 
937914 ```typescript TypeScript
938 const execShell = promisify(exec);
915 type ToolOutcome = { output: string; isError: boolean };
939916 
940917 // Run bash where the example was launched. In DOC_TEST_MODE the docs harness
941 // points it at a throwaway fixture directory instead, removed on exit.
942 const WORK_DIR = DOC_TEST_MODE
943 ? await mkdtemp(join(tmpdir(), "orchestration-"))
944 : process.cwd();
945 if (DOC_TEST_MODE) {
918 // points it at a throwaway fixture directory instead, removed when the program ends.
919 await using fixtureDir = DOC_TEST_MODE
920 ? await mkdtempDisposable(join(tmpdir(), "orchestration-"))
921 : null;
922 const WORK_DIR = fixtureDir?.path ?? process.cwd();
923 if (fixtureDir) {
946924 await writeFile(
947925 join(WORK_DIR, "sample.py"),
948 "def fib(n):\n" +
949 " return n if n < 2 else fib(n - 1) + fib(n - 2)\n\n" +
950 "print(fib(10))\n",
926 `def fib(n):
927 return n if n < 2 else fib(n - 1) + fib(n - 2)
928 
929 print(fib(10))
930 `,
951931 );
952 process.on("exit", () => rmSync(WORK_DIR, { recursive: true, force: true }));
953932 }
954933 
955934 // Run a shell command and return its output. No sandbox: example code only.
956 async function runBash(command: string): Promise<{ output: string; isError: boolean }> {
935 function runBash(command: string): Promise<ToolOutcome> {
957936 console.error(`[bash] ${command}`);
958 let stdout = "";
959 let stderr = "";
960 let exitCode = 0;
961 try {
962 ({ stdout, stderr } = await execShell(command, {
963 shell: "/bin/bash",
964 cwd: WORK_DIR,
965 timeout: BASH_TIMEOUT_SECONDS * 1000,
966 maxBuffer: 16 * 1024 * 1024,
967 }));
968 } catch (error) {
969 const failure = error as {
970 stdout?: string;
971 stderr?: string;
972 code?: number | string;
973 killed?: boolean;
974 };
975 if (failure.killed && failure.code !== "ERR_CHILD_PROCESS_STDIO_MAXBUFFER") {
976 return { output: `command timed out after ${BASH_TIMEOUT_SECONDS}s`, isError: true };
977 }
978 stdout = failure.stdout ?? "";
979 stderr = failure.stderr ?? "";
980 exitCode = typeof failure.code === "number" ? failure.code : 1;
981 }
982 let output = (stdout + stderr).trim() || "(no output)";
983 const codePoints = [...output];
984 if (codePoints.length > TOOL_RESULT_MAX_CHARS) {
985 output =
986 codePoints.slice(0, TOOL_RESULT_MAX_CHARS).join("") +
987 `\n(truncated at ${TOOL_RESULT_MAX_CHARS} chars)`;
988 }
989 if (exitCode !== 0) {
990 output = `(exit code ${exitCode})\n${output}`;
991 }
992 return { output, isError: exitCode !== 0 };
993 }
994 
995 async function handleBashBlock(
996 block: Anthropic.ToolUseBlock,
997 ): Promise<{ output: string; isError: boolean }> {
998 const input = block.input as { command?: string; restart?: boolean };
999 if (input.restart === true) {
937 const options = {
938 shell: "/bin/bash",
939 cwd: WORK_DIR,
940 timeout: BASH_TIMEOUT_SECONDS * 1000,
941 maxBuffer: 16 * 1024 * 1024,
942 };
943 return new Promise((resolve) => {
944 exec(command, options, (error, stdout, stderr) => {
945 if (error?.killed) {
946 resolve({ output: `command timed out after ${BASH_TIMEOUT_SECONDS}s`, isError: true });
947 return;
948 }
949 const exitCode = error === null ? 0 : typeof error.code === "number" ? error.code : 1;
950 let output = (stdout + stderr).trim() || "(no output)";
951 const codePoints = [...output];
952 if (codePoints.length > TOOL_RESULT_MAX_CHARS) {
953 const kept = codePoints.slice(0, TOOL_RESULT_MAX_CHARS).join("");
954 output = `${kept}\n(truncated at ${TOOL_RESULT_MAX_CHARS} chars)`;
955 }
956 if (exitCode !== 0) {
957 output = `(exit code ${exitCode})\n${output}`;
958 }
959 resolve({ output, isError: exitCode !== 0 });
960 });
961 });
962 }
963 
964 async function handleBashBlock(block: Anthropic.ToolUseBlock): Promise<ToolOutcome> {
965 const { command, restart } = block.input as { command?: string; restart?: boolean };
966 if (restart === true) {
1000967 return { output: "Shell restarted.", isError: false };
1001968 }
1002 if (!input.command) {
969 if (!command) {
1003970 return { output: "bash error: no command was provided.", isError: true };
1004971 }
1005 return runBash(input.command);
972 return runBash(command);
1006973 }
1007974 ```
1008975 
from line 980
1013980 if (docTestMode)
1014981 {
1015982 workDir = Directory.CreateTempSubdirectory("orchestration-").FullName;
1016 File.WriteAllText(Path.Combine(workDir, "sample.py"),
1017 "def fib(n):\n" +
1018 " return n if n < 2 else fib(n - 1) + fib(n - 2)\n\n" +
1019 "print(fib(10))\n");
983 await File.WriteAllTextAsync(Path.Combine(workDir, "sample.py"), """
984 def fib(n):
985 return n if n < 2 else fib(n - 1) + fib(n - 2)
986 
987 print(fib(10))
988 
989 """);
1020990 var fixtureDir = workDir;
1021991 AppDomain.CurrentDomain.ProcessExit += (_, _) =>
1022992 {
from line 1050
10801050 // Execute one bash tool call requested by the model.
10811051 async Task<(string Output, bool IsError)> HandleBashBlock(ToolUseBlock block)
10821052 {
1083 if (block.Input.TryGetValue("restart", out var restart) && restart.ValueKind == JsonValueKind.True)
1053 if (block.Input.GetValueOrDefault("restart").ValueKind is JsonValueKind.True)
10841054 {
10851055 return ("Shell restarted.", false);
10861056 }
1087 var command = block.Input.TryGetValue("command", out var rawCommand) && rawCommand.ValueKind == JsonValueKind.String
1088 ? rawCommand.GetString()!
1089 : "";
1090 if (command.Length == 0)
1091 {
1092 return ("bash error: no command was provided.", true);
1093 }
1094 return await RunBash(command);
1057 return block.Input.GetValueOrDefault("command") is { ValueKind: JsonValueKind.String } command
1058 && command.GetString() is { Length: > 0 } commandText
1059 ? await RunBash(commandText)
1060 : ("bash error: no command was provided.", true);
10951061 }
10961062 ```
10971063 
from line 1076
11101076 if err != nil {
11111077 log.Fatal(err)
11121078 }
1113 fixture := "def fib(n):\n" +
1114 " return n if n < 2 else fib(n - 1) + fib(n - 2)\n\n" +
1115 "print(fib(10))\n"
1079 const fixture = `def fib(n):
1080 return n if n < 2 else fib(n - 1) + fib(n - 2)
1081 
1082 print(fib(10))
1083 `
11161084 if err := os.WriteFile(filepath.Join(dir, "sample.py"), []byte(fixture), 0o644); err != nil {
11171085 log.Fatal(err)
11181086 }
from line 1109
11411109 if err == nil {
11421110 return output, false
11431111 }
1144 var exitErr *exec.ExitError
1145 if errors.As(err, &exitErr) {
1112 if exitErr, ok := errors.AsType[*exec.ExitError](err); ok {
11461113 return fmt.Sprintf("(exit code %d)\n%s", exitErr.ExitCode(), output), true
11471114 }
11481115 return fmt.Sprintf("(%s)\n%s", err, output), true
from line 1155
11881155 print(fib(10))
11891156 """);
11901157 Runtime.getRuntime().addShutdownHook(new Thread(() -> {
1158 // Best-effort cleanup; the OS tmp sweeper handles leftovers.
11911159 try (var paths = Files.walk(dir)) {
1192 paths.sorted(Comparator.reverseOrder()).forEach(p -> {
1193 try { Files.deleteIfExists(p); } catch (IOException ignored) {}
1160 paths.sorted(Comparator.reverseOrder()).forEach(path -> {
1161 try { Files.deleteIfExists(path); } catch (IOException _) {}
11941162 });
1195 } catch (IOException ignored) {
1196 // Best-effort cleanup; the OS tmp sweeper handles leftovers.
1197 }
1163 } catch (IOException _) {}
11981164 }));
11991165 return dir;
12001166 } catch (IOException error) {
from line 1184
12181184 CompletableFuture<String> outputReader = CompletableFuture.supplyAsync(() -> {
12191185 try (var stdout = process.getInputStream()) {
12201186 return new String(stdout.readAllBytes(), StandardCharsets.UTF_8);
1221 } catch (IOException error) {
1187 } catch (IOException _) {
12221188 return "";
12231189 }
12241190 });
from line 1193
12271193 outputReader.cancel(true);
12281194 return new ToolOutput("command timed out after " + BASH_TIMEOUT_SECONDS + "s", true);
12291195 }
1230 String output = outputReader.join().trim();
1196 String output = outputReader.join().strip();
12311197 if (output.isEmpty()) {
12321198 output = "(no output)";
1233 }
1234 if (output.length() > TOOL_RESULT_MAX_CHARS) {
1199 } else if (output.length() > TOOL_RESULT_MAX_CHARS) {
12351200 output = output.substring(0, TOOL_RESULT_MAX_CHARS)
12361201 + "\n(truncated at " + TOOL_RESULT_MAX_CHARS + " chars)";
12371202 }
12381203 int exitCode = process.exitValue();
1239 if (exitCode != 0) {
1240 return new ToolOutput("(exit code " + exitCode + ")\n" + output, true);
1241 }
1242 return new ToolOutput(output, false);
1204 return exitCode == 0
1205 ? new ToolOutput(output, false)
1206 : new ToolOutput("(exit code " + exitCode + ")\n" + output, true);
1207 }
1208 
1209 // A tool call's input fields, or an empty map if the model sent something other than an object.
1210 Map<String, JsonValue> toolInput(ToolUseBlock toolUse) {
1211 return toolUse._input() instanceof JsonObject object ? object.values() : Map.of();
12431212 }
12441213 
12451214 // Execute one bash tool call requested by the model.
1246 ToolOutput handleBashBlock(ToolUseBlock block) throws InterruptedException {
1247 Map<String, JsonValue> input = (Map<String, JsonValue>) block._input().asObject().orElse(Map.of());
1248 JsonValue restart = input.getOrDefault("restart", JsonValue.from(false));
1249 if (Boolean.TRUE.equals(restart.asBoolean().orElse(false))) {
1215 ToolOutput handleBashBlock(ToolUseBlock toolUse) throws InterruptedException {
1216 Map<String, JsonValue> input = toolInput(toolUse);
1217 if (input.get("restart") instanceof JsonBoolean restart && restart.value()) {
12501218 return new ToolOutput("Shell restarted.", false);
12511219 }
1252 JsonValue raw = input.get("command");
1253 String command = raw != null && raw.asString().isPresent() ? raw.asStringOrThrow() : "";
1220 String command = input.get("command") instanceof JsonString text ? text.value() : "";
12541221 if (command.isEmpty()) {
12551222 return new ToolOutput("bash error: no command was provided.", true);
12561223 }
from line 1486
15191486 }
15201487 if (response.stop_reason !== "tool_use") {
15211488 let text = response.content
1522 .filter((block): block is Anthropic.TextBlock => block.type === "text")
1523 .map((block) => block.text)
1489 .flatMap((block) => (block.type === "text" ? [block.text] : []))
15241490 .join("");
15251491 if (response.stop_reason === "max_tokens") {
15261492 text += "\n\n(warning: subagent response was truncated at max_tokens)";
from line 1499
15331499 if (block.type !== "tool_use") {
15341500 continue;
15351501 }
1536 let output: string;
1537 let isError: boolean;
1502 let outcome: ToolOutcome;
15381503 switch (block.name) {
15391504 case "report_findings":
15401505 report = JSON.stringify(block.input, null, 2);
1541 output = "Findings recorded.";
1542 isError = false;
1506 outcome = { output: "Findings recorded.", isError: false };
15431507 break;
15441508 case "bash":
1545 ({ output, isError } = await handleBashBlock(block));
1509 outcome = await handleBashBlock(block);
15461510 break;
15471511 default:
1548 output = `unknown tool: ${block.name}`;
1549 isError = true;
1512 outcome = { output: `unknown tool: ${block.name}`, isError: true };
15501513 }
15511514 toolResults.push({
15521515 type: "tool_result",
15531516 tool_use_id: block.id,
1554 content: output,
1555 is_error: isError,
1517 content: outcome.output,
1518 is_error: outcome.isError,
15561519 });
15571520 }
15581521 if (report !== null) {
from line 1530
15671530 ```csharp C#
15681531 // One subagent: a small nested agent loop with the bash tool plus report_findings.
15691532 // Subagents inherit the main loop's effort level.
1533 JsonSerializerOptions indented = new() { WriteIndented = true };
1534 
15701535 async Task<string> RunSubagent(string prompt)
15711536 {
15721537 const string subagentSystem =
from line 1542
15771542 for (var turn = 0; turn < maxSubagentTurns; turn++)
15781543 {
15791544 using var deadline = new CancellationTokenSource(TimeSpan.FromSeconds(requestTimeoutSeconds));
1580 var response = await client.Messages.Create(new MessageCreateParams
1545 var response = await client.Messages.CreateStreaming(new MessageCreateParams
15811546 {
15821547 Model = model,
1583 MaxTokens = requestMaxTokens,
1548 MaxTokens = 64000,
15841549 System = subagentSystem,
15851550 OutputConfig = new OutputConfig { Effort = effort },
15861551 Tools = [bashTool, reportTool],
15871552 Messages = messages,
1588 }, cancellationToken: deadline.Token);
1553 }, cancellationToken: deadline.Token).Aggregate();
15891554 messages.Add(new()
15901555 {
15911556 Role = Role.Assistant,
from line 1578
16131578 {
16141579 continue;
16151580 }
1616 string output;
1617 bool isError;
16181581 if (toolUse.Name == "report_findings")
16191582 {
1620 report = JsonSerializer.Serialize(
1621 toolUse.Input, new JsonSerializerOptions { WriteIndented = true });
1622 output = "Findings recorded.";
1623 isError = false;
1583 report = JsonSerializer.Serialize(toolUse.Input, indented);
16241584 }
1625 else if (toolUse.Name == "bash")
1585 var (output, isError) = toolUse.Name switch
16261586 {
1627 (output, isError) = await HandleBashBlock(toolUse);
1628 }
1629 else
1630 {
1631 output = $"unknown tool: {toolUse.Name}";
1632 isError = true;
1633 }
1587 "report_findings" => ("Findings recorded.", false),
1588 "bash" => await HandleBashBlock(toolUse),
1589 _ => ($"unknown tool: {toolUse.Name}", true),
1590 };
16341591 toolResults.Add(new ToolResultBlockParam(toolUse.ID) { Content = output, IsError = isError });
16351592 }
16361593 if (report is not null)
from line 1601
16441601 ```
16451602 
16461603 ```go Go
1604 // streamMessage sends one request and accumulates the streamed events into the
1605 // final message.
1606 func streamMessage(ctx context.Context, params anthropic.MessageNewParams) (anthropic.Message, error) {
1607 ctx, cancel := context.WithTimeout(ctx, requestTimeoutSeconds*time.Second)
1608 defer cancel()
1609 stream := client.Messages.NewStreaming(ctx, params)
1610 defer stream.Close()
1611 var message anthropic.Message
1612 for stream.Next() {
1613 if err := message.Accumulate(stream.Current()); err != nil {
1614 return message, err
1615 }
1616 }
1617 return message, stream.Err()
1618 }
1619 
1620 // textOf joins the text blocks of a response.
1621 func textOf(message anthropic.Message) string {
1622 var text strings.Builder
1623 for _, block := range message.Content {
1624 if textBlock, ok := block.AsAny().(anthropic.TextBlock); ok {
1625 text.WriteString(textBlock.Text)
1626 }
1627 }
1628 return text.String()
1629 }
1630 
16471631 // runSubagent runs one subagent: a small nested agent loop with the bash tool plus
16481632 // report_findings. Subagents inherit the main loop's effort level.
16491633 func runSubagent(ctx context.Context, model string, prompt string) (string, error) {
1650 subagentSystem := "You are one agent in a larger parallel fan-out, assigned a single subtask. " +
1634 const subagentSystem = "You are one agent in a larger parallel fan-out, assigned a single subtask. " +
16511635 "Investigate it directly, using bash to check facts rather than guessing, and finish " +
16521636 "by calling report_findings exactly once. Return findings, not narration."
16531637 messages := []anthropic.MessageParam{anthropic.NewUserMessage(anthropic.NewTextBlock(prompt))}
16541638 for range maxSubagentTurns {
1655 var response anthropic.Message
1656 err := func() error {
1657 ctx, cancel := context.WithTimeout(ctx, requestTimeoutSeconds*time.Second)
1658 defer cancel()
1659 stream := client.Messages.NewStreaming(ctx, anthropic.MessageNewParams{
1660 Model: model,
1661 MaxTokens: 64000,
1662 System: []anthropic.TextBlockParam{{Text: subagentSystem}},
1663 OutputConfig: anthropic.OutputConfigParam{Effort: effort},
1664 Tools: []anthropic.ToolUnionParam{bashTool, reportTool},
1665 Messages: messages,
1666 })
1667 defer stream.Close()
1668 for stream.Next() {
1669 if err := response.Accumulate(stream.Current()); err != nil {
1670 return err
1671 }
1672 }
1673 return stream.Err()
1674 }()
1639 response, err := streamMessage(ctx, anthropic.MessageNewParams{
1640 Model: model,
1641 MaxTokens: 64000,
1642 System: []anthropic.TextBlockParam{{Text: subagentSystem}},
1643 OutputConfig: anthropic.OutputConfigParam{Effort: effort},
1644 Tools: []anthropic.ToolUnionParam{bashTool, reportTool},
1645 Messages: messages,
1646 })
16751647 if err != nil {
16761648 return "", err
16771649 }
16781650 messages = append(messages, response.ToParam())
1679 if response.StopReason == anthropic.StopReasonPauseTurn {
1651 
1652 switch response.StopReason {
1653 case anthropic.StopReasonPauseTurn:
16801654 continue
1655 case anthropic.StopReasonMaxTokens:
1656 return textOf(response) + "\n\n(warning: subagent response was truncated at max_tokens)", nil
1657 case anthropic.StopReasonToolUse:
1658 default:
1659 return textOf(response), nil
16811660 }
1682 if response.StopReason != anthropic.StopReasonToolUse {
1683 var text strings.Builder
1684 for _, block := range response.Content {
1685 if textBlock, ok := block.AsAny().(anthropic.TextBlock); ok {
1686 text.WriteString(textBlock.Text)
1687 }
1688 }
1689 if response.StopReason == anthropic.StopReasonMaxTokens {
1690 text.WriteString("\n\n(warning: subagent response was truncated at max_tokens)")
1691 }
1692 return text.String(), nil
1693 }
1661 
16941662 var toolResults []anthropic.ContentBlockParamUnion
16951663 var report string
16961664 var reportRecorded bool
from line 1681
17131681 case "bash":
17141682 output, isError = handleBashBlock(ctx, toolUse)
17151683 default:
1716 output, isError = fmt.Sprintf("unknown tool: %s", toolUse.Name), true
1684 output, isError = "unknown tool: "+toolUse.Name, true
17171685 }
17181686 toolResults = append(toolResults, anthropic.NewToolResultBlock(toolUse.ID, output, isError))
17191687 }
from line 1696
17281696 ```
17291697 
17301698 ```java Java
1699 static final String SUBAGENT_SYSTEM = """
1700 You are one agent in a larger parallel fan-out, assigned a single subtask. \
1701 Investigate it directly, using bash to check facts rather than guessing, and finish \
1702 by calling report_findings exactly once. Return findings, not narration.""";
1703 
17311704 // One subagent: a small nested agent loop with the bash tool plus report_findings.
17321705 // Subagents inherit the main loop's effort level.
17331706 String runSubagent(Model model, String prompt) throws InterruptedException {
1734 String subagentSystem = "You are one agent in a larger parallel fan-out, assigned a single subtask. "
1735 + "Investigate it directly, using bash to check facts rather than guessing, and finish "
1736 + "by calling report_findings exactly once. Return findings, not narration.";
17371707 List<MessageParam> messages = new ArrayList<>();
17381708 messages.add(MessageParam.builder().role(MessageParam.Role.USER).content(prompt).build());
17391709 for (int turn = 0; turn < MAX_SUBAGENT_TURNS; turn++) {
17401710 MessageCreateParams params = MessageCreateParams.builder()
17411711 .model(model)
17421712 .maxTokens(64000L)
1743 .system(subagentSystem)
1713 .system(SUBAGENT_SYSTEM)
17441714 .outputConfig(OutputConfig.builder().effort(EFFORT).build())
17451715 .addTool(BASH_TOOL)
17461716 .addTool(REPORT_TOOL)
from line 1736
17661736 }
17671737 return text;
17681738 }
1739 List<ToolUseBlock> toolUses = response.content().stream()
1740 .flatMap(block -> block.toolUse().stream())
1741 .toList();
17691742 List<ContentBlockParam> toolResults = new ArrayList<>();
17701743 String report = null;
1771 for (ContentBlock block : response.content()) {
1772 if (block.toolUse().isEmpty()) {
1773 continue;
1774 }
1775 ToolUseBlock toolUse = block.toolUse().get();
1776 ToolOutput result;
1777 if (toolUse.name().equals("report_findings")) {
1778 report = toolUse._input().convert(JsonNode.class).toPrettyString();
1779 result = new ToolOutput("Findings recorded.", false);
1780 } else if (toolUse.name().equals("bash")) {
1781 result = handleBashBlock(toolUse);
1782 } else {
1783 result = new ToolOutput("unknown tool: " + toolUse.name(), true);
1784 }
1744 for (ToolUseBlock toolUse : toolUses) {
1745 ToolOutput result = switch (toolUse.name()) {
1746 case "report_findings" -> {
1747 report = toolUse._input().convert(JsonNode.class).toPrettyString();
1748 yield new ToolOutput("Findings recorded.", false);
1749 }
1750 case "bash" -> handleBashBlock(toolUse);
1751 default -> new ToolOutput("unknown tool: " + toolUse.name(), true);
1752 };
17851753 toolResults.add(ContentBlockParam.ofToolResult(ToolResultBlockParam.builder()
17861754 .toolUseId(toolUse.id())
17871755 .content(result.output())
from line 2041
20732041 // that never finished are recomputed. Delete the journal file to start fresh.
20742042 async Task<string> Journaled(string prompt, Func<Task<string>> compute)
20752043 {
2076 var key = Convert.ToHexString(SHA256.HashData(Encoding.UTF8.GetBytes(prompt))).ToLowerInvariant();
2044 var key = Convert.ToHexStringLower(SHA256.HashData(Encoding.UTF8.GetBytes(prompt)));
20772045 if ((await LoadJournal()).TryGetValue(key, out var cached))
20782046 {
20792047 Console.Error.WriteLine($"[journal] cache hit for {key[..12]}");
from line 2125
21572125 return Objects.requireNonNullElseGet(
21582126 JOURNAL_MAPPER.readValue(Files.readString(JOURNAL_PATH), new TypeReference<HashMap<String, String>>() {}),
21592127 HashMap::new);
2160 } catch (IOException error) {
2128 } catch (IOException _) {
21612129 return new HashMap<>();
21622130 }
21632131 }
from line 2258
22902258 "that contradicts them. Default to refuted if uncertain. Call report_findings with "
22912259 "summary 'refuted: <why>' or 'confirmed: <why>', citing the file:line or command "
22922260 "output that decided it.\n\n"
2293 f"Subtask: {subtask}\n\nResult to verify:\n{result}"
2261 f"""Subtask: {subtask}
2262 
2263 Result to verify:
2264 {result}"""
22942265 )
22952266 
22962267 
from line 2289
23182289 verdicts = list(pool.map(run_one, verify_prompts))
23192290 
23202291 joined = "\n\n".join(
2321 f"[agent {index + 1}: {task}]\n{result}\n\n[verify {index + 1}]\n{verdict}"
2292 f"""[agent {index + 1}: {task}]
2293 {result}
2294 
2295 [verify {index + 1}]
2296 {verdict}"""
23222297 for index, (task, result, verdict) in enumerate(zip(subtasks, results, verdicts))
23232298 )
23242299 if dropped > 0:
from line 2326
23512326 }
23522327 
23532328 function verifyPromptFor(subtask: string, result: string): string {
2354 return (
2329 const instructions =
23552330 "Adversarially verify the subagent result below: try to REFUTE it. Re-derive the " +
23562331 "claims yourself with bash rather than trusting the result, and look for evidence " +
23572332 "that contradicts them. Default to refuted if uncertain. Call report_findings with " +
23582333 "summary 'refuted: <why>' or 'confirmed: <why>', citing the file:line or command " +
2359 "output that decided it.\n\n" +
2360 `Subtask: ${subtask}\n\nResult to verify:\n${result}`
2361 );
2334 "output that decided it.";
2335 return `${instructions}
2336 
2337 Subtask: ${subtask}
2338 
2339 Result to verify:
2340 ${result}`;
23622341 }
23632342 
23642343 // Map with a concurrency limit: at most `limit` tasks are in flight at once.
from line 2361
23822361 // Run subtasks as parallel subagents, then run a second verification wave over
23832362 // the results, and return both. MAX_TOTAL_SUBTASKS bounds how many the model can
23842363 // queue; MAX_CONCURRENT bounds how many run at once.
2385 async function runWorkflow(
2386 model: string,
2387 rawSubtasks: unknown,
2388 ): Promise<{ output: string; isError: boolean }> {
2364 async function runWorkflow(model: string, rawSubtasks: unknown): Promise<ToolOutcome> {
23892365 const allSubtasks = normalizeSubtasks(rawSubtasks);
23902366 const subtasks = allSubtasks.slice(0, MAX_TOTAL_SUBTASKS);
23912367 const dropped = allSubtasks.length - subtasks.length;
from line 2385
24092385 const verifyPrompts = subtasks.map((task, index) => verifyPromptFor(task, results[index]));
24102386 const verdicts = await mapWithLimit(verifyPrompts, MAX_CONCURRENT, runOne);
24112387 
2412 let joined = subtasks
2388 const report = subtasks
24132389 .map(
2414 (task, index) =>
2415 `[agent ${index + 1}: ${task}]\n${results[index]}\n\n[verify ${index + 1}]\n${verdicts[index]}`,
2390 (task, index) => `[agent ${index + 1}: ${task}]
2391 ${results[index]}
2392 
2393 [verify ${index + 1}]
2394 ${verdicts[index]}`,
24162395 )
24172396 .join("\n\n");
2418 if (dropped > 0) {
2419 joined =
2420 `(note: ${dropped} subtasks beyond MAX_TOTAL_SUBTASKS=${MAX_TOTAL_SUBTASKS} were not ` +
2421 "run; rerun them in a follow-up Workflow call)\n\n" +
2422 joined;
2423 }
2424 return { output: joined, isError: false };
2397 const droppedNote =
2398 dropped > 0
2399 ? `(note: ${dropped} subtasks beyond MAX_TOTAL_SUBTASKS=${MAX_TOTAL_SUBTASKS} were ` +
2400 "not run; rerun them in a follow-up Workflow call)\n\n"
2401 : "";
2402 return { output: droppedNote + report, isError: false };
24252403 }
24262404 ```
24272405 
from line 2408
24302408 // JSON-encoded as a single string, or a newline-separated list.
24312409 List<string> NormalizeSubtasks(JsonElement raw)
24322410 {
2433 List<string> tasks = [];
2434 if (raw.ValueKind == JsonValueKind.Array)
2411 IEnumerable<string?> tasks = raw.ValueKind switch
24352412 {
2436 tasks = raw.EnumerateArray()
2413 JsonValueKind.Array => raw.EnumerateArray()
24372414 .Where(item => item.ValueKind == JsonValueKind.String)
2438 .Select(item => item.GetString()!)
2439 .ToList();
2440 }
2441 else if (raw.ValueKind == JsonValueKind.String)
2415 .Select(item => item.GetString()),
2416 JsonValueKind.String => ParseSubtaskString(raw.GetString()!),
2417 _ => [],
2418 };
2419 return [.. tasks.OfType<string>().Select(task => task.Trim()).Where(task => task.Length > 0)];
2420 }
2421 
2422 IEnumerable<string?> ParseSubtaskString(string single)
2423 {
2424 try
24422425 {
2443 var single = raw.GetString()!;
2444 try
2445 {
2446 tasks = JsonSerializer.Deserialize<List<string>>(single) ?? [];
2447 }
2448 catch (JsonException)
2449 {
2450 tasks = [.. single.Split('\n')];
2451 }
2452 }
2453 return tasks.Where(task => task != null).Select(task => task.Trim()).Where(task => task.Length > 0).ToList();
2454 }
2455 
2456 string VerifyPromptFor(string subtask, string result) =>
2457 "Adversarially verify the subagent result below: try to REFUTE it. Re-derive the "
2458 + "claims yourself with bash rather than trusting the result, and look for evidence "
2459 + "that contradicts them. Default to refuted if uncertain. Call report_findings with "
2460 + "summary 'refuted: <why>' or 'confirmed: <why>', citing the file:line or command "
2461 + "output that decided it.\n\n"
2462 + $"Subtask: {subtask}\n\nResult to verify:\n{result}";
2426 return JsonSerializer.Deserialize<List<string?>>(single) ?? [];
2427 }
2428 catch (JsonException)
2429 {
2430 return single.Split('\n');
2431 }
2432 }
2433 
2434 string VerifyPromptFor(string subtask, string result) => $"""
2435 Adversarially verify the subagent result below: try to REFUTE it. Re-derive the claims yourself with bash rather than trusting the result, and look for evidence that contradicts them. Default to refuted if uncertain. Call report_findings with summary 'refuted: <why>' or 'confirmed: <why>', citing the file:line or command output that decided it.
2436 
2437 Subtask: {subtask}
2438 
2439 Result to verify:
2440 {result}
2441 """;
24632442 
24642443 // Run subtasks as parallel subagents, then run a second verification wave over
24652444 // the results, and return both. maxTotalSubtasks bounds how many the model can
from line 2446
24672446 async Task<(string Output, bool IsError)> RunWorkflow(JsonElement rawSubtasks)
24682447 {
24692448 var allSubtasks = NormalizeSubtasks(rawSubtasks);
2470 var subtasks = allSubtasks.Take(maxTotalSubtasks).ToList();
2449 List<string> subtasks = [.. allSubtasks.Take(maxTotalSubtasks)];
24712450 var dropped = allSubtasks.Count - subtasks.Count;
24722451 if (subtasks.Count == 0)
24732452 {
from line 2475
24962475 
24972476 var results = await Task.WhenAll(subtasks.Select(RunOne));
24982477 Console.Error.WriteLine($"[workflow] verifying {results.Length} results");
2499 var verifyPrompts = subtasks.Select((task, index) => VerifyPromptFor(task, results[index])).ToList();
2500 var verdicts = await Task.WhenAll(verifyPrompts.Select(RunOne));
2501 
2502 var joined = string.Join(
2503 "\n\n",
2504 subtasks.Select((task, index) =>
2505 $"[agent {index + 1}: {task}]\n{results[index]}\n\n[verify {index + 1}]\n{verdicts[index]}"));
2478 var verdicts = await Task.WhenAll(subtasks.Zip(results, VerifyPromptFor).Select(RunOne));
2479 
2480 var joined = string.Join("\n\n", subtasks.Select((task, index) => $"""
2481 [agent {index + 1}: {task}]
2482 {results[index]}
2483 
2484 [verify {index + 1}]
2485 {verdicts[index]}
2486 """));
25062487 if (dropped > 0)
25072488 {
2508 joined = $"(note: {dropped} subtasks beyond maxTotalSubtasks={maxTotalSubtasks} were not run; "
2509 + "rerun them in a follow-up Workflow call)\n\n" + joined;
2489 joined = $"""
2490 (note: {dropped} subtasks beyond maxTotalSubtasks={maxTotalSubtasks} were not run; rerun them in a follow-up Workflow call)
2491 
2492 {joined}
2493 """;
25102494 }
25112495 return (joined, false);
25122496 }
from line 2525
25412525 "that contradicts them. Default to refuted if uncertain. Call report_findings with " +
25422526 "summary 'refuted: <why>' or 'confirmed: <why>', citing the file:line or command " +
25432527 "output that decided it.\n\n" +
2544 "Subtask: " + subtask + "\n\nResult to verify:\n" + result
2528 fmt.Sprintf(`Subtask: %s
2529 
2530 Result to verify:
2531 %s`, subtask, result)
25452532 }
25462533 
25472534 // mapWithLimit runs task over items with at most limit goroutines in flight.
from line 2537
25502537 semaphore := make(chan struct{}, limit)
25512538 var waitGroup sync.WaitGroup
25522539 for index, item := range items {
2553 waitGroup.Add(1)
25542540 semaphore <- struct{}{}
2555 go func() {
2556 defer waitGroup.Done()
2541 waitGroup.Go(func() {
25572542 defer func() { <-semaphore }()
25582543 results[index] = task(item)
2559 }()
2544 })
25602545 }
25612546 waitGroup.Wait()
25622547 return results
from line 2552
25672552 // queue; maxConcurrent bounds how many run at once.
25682553 func runWorkflow(ctx context.Context, model string, rawSubtasks json.RawMessage) (string, bool) {
25692554 allSubtasks := normalizeSubtasks(rawSubtasks)
2570 subtasks := allSubtasks
2571 if len(subtasks) > maxTotalSubtasks {
2572 subtasks = subtasks[:maxTotalSubtasks]
2573 }
2555 subtasks := allSubtasks[:min(len(allSubtasks), maxTotalSubtasks)]
25742556 dropped := len(allSubtasks) - len(subtasks)
25752557 if len(subtasks) == 0 {
25762558 return "Workflow error: no usable subtasks were provided.", true
from line 2578
25962578 
25972579 sections := make([]string, len(subtasks))
25982580 for index, task := range subtasks {
2599 sections[index] = fmt.Sprintf("[agent %d: %s]\n%s\n\n[verify %d]\n%s",
2581 sections[index] = fmt.Sprintf(`[agent %d: %s]
2582 %s
2583 
2584 [verify %d]
2585 %s`,
26002586 index+1, task, results[index], index+1, verdicts[index])
26012587 }
26022588 joined := strings.Join(sections, "\n\n")
from line 2598
26122598 ```java Java
26132599 // Accept the subtasks input in whatever shape the model emits: an array, the array
26142600 // JSON-encoded as a single string, or a newline-separated list.
2615 List<String> normalizeSubtasks(JsonValue raw) {
2616 List<String> tasks = new ArrayList<>();
2617 if (raw.asArray().isPresent()) {
2618 for (JsonValue item : (List<JsonValue>) raw.asArray().get()) {
2619 tasks.add(item.asString().isPresent() ? item.asStringOrThrow() : item.toString());
2620 }
2621 } else if (raw.asString().isPresent()) {
2622 String single = raw.asStringOrThrow();
2623 try {
2624 String[] parsed = new ObjectMapper().readValue(single, String[].class);
2625 if (parsed != null) {
2626 for (String task : parsed) {
2627 tasks.add(task);
2628 }
2629 }
2630 } catch (JsonProcessingException error) {
2631 for (String task : single.split("\n")) {
2632 tasks.add(task);
2633 }
2634 }
2635 }
2601 List<String> normalizeSubtasks(JsonValue rawSubtasks) {
2602 List<String> tasks = switch (rawSubtasks) {
2603 case JsonArray array -> array.values().stream()
2604 .map(item -> item instanceof JsonString text ? text.value() : item.toString())
2605 .toList();
2606 case JsonString string -> parseSubtaskString(string.value());
2607 default -> List.of();
2608 };
26362609 return tasks.stream()
2637 .filter(task -> task != null)
2638 .map(String::trim)
2610 .filter(Objects::nonNull)
2611 .map(String::strip)
26392612 .filter(task -> !task.isEmpty())
26402613 .toList();
26412614 }
26422615 
2616 List<String> parseSubtaskString(String encoded) {
2617 try {
2618 String[] parsed = new ObjectMapper().readValue(encoded, String[].class);
2619 return parsed == null ? List.of() : Arrays.asList(parsed);
2620 } catch (JsonProcessingException _) {
2621 return encoded.lines().toList();
2622 }
2623 }
2624 
26432625 String verifyPromptFor(String subtask, String result) {
2644 return "Adversarially verify the subagent result below: try to REFUTE it. Re-derive the "
2645 + "claims yourself with bash rather than trusting the result, and look for evidence "
2646 + "that contradicts them. Default to refuted if uncertain. Call report_findings with "
2647 + "summary 'refuted: <why>' or 'confirmed: <why>', citing the file:line or command "
2648 + "output that decided it.\n\n"
2649 + "Subtask: " + subtask + "\n\nResult to verify:\n" + result;
2650 }
2651 
2626 return """
2627 Adversarially verify the subagent result below: try to REFUTE it. Re-derive the \
2628 claims yourself with bash rather than trusting the result, and look for evidence \
2629 that contradicts them. Default to refuted if uncertain. Call report_findings with \
2630 summary 'refuted: <why>' or 'confirmed: <why>', citing the file:line or command \
2631 output that decided it.
2632 
2633 Subtask: %s
2634 
2635 Result to verify:
2636 %s""".formatted(subtask, result);
2637 }
2638 
2639 // invokeAll waits for every job, so each future is finished when it is read here.
26522640 List<String> runAll(ExecutorService pool, List<String> prompts, Model model) throws InterruptedException {
26532641 List<Callable<String>> jobs = prompts.stream()
26542642 .<Callable<String>>map(prompt -> () -> journaled(prompt, () -> runSubagent(model, prompt)))
26552643 .toList();
2656 List<String> results = new ArrayList<>();
2657 for (Future<String> future : pool.invokeAll(jobs)) {
2658 try {
2659 results.add(future.get());
2660 } catch (ExecutionException | CancellationException error) {
2661 // Isolation boundary: one bad subagent should not end the run.
2662 Throwable cause = error.getCause() != null ? error.getCause() : error;
2663 results.add("(subagent failed: " + cause + ")");
2664 }
2665 }
2666 return results;
2644 return pool.invokeAll(jobs).stream()
2645 .map(future -> switch (future.state()) {
2646 case SUCCESS -> future.resultNow();
2647 // Isolation boundary: one bad subagent should not end the run.
2648 case FAILED -> "(subagent failed: " + future.exceptionNow() + ")";
2649 case CANCELLED, RUNNING -> "(subagent failed: cancelled)";
2650 })
2651 .toList();
26672652 }
26682653 
26692654 // Run subtasks as parallel subagents, then run a second verification wave over
from line 2674
26892674 verdicts = runAll(pool, verifyPrompts, model);
26902675 }
26912676 String joined = IntStream.range(0, subtasks.size())
2692 .mapToObj(index -> "[agent " + (index + 1) + ": " + subtasks.get(index) + "]\n" + results.get(index)
2693 + "\n\n[verify " + (index + 1) + "]\n" + verdicts.get(index))
2677 .mapToObj(index -> """
2678 [agent %d: %s]
2679 %s
2680 
2681 [verify %d]
2682 %s""".formatted(index + 1, subtasks.get(index), results.get(index), index + 1, verdicts.get(index)))
26942683 .collect(Collectors.joining("\n\n"));
26952684 if (dropped > 0) {
2696 joined = "(note: " + dropped + " subtasks beyond MAX_TOTAL_SUBTASKS=" + MAX_TOTAL_SUBTASKS
2697 + " were not run; rerun them in a follow-up Workflow call)\n\n" + joined;
2685 joined = """
2686 (note: %d subtasks beyond MAX_TOTAL_SUBTASKS=%d were not run; rerun them in a follow-up Workflow call)
2687 
2688 %s""".formatted(dropped, MAX_TOTAL_SUBTASKS, joined);
26982689 }
26992690 return new ToolOutput(joined, false);
27002691 }
from line 2719
27282719 . 'that contradicts them. Default to refuted if uncertain. Call report_findings with '
27292720 . "summary 'refuted: <why>' or 'confirmed: <why>', citing the file:line or command "
27302721 . "output that decided it.\n\n"
2731 . "Subtask: {$subtask}\n\nResult to verify:\n{$result}";
2722 . <<<PROMPT
2723 Subtask: {$subtask}
2724 
2725 Result to verify:
2726 {$result}
2727 PROMPT;
27322728 }
27332729 
27342730 /**
from line 2791
27952791 "claims yourself with bash rather than trusting the result, and look for evidence " \
27962792 "that contradicts them. Default to refuted if uncertain. Call report_findings with " \
27972793 "summary 'refuted: <why>' or 'confirmed: <why>', citing the file:line or command " \
2798 "output that decided it.\n\n" \
2799 "Subtask: #{subtask}\n\nResult to verify:\n#{result}"
2794 "output that decided it.\n\n" +
2795 <<~PROMPT.chomp
2796 Subtask: #{subtask}
2797 
2798 Result to verify:
2799 #{result}
2800 PROMPT
28002801 end
28012802 
28022803 # Map with a concurrency limit: at most `limit` threads are in flight at once.
from line 2839
28382839 verdicts = map_with_limit(verify_prompts, MAX_CONCURRENT, &run_one)
28392840 
28402841 joined = subtasks.each_with_index.map do |task, index|
2841 "[agent #{index + 1}: #{task}]\n#{results[index]}\n\n[verify #{index + 1}]\n#{verdicts[index]}"
2842 <<~ENTRY.chomp
2843 [agent #{index + 1}: #{task}]
2844 #{results[index]}
2845 
2846 [verify #{index + 1}]
2847 #{verdicts[index]}
2848 ENTRY
28422849 end.join("\n\n")
28432850 if dropped > 0
28442851 joined =
from line 3025
30183025 if response.stop_reason != "tool_use":
30193026 text = "".join(block.text for block in response.content if block.type == "text")
30203027 if response.stop_reason == "max_tokens":
3021 # Drop the truncated assistant message so later turns don't build on it.
3028 # Drop the truncated assistant message so later turns do not build on it.
30223029 self.messages.pop()
30233030 text += "\n\n(warning: response was truncated at max_tokens)"
30243031 return text
from line 3056
30493056 ```typescript TypeScript
30503057 // An agent loop whose orchestration mode is toggled with mid-conversation system messages.
30513058 class ModeAgent {
3052 private readonly model: string;
3053 private modeOn: boolean;
3054 private readonly messages: Anthropic.MessageParam[] = [];
3055 private modeAnnounced = false;
3056 private exitPending = false;
3057 private turnsSinceReminder = 0;
3059 readonly #model: string;
3060 readonly #messages: Anthropic.MessageParam[] = [];
3061 #modeOn: boolean;
3062 #modeAnnounced = false;
3063 #exitPending = false;
3064 #turnsSinceReminder = 0;
30583065 
30593066 constructor(model: string, modeOn = true) {
3060 this.model = model;
3061 this.modeOn = modeOn;
3067 this.#model = model;
3068 this.#modeOn = modeOn;
30623069 }
30633070 
30643071 // Turn the mode on or off. The notice is delivered with the next user turn.
30653072 setMode(modeOn: boolean): void {
3066 if (modeOn === this.modeOn) {
3073 if (modeOn === this.#modeOn) {
30673074 return;
30683075 }
3069 if (!modeOn) {
3070 if (this.modeAnnounced) {
3071 this.exitPending = true;
3072 }
3073 } else {
3074 this.exitPending = false;
3075 }
3076 this.modeOn = modeOn;
3076 // An exit notice is only owed if the model was told the mode was on.
3077 this.#exitPending = !modeOn && this.#modeAnnounced;
3078 this.#modeOn = modeOn;
30773079 }
30783080 
30793081 // System messages owed on this turn: an exit notice, the full mode text on entry,
30803082 // or a one-line refresher every TURNS_BETWEEN_REFRESHERS user turns.
3081 private dueSystemMessages(): Anthropic.MessageParam[] {
3082 const due: Array<{ role: "system"; content: string }> = [];
3083 if (this.exitPending) {
3084 this.exitPending = false;
3085 this.modeAnnounced = false;
3083 #dueSystemMessages(): Anthropic.MessageParam[] {
3084 const due: Anthropic.MessageParam[] = [];
3085 if (this.#exitPending) {
3086 this.#exitPending = false;
3087 this.#modeAnnounced = false;
30863088 due.push({ role: "system", content: MODE_EXIT });
30873089 }
3088 if (this.modeOn) {
3089 if (!this.modeAnnounced) {
3090 this.modeAnnounced = true;
3091 this.turnsSinceReminder = 0;
3090 if (this.#modeOn) {
3091 if (!this.#modeAnnounced) {
3092 this.#modeAnnounced = true;
3093 this.#turnsSinceReminder = 0;
30923094 due.push({ role: "system", content: MODE_ENTER });
3093 } else if (this.turnsSinceReminder >= TURNS_BETWEEN_REFRESHERS) {
3094 this.turnsSinceReminder = 0;
3095 } else if (this.#turnsSinceReminder >= TURNS_BETWEEN_REFRESHERS) {
3096 this.#turnsSinceReminder = 0;
30953097 due.push({ role: "system", content: MODE_REFRESH });
30963098 }
30973099 }
3098 // The published SDK types message roles as "user" | "assistant"; typed support for
3099 // mid-conversation system messages ships with the SDK release that includes them.
3100 return due as unknown as Anthropic.MessageParam[];
3100 return due;
31013101 }
31023102 
31033103 async turn(userInput: string): Promise<string> {
31043104 // Mid-conversation system messages follow the user turn they apply to, which keeps
31053105 // the cached prefix ahead of them untouched.
3106 this.messages.push({ role: "user", content: userInput });
3107 this.messages.push(...this.dueSystemMessages());
3108 this.turnsSinceReminder += 1;
3106 this.#messages.push({ role: "user", content: userInput }, ...this.#dueSystemMessages());
3107 this.#turnsSinceReminder += 1;
31093108 
31103109 for (let turn = 0; turn < MAX_MAIN_TURNS; turn++) {
31113110 const response = await client.messages
31123111 .stream(
31133112 {
3114 model: this.model,
3113 model: this.#model,
31153114 max_tokens: 64000,
31163115 system: SYSTEM_PROMPT, // static for the whole session
31173116 output_config: { effort: EFFORT },
31183117 tools: [WORKFLOW_TOOL, BASH_TOOL],
3119 messages: this.messages,
3118 messages: this.#messages,
31203119 },
31213120 { signal: AbortSignal.timeout(REQUEST_TIMEOUT_SECONDS * 1000) },
31223121 )
31233122 .finalMessage();
3124 this.messages.push({ role: "assistant", content: response.content });
3123 this.#messages.push({ role: "assistant", content: response.content });
31253124 
31263125 if (response.stop_reason === "pause_turn") {
31273126 continue;
31283127 }
31293128 if (response.stop_reason !== "tool_use") {
31303129 let text = response.content
3131 .filter((block): block is Anthropic.TextBlock => block.type === "text")
3132 .map((block) => block.text)
3130 .flatMap((block) => (block.type === "text" ? [block.text] : []))
31333131 .join("");
31343132 if (response.stop_reason === "max_tokens") {
31353133 // Drop the truncated assistant message so later turns do not build on it.
3136 this.messages.pop();
3134 this.#messages.pop();
31373135 text += "\n\n(warning: response was truncated at max_tokens)";
31383136 }
31393137 return text;
from line 3142
31443142 if (block.type !== "tool_use") {
31453143 continue;
31463144 }
3147 let output: string;
3148 let isError: boolean;
3145 let outcome: ToolOutcome;
31493146 switch (block.name) {
31503147 case "Workflow": {
3151 const input = block.input as { subtasks?: unknown };
3152 ({ output, isError } = await runWorkflow(this.model, input.subtasks ?? []));
3148 const { subtasks } = block.input as { subtasks?: unknown };
3149 outcome = await runWorkflow(this.#model, subtasks);
31533150 break;
31543151 }
31553152 case "bash":
3156 ({ output, isError } = await handleBashBlock(block));
3153 outcome = await handleBashBlock(block);
31573154 break;
31583155 default:
3159 output = `unknown tool: ${block.name}`;
3160 isError = true;
3156 outcome = { output: `unknown tool: ${block.name}`, isError: true };
31613157 }
31623158 toolResults.push({
31633159 type: "tool_result",
31643160 tool_use_id: block.id,
3165 content: output,
3166 is_error: isError,
3161 content: outcome.output,
3162 is_error: outcome.isError,
31673163 });
31683164 }
3169 this.messages.push({ role: "user", content: toolResults });
3165 this.#messages.push({ role: "user", content: toolResults });
31703166 }
31713167 return "(hit the main loop turn limit before finishing)";
31723168 }
from line 3184
31883184 {
31893185 return;
31903186 }
3191 if (!nextModeOn)
3192 {
3193 if (modeAnnounced)
3194 {
3195 exitPending = true;
3196 }
3197 }
3198 else
3199 {
3200 exitPending = false;
3201 }
3187 // An exit notice is owed only if the model was told the mode was on.
3188 exitPending = !nextModeOn && modeAnnounced;
32023189 modeOn = nextModeOn;
32033190 }
32043191 
3205 // The Role property is an open enum, so the mid-conversation "system" role can be assigned
3206 // as a raw string; a dedicated constant ships with the SDK release.
3207 MessageParam SystemMessage(string content) => new() { Role = "system", Content = content };
3192 MessageParam SystemMessage(string content) => new() { Role = Role.System, Content = content };
32083193 
32093194 // System messages owed on this turn: an exit notice, the full mode text on entry,
32103195 // or a one-line refresher every turnsBetweenRefreshers user turns.
from line 3231
32463231 for (var turn = 0; turn < maxMainTurns; turn++)
32473232 {
32483233 using var deadline = new CancellationTokenSource(TimeSpan.FromSeconds(requestTimeoutSeconds));
3249 var response = await client.Messages.Create(new MessageCreateParams
3234 var response = await client.Messages.CreateStreaming(new MessageCreateParams
32503235 {
32513236 Model = model,
3252 MaxTokens = requestMaxTokens,
3237 MaxTokens = 64000,
32533238 System = systemPrompt, // static for the whole session
32543239 OutputConfig = new OutputConfig { Effort = effort },
32553240 Tools = [workflowTool, bashTool],
32563241 Messages = messages,
3257 }, cancellationToken: deadline.Token);
3242 }, cancellationToken: deadline.Token).Aggregate();
32583243 messages.Add(new()
32593244 {
32603245 Role = Role.Assistant,
from line 3256
32713256 response.Content.Select(block => block.TryPickText(out var textBlock) ? textBlock.Text : ""));
32723257 if (response.StopReason == StopReason.MaxTokens)
32733258 {
3274 // Drop the truncated assistant message so the next turn does not build on it.
3259 // Drop the truncated assistant message so later turns do not build on it.
32753260 messages.RemoveAt(messages.Count - 1);
32763261 text += "\n\n(warning: response was truncated at max_tokens)";
32773262 }
from line 3270
32853270 {
32863271 continue;
32873272 }
3288 string output;
3289 bool isError;
3290 if (toolUse.Name == "Workflow")
3273 var (output, isError) = toolUse.Name switch
32913274 {
3292 toolUse.Input.TryGetValue("subtasks", out var rawSubtasks);
3293 (output, isError) = await RunWorkflow(rawSubtasks);
3294 }
3295 else if (toolUse.Name == "bash")
3296 {
3297 (output, isError) = await HandleBashBlock(toolUse);
3298 }
3299 else
3300 {
3301 output = $"unknown tool: {toolUse.Name}";
3302 isError = true;
3303 }
3275 "Workflow" => await RunWorkflow(toolUse.Input.GetValueOrDefault("subtasks")),
3276 "bash" => await HandleBashBlock(toolUse),
3277 _ => ($"unknown tool: {toolUse.Name}", true),
3278 };
33043279 toolResults.Add(new ToolResultBlockParam(toolUse.ID) { Content = output, IsError = isError });
33053280 }
33063281 messages.Add(new() { Role = Role.User, Content = toolResults });
from line 3318
33433318 // dueSystemMessages returns the system messages owed on this turn: an exit notice, the
33443319 // full mode text on entry, or a one-line refresher every turnsBetweenRefreshers user turns.
33453320 func (agent *modeAgent) dueSystemMessages() []anthropic.MessageParam {
3346 // MessageParamRole is an open string type, so the mid-conversation "system" role can
3347 // be expressed directly; a dedicated constant ships with the SDK release.
33483321 systemMessage := func(content string) anthropic.MessageParam {
33493322 return anthropic.MessageParam{
3350 Role: anthropic.MessageParamRole("system"),
3323 Role: anthropic.MessageParamRoleSystem,
33513324 Content: []anthropic.ContentBlockParamUnion{anthropic.NewTextBlock(content)},
33523325 }
33533326 }
from line 3352
33793352 agent.turnsSinceReminder++
33803353 
33813354 for range maxMainTurns {
3382 var response anthropic.Message
3383 err := func() error {
3384 ctx, cancel := context.WithTimeout(ctx, requestTimeoutSeconds*time.Second)
3385 defer cancel()
3386 stream := client.Messages.NewStreaming(ctx, anthropic.MessageNewParams{
3387 Model: agent.model,
3388 MaxTokens: 64000,
3389 System: []anthropic.TextBlockParam{{Text: systemPrompt}}, // static for the whole session
3390 OutputConfig: anthropic.OutputConfigParam{Effort: effort},
3391 Tools: []anthropic.ToolUnionParam{workflowTool, bashTool},
3392 Messages: agent.messages,
3393 })
3394 defer stream.Close()
3395 for stream.Next() {
3396 if err := response.Accumulate(stream.Current()); err != nil {
3397 return err
3398 }
3399 }
3400 return stream.Err()
3401 }()
3355 response, err := streamMessage(ctx, anthropic.MessageNewParams{
3356 Model: agent.model,
3357 MaxTokens: 64000,
3358 System: []anthropic.TextBlockParam{{Text: systemPrompt}}, // static for the whole session
3359 OutputConfig: anthropic.OutputConfigParam{Effort: effort},
3360 Tools: []anthropic.ToolUnionParam{workflowTool, bashTool},
3361 Messages: agent.messages,
3362 })
34023363 if err != nil {
34033364 return "", err
34043365 }
34053366 agent.messages = append(agent.messages, response.ToParam())
34063367 
3407 if response.StopReason == anthropic.StopReasonPauseTurn {
3368 switch response.StopReason {
3369 case anthropic.StopReasonPauseTurn:
34083370 continue
3409 }
3410 if response.StopReason != anthropic.StopReasonToolUse {
3411 var text strings.Builder
3412 for _, block := range response.Content {
3413 if textBlock, ok := block.AsAny().(anthropic.TextBlock); ok {
3414 text.WriteString(textBlock.Text)
3415 }
3416 }
3417 if response.StopReason == anthropic.StopReasonMaxTokens {
3418 // Drop the truncated assistant message rather than leave a clipped turn in history.
3419 agent.messages = agent.messages[:len(agent.messages)-1]
3420 text.WriteString("\n\n(warning: response was truncated at max_tokens)")
3421 }
3422 return text.String(), nil
3371 case anthropic.StopReasonMaxTokens:
3372 // Drop the truncated assistant message so later turns do not build on it.
3373 agent.messages = agent.messages[:len(agent.messages)-1]
3374 return textOf(response) + "\n\n(warning: response was truncated at max_tokens)", nil
3375 case anthropic.StopReasonToolUse:
3376 default:
3377 return textOf(response), nil
34233378 }
34243379 
34253380 var toolResults []anthropic.ContentBlockParamUnion
from line 3398
34433398 case "bash":
34443399 output, isError = handleBashBlock(ctx, toolUse)
34453400 default:
3446 output, isError = fmt.Sprintf("unknown tool: %s", toolUse.Name), true
3401 output, isError = "unknown tool: "+toolUse.Name, true
34473402 }
34483403 toolResults = append(toolResults, anthropic.NewToolResultBlock(toolUse.ID, output, isError))
34493404 }
from line 3413
34583413 // An agent loop whose orchestration mode is toggled with mid-conversation system messages.
34593414 class ModeAgent {
34603415 private final Model model;
3416 private final List<MessageParam> messages = new ArrayList<>();
34613417 private boolean modeOn;
3462 private final List<MessageParam> messages = new ArrayList<>();
3463 private boolean modeAnnounced = false;
3464 private boolean exitPending = false;
3465 private int turnsSinceReminder = 0;
3418 private boolean modeAnnounced;
3419 private boolean exitPending;
3420 private int turnsSinceReminder;
34663421 
34673422 ModeAgent(Model model) {
34683423 this(model, true);
from line 3428
34733428 this.modeOn = modeOn;
34743429 }
34753430 
3476 // Turn the mode on or off. The notice is delivered with the next user turn.
3431 // Turn the mode on or off. The notice is delivered with the next user turn:
3432 // an exit notice is owed only if the model was told the mode was on.
34773433 void setMode(boolean modeOn) {
34783434 if (modeOn == this.modeOn) {
34793435 return;
34803436 }
3481 if (!modeOn) {
3482 if (modeAnnounced) {
3483 exitPending = true;
3484 }
3485 } else {
3486 exitPending = false;
3487 }
3437 exitPending = !modeOn && modeAnnounced;
34883438 this.modeOn = modeOn;
34893439 }
34903440 
from line 3460
35103460 return due;
35113461 }
35123462 
3513 // MessageParam.Role is an open enum, so the mid-conversation "system" role can be
3514 // expressed with Role.of; a dedicated constant ships with the SDK release.
35153463 private MessageParam systemMessage(String content) {
35163464 return MessageParam.builder()
3517 .role(MessageParam.Role.of("system"))
3465 .role(MessageParam.Role.SYSTEM)
35183466 .content(content)
35193467 .build();
35203468 }
from line 3502
35543502 .map(TextBlock::text)
35553503 .collect(Collectors.joining());
35563504 if (StopReason.MAX_TOKENS.equals(stopReason)) {
3557 // Drop the truncated assistant message so it does not poison later turns.
3505 // Drop the truncated assistant message so later turns do not build on it.
35583506 messages.removeLast();
35593507 text += "\n\n(warning: response was truncated at max_tokens)";
35603508 }
35613509 return text;
35623510 }
35633511 
3512 List<ToolUseBlock> toolUses = response.content().stream()
3513 .flatMap(block -> block.toolUse().stream())
3514 .toList();
35643515 List<ContentBlockParam> toolResults = new ArrayList<>();
3565 for (ContentBlock block : response.content()) {
3566 if (block.toolUse().isEmpty()) {
3567 continue;
3568 }
3569 ToolUseBlock toolUse = block.toolUse().get();
3516 for (ToolUseBlock toolUse : toolUses) {
35703517 ToolOutput result = switch (toolUse.name()) {
3571 case "Workflow" -> {
3572 Map<String, JsonValue> input =
3573 (Map<String, JsonValue>) toolUse._input().asObject().orElse(Map.of());
3574 JsonValue rawSubtasks = input.getOrDefault("subtasks", JsonValue.from(List.of()));
3575 yield runWorkflow(model, rawSubtasks);
3576 }
3518 case "Workflow" -> runWorkflow(model,
3519 toolInput(toolUse).getOrDefault("subtasks", JsonValue.from(List.of())));
35773520 case "bash" -> handleBashBlock(toolUse);
35783521 default -> new ToolOutput("unknown tool: " + toolUse.name(), true);
35793522 };
from line 3598
36553598 }
36563599 }
36573600 if ($stopReason === 'max_tokens') {
3658 // Drop the truncated assistant message so the next turn does not build on it.
3601 // Drop the truncated assistant message so later turns do not build on it.
36593602 array_pop($this->messages);
36603603 $text .= "\n\n(warning: response was truncated at max_tokens)";
36613604 }
from line 3707
37643707 unless response.stop_reason == :tool_use
37653708 text = response.content.select { |block| block.type == :text }.map(&:text).join
37663709 if response.stop_reason == :max_tokens
3767 @messages.pop # drop the truncated assistant message from the history
3710 # Drop the truncated assistant message so later turns do not build on it.
3711 @messages.pop
37683712 text += "\n\n(warning: response was truncated at max_tokens)"
37693713 end
37703714 return text
from line 3844
39003844 
39013845 ```java Java
39023846 void main(String[] args) throws InterruptedException {
3903 String task = args.length > 0
3904 ? args[0]
3905 : "Explore the current directory, then give a thorough review: what it does, "
3906 + "code-quality issues, and concrete improvements.";
3847 String task = args.length > 0 ? args[0] : """
3848 Explore the current directory, then give a thorough review: what it does, \
3849 code-quality issues, and concrete improvements.""";
39073850 ModeAgent agent = new ModeAgent(MODEL);
39083851 IO.println(agent.turn(task));
39093852 agent.setMode(false);
39103853 

build-with-claude/mid-conversation-system-messages Changed · +24 / -3 lines

from line 1631
16311631* **Per-turn context that must be authoritative.** You want to inject a freshness note, a session deadline, or a tool-availability change with system-level weight, and it changes too often to live in the cached prefix.
16321632* **Per-turn reminders that shouldn't pile up.** A harness nudges the model after each batch of tool results ("request independent reads together", "the user hasn't heard from you in a while") and wants the model to see only the newest copy. A [turn-scoped system message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#turn-scoped-system-messages) renders for one turn and then costs nothing, without deleting anything from the history.
16331633* **State changes your application observes.** Your application notices something Claude should treat as an operator-level fact: files changed on disk, the user toggled an auto-approve setting, available tools changed, or the remaining token budget dropped below a threshold.
1634* **User input that should not interrupt an agentic loop.** A user types a follow-up while Claude is still executing tools for the previous request. Relaying it as a system message after the next tool result lets Claude fold the new input into the work it is already doing, instead of treating it as a fresh request to switch to. See [Placement after tool results](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#placement-after-tool-results).
1634* **User input that should not interrupt an agentic loop.** A user types a follow-up while Claude is still executing tools for the previous request. Relaying it as a system message after the next tool result lets Claude fold the new input into the work it is already doing, instead of treating it as a fresh request to switch to. On Claude Sonnet 5.5 and Claude Haiku 5.5, append the user's words as a `text` block after the last `tool_result` in the same `user` message instead. See [Placement after tool results](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#placement-after-tool-results).
16351635* **Mode switches that grant standing permissions.** A session-level mode can use a mid-conversation system message to grant standing consent to an expensive capability, such as automatically launching multiagent workflows, with a short refresher every several turns and an exit notice when the mode is turned off. For a worked example, see [Build an orchestration mode](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-effort-example).
16361636 
16371637In all of these cases you could put the instruction in a regular `user` message, and Claude does follow instructions that arrive in user turns. The difference is priority: a `user` message is treated as coming from the end user, while a `system` message is treated as coming from you, the application operator. When the two conflict, system instructions take precedence, so use the `system` role for operator-level facts and constraints that should hold even if the end user asks for something different. A mid-conversation system message keeps that operator-level priority without paying the cache-miss cost of editing the top-level `system` field.
from line 1953
19531953 
19541954This example enables [automatic caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#automatic-caching) with the top-level `cache_control` field. Prompt caching is opt-in: if a request has no `cache_control` field (automatic or an [explicit breakpoint](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#explicit-cache-breakpoints)), nothing is cached and every request pays the regular input token price for the full conversation. With caching enabled, appending the system message leaves the already-cached turns unchanged, so the request that carries the new instruction still reads them from cache instead of processing them again. Caching also requires the conversation to meet the [minimum cacheable prompt length](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#cache-limitations); an example as short as this one falls below it, so `cache_creation_input_tokens` and `cache_read_input_tokens` stay at 0 until the conversation grows.
19551955 
1956A mid-conversation system message must immediately follow a `user` turn (or an `assistant` turn ending in a server tool result), and must either be the last entry in `messages` or be immediately followed by an `assistant` turn. A `user` message that carries `tool_result` blocks counts: in an agentic loop you can place the system message right after the tool results, before Claude's next turn. Any other position, including between an `assistant` `tool_use` block and the `tool_result` that answers it, returns a 400 error.
1956A mid-conversation system message that carries content (`text`, `tool_addition`, or `tool_removal` blocks) must immediately follow a `user` turn (or an `assistant` turn ending in a server tool result), and must either be the last entry in `messages` or be immediately followed by an `assistant` turn. A `user` message that carries `tool_result` blocks counts: in an agentic loop you can place the system message right after the tool results, before Claude's next turn. Any other position for a message that carries content, including between an `assistant` `tool_use` block and the `tool_result` that answers it, returns a 400 error. A message with empty `content` that only sets [`output_config.effort`](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) is accepted anywhere in `messages`. Consecutive `system` messages are judged together, so an effort-only message next to one that carries content follows the rule for content.
19571957 
19581958### Placement after tool results
19591959 
1960In an [agentic loop](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview), the system message goes after the `user` message that delivers the tool results. This is also where your application can relay input that the user typed while Claude was working, so the new context is absorbed without restarting the turn:
1960In an [agentic loop](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview), the system message goes after the `user` message that delivers the tool results. On most models, this is also where your application can relay input that the user typed while Claude was working, so the new context is absorbed without restarting the turn (for Claude Sonnet 5.5 and Claude Haiku 5.5, see the end of this section):
19611961 
19621962```json
19631963[
from line 1982
19821982Phrase the system content as context rather than as a command that overrides the user. State the fact ("new input arrived from the user: X", "the remaining token budget is now Y") and let Claude act on it. Claude is trained to resist instructions that appear to work against the user, and that protection still applies to the system role, so language such as "ignore what the user said" is less effective than stating what changed.
19831983 
19841984This pattern is for relaying input from the conversation's own end user. Do not use it to pass tool output, retrieved documents, or other third-party content; keep that content in `tool_result` blocks (see [Limitations](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#limitations)).
1985 
1986On Claude Sonnet 5.5 and Claude Haiku 5.5, deliver the user's words in the user turn instead. These models are trained to resist prompt injection through tool results, so they can treat a system message placed right after a tool result as untrusted text and ignore it. Append the user's words as a `text` block after the last `tool_result` in the same `user` message. Keep harness notices, such as reminders, in a separate mid-conversation system message after that `user` message, and never put a notice and the user's words in the same block:
1987 
1988```json
1989[
1990 { "role": "user", "content": "Run the test suite and fix any failures." },
1991 {
1992 "role": "assistant",
1993 "content": [{ "type": "tool_use", "id": "toolu_01", "name": "run_tests", "input": {} }]
1994 },
1995 {
1996 "role": "user",
1997 "content": [
1998 { "type": "tool_result", "tool_use_id": "toolu_01", "content": "12 passed, 0 failed" },
1999 { "type": "text", "text": "Also update the changelog before you finish." }
2000 ]
2001 }
2002]
2003```
2004 
2005For more on this behavior, see [Prompting Claude Sonnet 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#mid-turn-user-messages-and-task-budgets) and [Prompting Claude Haiku 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#mid-turn-user-messages).
19852006 
19862007### Turn-scoped system messages
19872008 

build-with-claude/preserved-thinking Changed · +8 / -6 lines

from line 34
3434 
3535## Switching models mid-conversation
3636 
37Claude Fable 5.1 and Claude Mythos 5.1 read thinking blocks produced by each other and by earlier Claude models. No earlier model reads thinking blocks from Claude Fable 5.1 or Claude Mythos 5.1.
37Claude Fable 5.1 and Claude Mythos 5.1 read thinking blocks produced by each other and by earlier Claude models. No other model reads thinking blocks from Claude Fable 5.1 or Claude Mythos 5.1.
3838 
3939Claude Opus 5.5 reads thinking blocks from Claude Opus 5, from earlier Opus, Sonnet, and Haiku models, and, on the Claude API and Google Cloud, from Claude Sonnet 5.5 and Claude Haiku 5.5, but not from Claude Fable or Claude Mythos models. On the Claude API, Claude Fable 5.1 and Claude Mythos 5.1 read thinking blocks from Claude Opus 5.5; no other model does. So a conversation that moves from Claude Opus 5 onto Claude Opus 5.5 keeps its reasoning, and so does one that moves from Claude Opus 5.5 up to Claude Fable 5.1 or Claude Mythos 5.1 on the Claude API. One that moves from Claude Fable 5.1 or Claude Mythos 5.1 to Claude Opus 5.5, or from Claude Opus 5.5 to any model other than those two, runs the turns after the switch without the previous model's reasoning. The blocks are dropped, not rejected, as described later in this section.
4040 
4141Claude Sonnet 5.5 reads thinking blocks from Claude Sonnet 5, Claude Opus 4.8, Claude Haiku 4.5, and earlier models, and, on the Claude API and Google Cloud, from Claude Haiku 5.5, but not from Claude Opus 5, Claude Opus 5.5, or any Claude Fable or Claude Mythos model. On the Claude API and Google Cloud, Claude Opus 5.5 reads thinking blocks from Claude Sonnet 5.5; no other model does. So a conversation that moves from Claude Sonnet 5 onto Claude Sonnet 5.5 keeps its reasoning, and so does one that moves from Claude Sonnet 5.5 up to Claude Opus 5.5 on the Claude API and Google Cloud. One that moves onto Claude Sonnet 5.5 from Claude Opus 5, Claude Opus 5.5, or a Claude Fable or Claude Mythos model runs the turns after the switch without the previous model's reasoning. So does any other move away from Claude Sonnet 5.5, for example a [server-side fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#server-side-fallback) to Claude Sonnet 5.
4242 
43Claude Haiku 5.5 reads thinking blocks from Claude Sonnet 5, Claude Opus 4.8, Claude Haiku 4.5, and earlier models, but not from Claude Opus 5, Claude Opus 5.5, Claude Sonnet 5.5, or any Claude Fable or Claude Mythos model. On the Claude API and Google Cloud, Claude Opus 5.5 and Claude Sonnet 5.5 read thinking blocks from Claude Haiku 5.5; no other model does. So a conversation that moves from Claude Haiku 4.5 onto Claude Haiku 5.5 keeps its reasoning, and so does one that moves from Claude Haiku 5.5 up to Claude Opus 5.5 or Claude Sonnet 5.5 on the Claude API and Google Cloud. One that moves onto Claude Haiku 5.5 from Claude Opus 5, Claude Opus 5.5, Claude Sonnet 5.5, or a Claude Fable or Claude Mythos model runs the turns after the switch without the previous model's reasoning, and so does any other move away from Claude Haiku 5.5.
44 
4345* **A conversation that moves to Claude Fable 5.1 from an earlier model, or from Claude Opus 5.5 on the Claude API, keeps its reasoning.** The earlier model's thinking blocks stay readable, so the model thinks as usual from the first turn after the switch.
44* **A conversation that moves down to an earlier model loses Claude Fable 5.1's reasoning for that request.** This occurs when a router sends a turn to a cheaper model, after a [classifier refusal fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback), or during a [server-side fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#server-side-fallback). The API removes the unreadable blocks before the prompt reaches the model. They aren't billed and don't count toward `input_tokens`.
46* **A conversation that moves from Claude Fable 5.1 to any model other than Claude Mythos 5.1 loses Claude Fable 5.1's reasoning for that request.** This occurs when a router sends a turn to a cheaper model, after a [classifier refusal fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback), or during a [server-side fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#server-side-fallback). The API removes the unreadable blocks before the prompt reaches the model. They aren't billed and don't count toward `input_tokens`.
4547 
4648Keep sending the full history on every request, thinking blocks included, and let the API drop what the current model can't read. The API never edits your `messages` array, so the dropped blocks stay in your history. When the same history goes back to Claude Fable 5.1, its blocks are readable again, along with the earlier model's thinking. The reasoning is lost for good only if your client removes the blocks itself, for example a harness that strips thinking on a model switch or rebuilds the history from what each model used.
4749 
from line 1261
12591261 
12601262Store the `content` array from each response and send it back unchanged as the assistant turn: every block type, in the order received, including `thinking` blocks whose `thinking` field is empty. A serializer that drops unknown block types, drops empty fields, or reorders blocks edits the prefix for every later turn.
12611263 
1262On Claude Fable 5.1 and Claude Haiku 5.5, the `thinking` field is empty by default and the `signature` carries the reasoning, so a serializer that skips empty blocks removes thinking. If it removes all of them, nothing fails and the model loses its earlier reasoning on every turn. If you parse the stream yourself, keep the block even when no thinking text arrives: it opens, receives its `signature` in a `signature_delta` event, and closes. A block sent back with an empty `signature` fails.
1264On Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, and Claude Haiku 5.5, the `thinking` field is empty by default and the `signature` carries the reasoning, so a serializer that skips empty blocks removes thinking. If it removes all of them, nothing fails and the model loses its earlier reasoning on every turn. If you parse the stream yourself, keep the block even when no thinking text arrives: it opens, receives its `signature` in a `signature_delta` event, and closes. A block sent back with an empty `signature` fails.
12631265 
12641266### Add instructions with a mid-conversation system message
12651267 
from line 1419
14171419 
14181420### Change effort with a per-message `output_config`
14191421 
1420Changing top-level `output_config.effort` between requests doesn't invalidate thinking, because effort isn't part of the prefix. Changing top-level effort does restart the prompt cache. On Claude Fable 5.1 and on Claude Haiku 5.5 with adaptive thinking, use [per-message effort](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) instead: append a `role: "system"` message with empty `content` and the new level. It needs the beta header `mid-conversation-output-config-2026-07-01`.
1422Changing top-level `output_config.effort` between requests doesn't invalidate thinking, because effort isn't part of the prefix. Changing top-level effort does restart the prompt cache. Where the model and platform support [per-message effort](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta), use it instead: append a `role: "system"` message with empty `content` and the new level. It needs the beta header `mid-conversation-output-config-2026-07-01`. On Claude Haiku 5.5, per-message effort is available on the Claude API and Google Cloud and needs adaptive thinking.
14211423 
14221424```json
14231425{ "role": "system", "content": [], "output_config": { "effort": "low" } }
14241426```
14251427 
1426The new level takes effect from the next `user` turn. Once sent, the message is part of `messages` and therefore part of the prefix for later thinking: leave it in place on later requests, and append another one to change effort again.
1428When the new level takes effect depends on where you put the message. Directly after a `user` message with new input, it starts with Claude's reply to that message. Anywhere else, such as after an `assistant` message or after a `user` message that holds only `tool_result` blocks, it starts with the next `user` message with new input, so a tool loop already under way keeps the current level until then. Once sent, the message is part of `messages` and therefore part of the prefix for later thinking: leave it in place on later requests, and append another one to change effort again.
14271429 
14281430### Trim context on the server
14291431 
from line 1719
17171719 </Accordion>
17181720 
17191721 <Accordion title="Does changing effort or other thinking settings between requests invalidate earlier thinking?">
1720 No. `output_config.effort`, `max_tokens`, and the `thinking` configuration aren't part of the checked prefix, which covers only `system`, `tools`, and `messages`. A top-level effort change invalidates most of the prompt cache. On Claude Fable 5.1 and on Claude Haiku 5.5 with adaptive thinking, a [per-message effort](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#effort-changes) change keeps the prompt cache and is used as the new effort level until changed again.
1722 No. `output_config.effort`, `max_tokens`, and the `thinking` configuration aren't part of the checked prefix, which covers only `system`, `tools`, and `messages`. A top-level effort change invalidates most of the prompt cache. Where the model and platform support [per-message effort](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#effort-changes), a per-message effort change keeps the prompt cache and is used as the new effort level until changed again.
17211723 </Accordion>
17221724 
17231725 <Accordion title="My tool list changes mid-session. How do I avoid invalidating the conversation?">

build-with-claude/prompt-caching Changed · +13 / -13 lines

from line 268
268268 * 1-hour cache write tokens are 2 times the base input tokens price
269269 * Cache read tokens are 0.1 times the base input tokens price (see the table footnote for per-model exceptions)
270270 
271 These multipliers stack with other pricing modifiers such as the Batch API discount and data residency. See [pricing](https://platform.claude.com/docs/en/about-claude/pricing) for full details.
271 These multipliers stack with other pricing modifiers such as the Batch API discount, [long context pricing](https://platform.claude.com/docs/en/about-claude/pricing#long-context-pricing), and data residency. On a model priced by prompt length, cache reads and cache writes count toward a request's prompt length. See [pricing](https://platform.claude.com/docs/en/about-claude/pricing) for full details.
272272</Note>
273273 
274274***
from line 658
658658 
659659The following table shows which parts of the cache are invalidated by different types of changes. ✘ indicates that the cache is invalidated, while ✓ indicates that the cache remains valid.
660660 
661| What changes | Tools cache | System cache | Messages cache | Impact |
662| --------------------------------------------------------- | -------------- | -------------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
663| **Tool definitions** | ✘ | ✘ | ✘ | Modifying tool definitions (names, descriptions, parameters) invalidates the entire cache |
664| **Web search toggle** | ✓ | ✘ | ✘ | Enabling/disabling web search modifies the system prompt |
665| **Citations toggle** | ✓ | ✘ | ✘ | Enabling/disabling citations modifies the system prompt |
666| **Speed setting** | ✓ | ✘ | ✘ | Switching between [`speed: "fast"` and standard speed](https://platform.claude.com/docs/en/build-with-claude/fast-mode) invalidates system and message caches |
667| **Tool choice** | ✓ | ✓ | ✘ | Changes to `tool_choice` parameter only affect message blocks |
668| **Images** | ✓ | ✓ | ✘ | Adding/removing images anywhere in the prompt affects message blocks |
669| **Thinking parameters** | Model-specific | Model-specific | ✘ | The thinking configuration (mode, and `budget_tokens` in extended mode) is rendered into the prompt, so changing it always invalidates message blocks; tool and system caches are also invalidated on models that render the configuration ahead of them. See [Thinking and prompt caching](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-prompt-caching). |
670| **Effort setting** | Model-specific | Model-specific | ✘ | Changing the [`output_config.effort`](https://platform.claude.com/docs/en/build-with-claude/effort) value always invalidates message blocks, with the same model-specific effect on tool and system caches as thinking parameters. Setting effort explicitly to the model's default is equivalent to omitting it and does not invalidate. On models that support [per-message effort](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta), an effort change carried in a `role: "system"` message inside `messages` leaves the cached prefix intact. |
671| **Non-tool results passed to extended thinking requests** | ✓ | ✓ | Model-specific | On Opus 4.5+, Sonnet 4.6+, and Haiku 5.5, thinking blocks are preserved by default, so the cache remains valid (✓). On earlier Opus/Sonnet models and Haiku models through Claude Haiku 4.5, all previously-cached thinking blocks are stripped from context, and any messages that follow those thinking blocks are removed from the cache (✘). For more details, see [Caching with thinking blocks](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#caching-with-thinking-blocks). |
672| **Dropped thinking blocks** | ✓ | ✓ | ✘ | When the API drops a Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5.5, Claude Sonnet 5.5, or Claude Haiku 5.5 thinking block that isn't [preserved](https://platform.claude.com/docs/en/build-with-claude/thinking#preserved-thinking) on that request (for example, one you replay to a model that can't read it), the cached prefix changes from that block's position onward on that request. Blocks the receiving model can read, passed back unchanged, keep the cache intact. |
661| What changes | Tools cache | System cache | Messages cache | Impact |
662| ----------------------------------------------------- | -------------- | -------------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
663| **Tool definitions** | ✘ | ✘ | ✘ | Modifying tool definitions (names, descriptions, parameters) invalidates the entire cache |
664| **Web search toggle** | ✓ | ✘ | ✘ | Enabling/disabling web search modifies the system prompt |
665| **Citations toggle** | ✓ | ✘ | ✘ | Enabling/disabling citations modifies the system prompt |
666| **Speed setting** | ✓ | ✘ | ✘ | Switching between [`speed: "fast"` and standard speed](https://platform.claude.com/docs/en/build-with-claude/fast-mode) invalidates system and message caches |
667| **Tool choice** | ✓ | ✓ | ✘ | Changes to `tool_choice` parameter only affect message blocks |
668| **Images** | ✓ | ✓ | ✘ | Adding/removing images anywhere in the prompt affects message blocks |
669| **Thinking parameters** | Model-specific | Model-specific | ✘ | The thinking configuration (mode, and `budget_tokens` in extended mode) is rendered into the prompt, so changing it always invalidates message blocks; tool and system caches are also invalidated on models that render the configuration ahead of them. See [Thinking and prompt caching](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-prompt-caching). |
670| **Effort setting** | Model-specific | Model-specific | ✘ | Changing the [`output_config.effort`](https://platform.claude.com/docs/en/build-with-claude/effort) value always invalidates message blocks, with the same model-specific effect on tool and system caches as thinking parameters. Setting effort explicitly to the model's default is equivalent to omitting it and does not invalidate. On models that support [per-message effort](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta), an effort change carried in a `role: "system"` message inside `messages` leaves the cached prefix intact. |
671| **Non-tool results passed to requests with thinking** | ✓ | ✓ | Model-specific | On Opus 4.5+, Sonnet 4.6+, and Haiku 5.5, thinking blocks are preserved by default, so the cache remains valid (✓). On earlier Opus/Sonnet models and Haiku models through Claude Haiku 4.5, all previously-cached thinking blocks are stripped from context, and any messages that follow those thinking blocks are removed from the cache (✘). For more details, see [Caching with thinking blocks](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#caching-with-thinking-blocks). |
672| **Dropped thinking blocks** | ✓ | ✓ | ✘ | When the API drops a Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5.5, Claude Sonnet 5.5, or Claude Haiku 5.5 thinking block that isn't [preserved](https://platform.claude.com/docs/en/build-with-claude/thinking#preserved-thinking) on that request (for example, one you replay to a model that can't read it), the cached prefix changes from that block's position onward on that request. Blocks the receiving model can read, passed back unchanged, keep the cache intact. |
673673 
674674On models that support [mid-conversation tool changes](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#mid-conversation-tool-changes), the `inline-tools-2026-09-15` beta header lets you add a tool, or change a tool's definition, partway through a conversation without editing `tools`. Send the definition in a `tool_addition` block in a mid-conversation system message and leave `tools` exactly as you first sent it. The cached prefix still matches, so only the appended message is processed as new input. The one exception is a `tools` array with no non-deferred tool, where the first tool defined this way costs one full cache miss on that request. See [Define tools in a message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#define-tools-in-a-message-beta).
675675 

build-with-claude/prompt-engineering/prompting-claude-haiku-5-5 Changed · +7 / -3 lines

from line 19
1919* Requests return `stop_reason: "refusal"`: [Safeguard refusals](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#safeguard-refusals)
2020 
2121<Note>
22 For the five breaking API changes when migrating from Claude Haiku 4.5, see the [migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide).
22 For the breaking API changes when migrating from Claude Haiku 4.5, see the [migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide).
23 
24 Several sections on this page suggest text to add to your system prompt. Add it to new conversations only. A request that sends thinking blocks back after `system` changed can return a 400 error, so a stored conversation resumed with the new system prompt can fail. To change instructions partway through a conversation, [put them in the newest turn](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#changing-context). On the Claude API, Amazon Bedrock, and Google Cloud, you can instead [append a system message to the conversation](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#new-instructions).
2325</Note>
2426 
2527## Use effort to control thinking
from line 29
2729[Effort](https://platform.claude.com/docs/en/build-with-claude/effort) is the main control for how much Claude Haiku 5.5 thinks. It replaces the thinking budget (`budget_tokens`) that Claude Haiku 4.5 used, so there's no old setting to carry over. Compare two or three of these levels on your own evals:
2830 
2931* `low` is the cheapest and fastest level. Use it for chat, short tool tasks, and simple, high-volume requests. In long agent prompts, the model is more likely to skip a search, stop early, or skip a check at this level.
30* `medium` is the default on the Claude API and in Claude Code. Start here for most work, including agentic coding.
32* `medium` is the default. Start here for most work, including agentic coding.
3133* `high` suits knowledge work, longer agent tasks, and strict instruction following.
3234* `xhigh` and `max` are for work where a quality gain on your evals justifies the cost. Thinking and replies get much longer at these levels, so also run your evals on Claude Sonnet 5.5 and compare performance, cost, and speed.
3335 
from line 38
3638* Thinking is on by default and counts toward `max_tokens`, which can go up to 128,000. A `max_tokens` value sized for Claude Haiku 4.5 requests that ran without thinking can cut the reply off, so leave room for thinking.
3739* To get less thinking, lower the effort level. In Anthropic's testing, telling the model in the prompt to answer directly didn't stop it from thinking. You can also turn thinking off with `thinking: {"type": "disabled"}`. This works at `low`, `medium`, and `high` only. At `xhigh` and `max`, the request returns a 400 error.
3840* At `xhigh` effort in multi-turn chats, the model sometimes writes its whole answer in its thinking and ends the turn with no visible text. If you see this behavior, check each response for an empty reply.
39* Changing the top-level `effort` value between requests invalidates the prompt cache for the conversation's messages. To run individual turns at a different level, use a [per-message effort change](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) (beta), which keeps the cache. It needs the `mid-conversation-output-config-2026-07-01` beta header and adaptive thinking, which is the default. With thinking off, a per-message effort change returns a 400 error.
41* Changing the top-level `effort` value between requests invalidates the prompt cache for the conversation's messages. On the Claude API and Google Cloud, you can change the level partway through a conversation with a [per-message effort change](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) (beta), which keeps the cache. A change holds for every later turn until another one replaces it. It needs the `mid-conversation-output-config-2026-07-01` beta header and adaptive thinking, which is the default. With thinking off, a per-message effort change returns a 400 error.
4042 
4143## Accurate search results
4244 
from line 47
4547```text wrap
4648The current date is {{current_date}}.
4749```
50 
51Render the date once, when the conversation starts, and send the same `system` and `tools` on every later request in that conversation (see [Keep earlier turns unchanged](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide#keep-earlier-turns-unchanged)). When a conversation continues on a later day, [give the new date in the newest turn](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#changing-context) instead.
4852 
4953The model also sometimes needs an extra nudge to search. This happens most at `low` effort and with long system prompts. To fix it, add this text directly after the date:
5054 

build-with-claude/refusals-and-fallback Changed · +10 / -9 lines

from line 15
1515* [SDK middleware](https://platform.claude.com/docs/en/cli-sdks-libraries/middleware): the SDK helper that wraps all of this.
1616* [Fallback and billing cookbook](https://platform.claude.com/cookbook/fable-5-fallback-billing-guide): a worked end-to-end example.
1717 
18The simplest setup, in beta on the Claude API: set `fallbacks` to `"default"`, and the API retries a declined request on the fallback model Anthropic recommends for its refusal category. For categories with no recommended fallback, the refusal stands.
18The simplest setup, in beta on the Claude API: set `fallbacks` to `"default"`, and the API retries a declined request on the fallback model Anthropic recommends for its refusal category. For categories with no recommended fallback, the refusal stands. Claude Haiku 5.5 has no server-side fallback, so for it use [client-side fallback with the SDK middleware](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#client-side-fallback) or [write the retry yourself](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#manual-retry).
1919 
2020<CodeGroup>
2121 ```bash cURL
from line 211
211211 
212212**Mid-stream refusals:** A mid-stream refusal bills the input tokens and the output already streamed at normal rates.
213213 
214**Fallback:** When you use fallback, the refusal that triggered it is billed, in addition to the fallback request, when it arrived mid-stream or is in one of the billed categories. [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) compensates for the fallback request's prompt-cache miss, so you don't pay to cache the conversation twice. For how server-side fallback reports each attempt, see [Billing and rate limits](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#billing-and-rate-limits).
214**Fallback:** When you use fallback, the refusal that triggered it is billed, in addition to the fallback request, when it arrived mid-stream or is in one of the billed categories. [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) compensates for the fallback request's prompt-cache miss, so you don't pay to cache the conversation twice. A Claude Haiku 5.5 refusal carries no fallback credit, so a fallback after one pays the full cost of writing the fallback model's prompt cache. For how server-side fallback reports each attempt, see [Billing and rate limits](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#billing-and-rate-limits).
215215 
216216The billed categories may change as Anthropic keeps measuring and refining its safeguards' false positive rates. The **Billed before any output** column in the [refusal category table](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#refusal-response) lists the billed categories.
217217 
from line 225
225225| Any platform, using an Anthropic SDK | [The SDK middleware](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#client-side-fallback) | Configure once on the client. Retries happen automatically. |
226226| Raw HTTP or custom retry logic | [A manual retry](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#manual-retry) with [fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) | Full control. Fallback credit keeps the cost down. |
227227 
228Server-side fallback and the SDK middleware apply fallback credit for you. You only need the [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) page when you build the retry yourself.
228Server-side fallback and the SDK middleware apply fallback credit for you. You only need the [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) page when you build the retry yourself. Claude Haiku 5.5 has neither server-side fallback nor fallback credit, so use the SDK middleware or a manual retry.
229229 
230230## Server-side fallback
231231 
from line 776
776776 
777777## Client-side fallback with the SDK middleware
778778 
779The SDK includes a refusal-fallback middleware. You configure it once on the client with your list of fallback models. Calls through `client.beta.messages` (csharp, go: `client.Beta.Messages`; java: `client.beta().messages()`; php: `$client->beta->messages`) then retry refused requests automatically, on any platform. The middleware also sends the `fallback-credit-2026-07-01` beta header on every request it handles, so retries are repriced without per-request setup.
779The SDK includes a refusal-fallback middleware. You configure it once on the client with your list of fallback models. Calls through `client.beta.messages` (csharp, go: `client.Beta.Messages`; java: `client.beta().messages()`; php: `$client->beta->messages`) then retry refused requests automatically, on any platform. The middleware also sends the `fallback-credit-2026-07-01` beta header on every request it handles, so retries are repriced without per-request setup. A Claude Haiku 5.5 refusal carries no fallback credit, so a retry after one pays the full cost of writing the fallback model's prompt cache.
780780 
781781### Setting it up
782782 
from line 1128
11281128 
11291129* Retries walk your fallback list in order. A fallback model that itself refuses passes the request to the next entry.
11301130* When every model in the list has declined, the middleware returns the final refusal (the last model's refusal response) rather than raising an error.
1131* Thinking blocks from Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, or Claude Fable 5 pass through unchanged. Each retry re-sends your original request body, and the only blocks the middleware removes from conversation history on later requests are the `fallback` boundary blocks it added itself. The fallback model can't read Claude Fable 5.1 blocks, which are [preserved only for that model or a newer one](https://platform.claude.com/docs/en/build-with-claude/thinking#preserved-for-model), so the API drops them. The API also drops Claude Opus 5.5 blocks for every fallback model except Claude Fable 5.1 and Claude Mythos 5.1 (see [Switching models mid-conversation](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models)). It drops Claude Sonnet 5.5 blocks too, for every fallback model except Claude Opus 5.5 on the Claude API and Google Cloud.
1131* Thinking blocks from Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, Claude Haiku 5.5, or Claude Fable 5 pass through unchanged. Each retry re-sends your original request body, and the only blocks the middleware removes from conversation history on later requests are the `fallback` boundary blocks it added itself. The fallback model can't read Claude Fable 5.1 blocks, which only [Claude Fable 5.1 and Claude Mythos 5.1 read](https://platform.claude.com/docs/en/build-with-claude/thinking#preserved-for-model), so the API drops them. The API also drops Claude Opus 5.5 blocks for every fallback model except Claude Fable 5.1 and Claude Mythos 5.1 (see [Switching models mid-conversation](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models)). It drops Claude Sonnet 5.5 blocks too, for every fallback model except Claude Opus 5.5 on the Claude API and Google Cloud. Claude Opus 5.5 and Claude Sonnet 5.5 read Claude Haiku 5.5 blocks on the Claude API and Google Cloud, and the API drops them for a fallback model that can't read them.
11321132* Responses served through the middleware include a `fallback` content block at each model boundary, the same as server-side fallback responses. The middleware manages those blocks for you on later requests.
11331133* The model that accepted is recorded in `BetaFallbackState`, so follow-up requests that share the state stay pinned to it rather than re-asking a model that refused.
11341134 
from line 1148
11481148 <Step title="Re-send on a fallback model">
11491149 Send the same request with `model` set to a fallback model, such as Claude Opus 4.8. If the refused request sent `thinking: {"type": "between_tools"}`, change `thinking` first: only Claude Sonnet 5.5 accepts that value, so omit `thinking` or set a value the fallback model accepts. [Server-side fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#server-side-fallback) under the `2026-07-01` header makes this change for you when it falls back to Claude Sonnet 5. Another model can normally serve a request that Claude Fable 5.1 or Claude Fable 5 declines. How you handle the conversation history depends on whether you redeem a [fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit):
11501150 
1151 * **Not redeeming a credit:** you can leave the earlier `thinking` and `redacted_thinking` blocks in place or strip them to save input tokens. The fallback model normally can't use them either way: it ignores Claude Fable 5 blocks, and Claude Fable 5.1 blocks are [preserved only for that model or a newer one](https://platform.claude.com/docs/en/build-with-claude/thinking#preserved-for-model), so the API drops them. The API also drops Claude Opus 5.5 blocks for every fallback model except Claude Fable 5.1 and Claude Mythos 5.1 (see [Switching models mid-conversation](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models)). It drops Claude Sonnet 5.5 blocks too, for every fallback model except Claude Opus 5.5 on the Claude API and Google Cloud.
1151 * **Not redeeming a credit:** you can leave the earlier `thinking` and `redacted_thinking` blocks in place or strip them to save input tokens. The fallback model normally can't use them either way: only Claude Fable 5.1 and Claude Mythos 5.1 [read Claude Fable 5.1 blocks](https://platform.claude.com/docs/en/build-with-claude/thinking#preserved-for-model), and only those two models, Claude Fable 5, and Claude Mythos 5 read Claude Fable 5 blocks, so for any other fallback model the API drops them. The API also drops Claude Opus 5.5 blocks for every fallback model except Claude Fable 5.1 and Claude Mythos 5.1 (see [Switching models mid-conversation](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models)). It drops Claude Sonnet 5.5 blocks too, for every fallback model except Claude Opus 5.5 on the Claude API and Google Cloud. A retry after a Claude Haiku 5.5 refusal is always in this case. Claude Opus 5.5 and Claude Sonnet 5.5 read Claude Haiku 5.5 blocks on the Claude API and Google Cloud, so leave those blocks in place when you fall back to either model there. For a fallback model that can't read them, the API drops them without billing them.
11521152 * **Redeeming a credit:** send the body unchanged, because redemption requires an exact match. The server handles the earlier model's thinking blocks on a redemption, so do not strip them (see [Fields that must match the refused request](https://platform.claude.com/docs/en/build-with-claude/fallback-credit#reference)).
11531153 </Step>
11541154 
from line 1157
11571157 </Step>
11581158</Steps>
11591159 
1160A manual retry writes the fallback model's prompt cache from scratch, which costs more than reading an existing cache. [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) refunds that cost; redeem it on every retry you build yourself. A Claude Haiku 5.5 refusal carries no fallback credit, so a retry after one writes the fallback model's cache at full price.
1160A manual retry writes the fallback model's prompt cache from scratch, which costs more than reading an existing cache. [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) refunds that cost; redeem it on every retry you build yourself. A Claude Haiku 5.5 refusal carries no fallback credit, so a retry after one pays the full cost of writing the fallback model's prompt cache.
11611161 
11621162## Refusals in Message Batches
11631163 
from line 1166
11661166Server-side fallback is not available for batches (a batch request that includes `fallbacks` produces a per-item errored result). To retry refused batch items:
11671167 
116811681. Collect the refused items from the results.
11692. Strip the Claude Fable 5.1 or Claude Fable 5 thinking blocks from any multi-turn histories.
11703. Resubmit them on a fallback model as a new batch or as direct requests.
11692. Leave the thinking blocks in multi-turn histories in place, or strip them to save input tokens. For a fallback model that can't read them, the API drops them without billing them, as in [a manual retry](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#manual-retry).
11703. If a refused item sent `thinking: {"type": "between_tools"}`, omit `thinking` or set a value the fallback model accepts. Only Claude Sonnet 5.5 accepts `between_tools`.
11714. Resubmit them on a fallback model as a new batch or as direct requests.
11711172 
11721173## Common pitfalls
11731174 

build-with-claude/thinking Changed · +7 / -7 lines

from line 469
469469 
470470Claude Sonnet 5.5 also has thinking on by default, and it rejects `thinking: {type: "disabled"}` with a 400 error. To turn off up-front thinking, send `thinking: {type: "between_tools"}` instead. It's the lowest thinking setting on Claude Sonnet 5.5, and it's accepted at [effort](https://platform.claude.com/docs/en/build-with-claude/effort) `high` or below. The model still returns its [progress updates between tool calls](https://platform.claude.com/docs/en/build-with-claude/thinking#progress-updates). Without tools, the response contains only text, as with `disabled` on Claude Sonnet 5. See [Running without up-front thinking](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#running-without-up-front-thinking) for prompting guidance.
471471 
472Claude Haiku 5.5 also has thinking on by default and accepts `thinking: {type: "disabled"}` at [effort](https://platform.claude.com/docs/en/build-with-claude/effort) `high` or below. At `xhigh` or `max` effort, that combination returns a 400 error. To get less thinking, lower the effort level first. With thinking off, the model can skip a tool call it needs when you also request JSON output. See [Use effort to control thinking](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#use-effort-to-control-thinking) and [Use adaptive thinking with JSON output and your own tools](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#json-output-with-your-own-tools).
472Claude Haiku 5.5 also has thinking on by default and accepts `thinking: {type: "disabled"}` at [effort](https://platform.claude.com/docs/en/build-with-claude/effort) `high` or below. At `xhigh` or `max` effort, that combination returns a 400 error. To get less thinking, lower the effort level first. With thinking off, the model can skip a tool call it needs when you also request JSON output. A [per-message `output_config.effort`](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) that differs from the level in effect also returns a 400 error while thinking is off. See [Use effort to control thinking](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#use-effort-to-control-thinking) and [Use adaptive thinking with JSON output and your own tools](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#json-output-with-your-own-tools).
473473 
474474Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, and Claude Mythos Preview reject `thinking: {type: "disabled"}`. Thinking can't be turned off on these models.
475475 
from line 511
511511 
512512* You're still charged for the full thinking tokens. Omitting reduces latency, not cost.
513513* If you pass thinking blocks back in multi-turn conversations, pass them unchanged. The server decrypts the `signature` to reconstruct the original thinking for prompt construction (see [Preserving thinking blocks](https://platform.claude.com/docs/en/build-with-claude/thinking#preserving-thinking-blocks)). Any text you place in the `thinking` field of a round-tripped omitted block is ignored.
514* `display` is invalid with `thinking.type: "disabled"` (there is nothing to display).
514* `display` is invalid with `thinking.type: "disabled"` (there is nothing to display) and with `thinking.type: "between_tools"` (Claude Sonnet 5.5 only), which accepts no other field. With `between_tools`, progress updates still come back with summary text, as they do under `display: "updates"` (see [Progress updates](https://platform.claude.com/docs/en/build-with-claude/thinking#progress-updates)).
515515* When using `thinking.type: "adaptive"` and the model skips thinking for a simple request, no thinking block is produced regardless of `display`.
516516* When streaming with `display: "omitted"`, no thinking text is streamed. Each thinking block streams a `thinking_delta` with an empty `thinking` string, then its `signature_delta`. With `display: "updates"`, only [progress-update blocks](https://platform.claude.com/docs/en/build-with-claude/thinking#progress-updates) stream `thinking_delta` events that carry text. See [Streaming thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#streaming-thinking) for the event sequence.
517517 
from line 913
913913 
914914Thinking works alongside [tool use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview), letting Claude reason through tool selection and process tool results. Two constraints apply:
915915 
9161. **Tool choice limitation (manual mode):** tool use with manual extended thinking (`thinking: {type: "enabled"}`) only supports `tool_choice: {"type": "auto"}` (the default) or `tool_choice: {"type": "none"}`. Using `tool_choice: {"type": "any"}` or `tool_choice: {"type": "tool", "name": "..."}` results in an error because these options force tool use, which is incompatible with manual extended thinking. Adaptive thinking, including on models where thinking is on by default, supports forced tool use, except on Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1, and Claude Mythos 5.1 (see [Response prefill and forced tool use](https://platform.claude.com/docs/en/build-with-claude/thinking#limits-and-feature-compatibility)).
9161. **Tool choice limitation (manual mode):** tool use with manual extended thinking (`thinking: {type: "enabled"}`) only supports `tool_choice: {"type": "auto"}` (the default) or `tool_choice: {"type": "none"}`. Using `tool_choice: {"type": "any"}` or `tool_choice: {"type": "tool", "name": "..."}` results in an error because these options force tool use, which is incompatible with manual extended thinking. Adaptive thinking, including on models where thinking is on by default, supports forced tool use, except on Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1, and Claude Mythos 5.1 (see [Response prefill and forced tool use](https://platform.claude.com/docs/en/build-with-claude/thinking#limits-and-feature-compatibility)). Where forced tool use is accepted, the response starts with the tool call and has no `thinking` block.
9179172. **Preserving thinking blocks:** when you return tool results, you must pass the thinking blocks from the assistant message back to the API, complete and unmodified. See [Preserving thinking blocks](https://platform.claude.com/docs/en/build-with-claude/thinking#preserving-thinking-blocks).
918918 
919919**A tool-use loop is one assistant turn.** From the model's perspective, an assistant turn doesn't complete until Claude finishes its full response, which may include multiple tool calls and results. This whole sequence is a single assistant turn:
from line 1075
10751075 
10761076The tradeoff is context usage: long conversations consume more context space on keep-all models, because retained thinking blocks count as input like any other conversation history (see [Thinking and the context window](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-the-context-window)). The behavior is automatic in both regimes. No code changes or beta headers are required, and you should keep passing complete, unmodified thinking blocks back as described in [Preserving thinking blocks](https://platform.claude.com/docs/en/build-with-claude/thinking#preserving-thinking-blocks). To override the default in either direction, use [thinking block clearing](https://platform.claude.com/docs/en/build-with-claude/context-editing#thinking-block-clearing).
10771077 
1078**Switching models mid-conversation.** Keep passing thinking blocks back unchanged when you switch models, for example after a [classifier refusal fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback). A thinking block is readable only by the model that produced it and certain other models, and the API ignores or drops the blocks the target model can't read. On Claude Fable 5.1 and Claude Mythos 5.1 the direction matters: they read every earlier model's thinking blocks and no earlier model reads theirs, so switching up to them keeps the conversation's reasoning and switching down drops it (see [how dropped blocks are billed and reported](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models)). Claude Opus 5.5 reads Claude Opus 5's thinking blocks and those of earlier Opus, Sonnet, and Haiku models, and, on the Claude API and Google Cloud, of Claude Haiku 5.5, but not those of the Claude Fable and Claude Mythos models; on the Claude API, Claude Fable 5.1 and Claude Mythos 5.1 read Claude Opus 5.5's blocks. A switch from Claude Opus 5.5 up to Claude Fable 5.1 on the Claude API keeps the earlier turns' reasoning; a switch from Claude Fable 5.1 to Claude Opus 5.5 drops it. Claude Sonnet 5.5 reads thinking blocks from Claude Sonnet 5, Claude Opus 4.8, Claude Haiku 4.5, and earlier models, and, on the Claude API and Google Cloud, from Claude Haiku 5.5, but not from Claude Opus 5, Claude Opus 5.5, or any Claude Fable or Claude Mythos model. On the Claude API and Google Cloud, Claude Opus 5.5 reads Claude Sonnet 5.5's blocks and no other model does: a switch from Claude Sonnet 5.5 up to Claude Opus 5.5 on the Claude API and Google Cloud keeps the earlier turns' reasoning, and any other switch away from Claude Sonnet 5.5 drops it. Strip prior `thinking` and `redacted_thinking` blocks yourself only to save input tokens on models that ignore rather than drop them, and never when redeeming a [fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit), which requires the body unchanged.
1078**Switching models mid-conversation.** Keep passing thinking blocks back unchanged when you switch models, for example after a [classifier refusal fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback). A thinking block is readable only by the model that produced it and certain other models, and the API ignores or drops the blocks the target model can't read. On Claude Fable 5.1 and Claude Mythos 5.1 the direction matters: they read earlier models' thinking blocks, and only the two of them read each other's, so switching to them keeps the conversation's reasoning and switching from them to any other model drops it (see [how dropped blocks are billed and reported](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models)). Claude Opus 5.5 reads Claude Opus 5's thinking blocks and those of earlier Opus, Sonnet, and Haiku models, and, on the Claude API and Google Cloud, of Claude Haiku 5.5, but not those of the Claude Fable and Claude Mythos models; on the Claude API, Claude Fable 5.1 and Claude Mythos 5.1 read Claude Opus 5.5's blocks. A switch from Claude Opus 5.5 up to Claude Fable 5.1 on the Claude API keeps the earlier turns' reasoning; a switch from Claude Fable 5.1 to Claude Opus 5.5 drops it. Claude Sonnet 5.5 reads thinking blocks from Claude Sonnet 5, Claude Opus 4.8, Claude Haiku 4.5, and earlier models, and, on the Claude API and Google Cloud, from Claude Haiku 5.5, but not from Claude Opus 5, Claude Opus 5.5, or any Claude Fable or Claude Mythos model. On the Claude API and Google Cloud, Claude Opus 5.5 reads Claude Sonnet 5.5's blocks and no other model does: a switch from Claude Sonnet 5.5 up to Claude Opus 5.5 on the Claude API and Google Cloud keeps the earlier turns' reasoning, and any other switch away from Claude Sonnet 5.5 drops it. Claude Haiku 5.5 reads thinking blocks from Claude Sonnet 5, Claude Opus 4.8, Claude Haiku 4.5, and earlier models, but not from Claude Opus 5, Claude Opus 5.5, Claude Sonnet 5.5, or any Claude Fable or Claude Mythos model. On the Claude API and Google Cloud, Claude Opus 5.5 and Claude Sonnet 5.5 read Claude Haiku 5.5's blocks and no other model does. Strip prior `thinking` and `redacted_thinking` blocks yourself only to save input tokens on models that ignore rather than drop them, and never when redeeming a [fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit), which requires the body unchanged.
10791079 
10801080## Preserved thinking
10811081 
from line 1170
11701170* When [streaming responses](https://platform.claude.com/docs/en/build-with-claude/thinking#streaming-thinking), the signature arrives as a `signature_delta` inside a `content_block_delta` event just before the `content_block_stop` event.
11711171* `signature` values are significantly longer in Claude 4 and later models than in previous models.
11721172* The `signature` field is opaque: don't interpret or parse it.
1173* `signature` values are compatible across platforms (the Claude API, [Amazon Bedrock](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock), and [Google Cloud](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)). Values generated on one platform work on another.
1173* `signature` values are compatible across platforms (the Claude API, [Amazon Bedrock](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock), and [Google Cloud](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)). Values generated on one platform work on another, except that thinking blocks from Claude Sonnet 5.5 and Claude Haiku 5.5 work only in the account that produced them or in an account linked to it. See [Thinking blocks stay with the account that produced them](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#account-bound-thinking).
11741174 
11751175## Redacted thinking blocks
11761176 
from line 1197
11971197 
11981198### Sampling parameters
11991199 
1200On Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Sonnet 5.5, Claude Sonnet 5, and Claude Haiku 5.5, non-default `temperature`, `top_p`, or `top_k` values return a 400 error on every request, regardless of whether thinking is used. On older models, the restriction applies only while thinking is on: `temperature` and `top_k` are incompatible with thinking, and `top_p` is allowed at values between 0.95 and 1.
1200On Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Sonnet 5.5, Claude Sonnet 5, and Claude Haiku 5.5, non-default `temperature`, `top_p`, or `top_k` values return a 400 error on every request, regardless of whether thinking is used. On these models, the `top_p` default is `0.99`, so a `top_p` of `1` returns a 400 error, and so does a request that includes both `temperature` and `top_p`, even at their defaults. See [Remove sampling parameters](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide#remove-sampling-parameters). On older models, the restriction applies only while thinking is on: `temperature` and `top_k` are incompatible with thinking, and `top_p` is allowed at values between 0.95 and 1.
12011201 
12021202### Response prefill and forced tool use
12031203 
1204You can't prefill the assistant response while thinking is on. Forced tool use (`tool_choice: {"type": "any"}` or `{"type": "tool", ...}`) is incompatible with manual extended thinking but works with adaptive thinking. The exceptions are Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1, and Claude Mythos 5.1, which reject forced tool use on every request with a 400 error. On those models, use `tool_choice: {"type": "auto"}` with [strict tool use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/strict-tool-use) or [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) instead. See [Thinking with tool use](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-with-tool-use).
1204Claude 4.6 and later models and Claude Mythos Preview reject a prefilled final assistant turn with a 400 error, whether or not thinking is on (see [Prefill not supported](https://platform.claude.com/docs/en/api/errors#prefill-not-supported)). On earlier models, you can't prefill the assistant response while thinking is on. Forced tool use (`tool_choice: {"type": "any"}` or `{"type": "tool", ...}`) is incompatible with manual extended thinking but works with adaptive thinking. The exceptions are Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1, and Claude Mythos 5.1, which reject forced tool use on every request with a 400 error. On those models, use `tool_choice: {"type": "auto"}` with [strict tool use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/strict-tool-use) or [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) instead. Where forced tool use is accepted, the response starts with the tool call and has no `thinking` block. To let the model think before it calls a tool, use `tool_choice: {"type": "auto"}` and say in the prompt when to use the tool. See [Thinking with tool use](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-with-tool-use).
12051205 
12061206### Output limits
12071207 

build-with-claude/thinking-steering-and-cost Changed · +10 / -4 lines

from line 37
37371. Set the effort level that matches your workload's default balance of quality and latency.
38382. Add prompt guidance only if Claude's triggering still doesn't match your needs at that level.
3939 
40<Note>
41 To get less thinking on Claude Sonnet 5.5 and Claude Haiku 5.5, lower the effort level. Asking Claude Sonnet 5.5 in the system prompt to think less doesn't reliably reduce its thinking, and in Anthropic's testing, telling Claude Haiku 5.5 in the prompt to answer directly didn't stop it from thinking. See [Prompting Claude Sonnet 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#calibrate-effort) and [Prompting Claude Haiku 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#use-effort-to-control-thinking).
42</Note>
43 
4044For broader prompting guidance with thinking, see [leverage thinking and interleaved thinking capabilities](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices#leverage-thinking-and-interleaved-thinking-capabilities).
4145 
4246### Effort levels
from line 92
8892 
8993### Per-message steering
9094 
91You can also steer thinking on a per-message basis from the user turn, independently of the system prompt. Appending `"Please think hard before responding."` to a user message encourages Claude to think on that turn; `"Answer directly without deliberating."` suppresses it.
95You can also steer thinking on a per-message basis from the user turn, independently of the system prompt. Appending `"Please think hard before responding."` to a user message encourages Claude to think on that turn; `"Answer directly without deliberating."` discourages it.
9296 
93Per-message steering is useful when only some requests in a conversation warrant extended reasoning. An agent harness, for example, can append the encouraging phrase on planning steps and the suppressing phrase on routine confirmations, without touching the system prompt or changing any request parameters between turns.
97Per-message steering is useful when only some requests in a conversation warrant extended reasoning. An agent harness, for example, can append the encouraging phrase on planning steps and the discouraging phrase on routine confirmations, without touching the system prompt or changing any request parameters between turns.
9498 
99On Claude Haiku 5.5, the discouraging phrase doesn't reliably reduce thinking. On the Claude API and Google Cloud, run routine turns on that model with less thinking by using a [per-message effort change](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) (beta) instead, which also keeps the prompt cache.
100 
95101### Verify steering on your workload
96102 
97103Prompt-based steering changes model behavior, so treat it like any other prompt change: measure before you ship. Run a representative sample of your traffic with and without the guidance, and compare how often thinking triggers (the presence of thinking blocks in responses), output token usage, latency, and answer quality on the cases that matter to you.
from line 126
120126 
121127Consecutive requests that keep the same thinking configuration and effort level preserve prompt caching; see [Thinking and prompt caching](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-prompt-caching) for the full rules. The resolved effort value is rendered into the prompt, so changing it between requests invalidates cache breakpoints, just as changing the legacy [`budget_tokens`](https://platform.claude.com/docs/en/build-with-claude/extended-thinking#extended-thinking-with-prompt-caching) parameter does on models that use it. Setting `effort` explicitly to the model's default is equivalent to omitting it and does not break the cache.
122128 
123The practical consequence: pick a thinking configuration and an effort level per conversation and keep them. If some turns need more or less thinking, steer with [per-message prompting](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost#tuning-thinking-behavior): guidance appended to the newest user message leaves earlier cache breakpoints intact, where a configuration or effort change does not.
129The practical consequence: pick a thinking configuration and a top-level effort level per conversation and keep them. If some turns need more or less thinking, use a [per-message effort change](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) (beta) on models that support it, or steer with [per-message prompting](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost#tuning-thinking-behavior). Both leave earlier cache breakpoints intact, where a change to the thinking configuration or the top-level effort level does not.
124130 
125131The following example demonstrates the invalidation with a multi-turn script you can run yourself:
126132 
Feedback