One read of Claude Developer Platformapi-20261007T190728Z
91 pages moved out of 754 read.
Pages moved
91
significant first
Pages read
754
in this capture
Captured
19:07 UTC
Corpus hash
4b90aa04d919
corpus-hash
What this read moved
26-50 of 91, page 2 of 4This capture is too large to show at once. Changes 26-50 of 91 are below, significant first; the rest are on the following screens.
build-with-claude/refusals-and-fallback Changed · +4 / -4 lines
from line 1
11---
22title: Refusals and fallback
33url: https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback
4description: How Claude Fable models, Claude Opus models, and Claude Sonnet 5.5 return classifier refusals and how to retry refused requests on a fallback model.
4description: How Claude Fable models, Claude Opus models, Claude Sonnet 5.5, and Claude Haiku 5.5 return classifier refusals and how to retry refused requests on a fallback model.
55---
66
7Claude Fable 5.1, Claude Fable 5, Claude Opus 5.5, Claude Opus 5, and Claude Sonnet 5.5 include safety classifiers that can decline a request. When that happens, you receive a normal response, not an error, with `stop_reason: "refusal"`. Its `stop_details.category` names the policy area (see [What a refusal looks like](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#refusal-response)). You can usually still get an answer by sending the same request to another Claude model. This page shows you how to recognize a refusal and how to set up that retry.
7Claude Fable 5.1, Claude Fable 5, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, and Claude Haiku 5.5 include safety classifiers that can decline a request. When that happens, you receive a normal response, not an error, with `stop_reason: "refusal"`. Its `stop_details.category` names the policy area (see [What a refusal looks like](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#refusal-response)). You can usually still get an answer by sending the same request to another Claude model. This page shows you how to recognize a refusal and how to set up that retry.
88
99Read this page when you build on any of these models and want declined requests to fall through to another model automatically. It also applies when you have seen `"refusal"` in a response and want to know what to do next.
1010
from line 232
232232Server-side fallback retries a refused request inside a single API call. In the default mode, when the primary model declines and the refusal category has a recommended fallback, the API runs the same request on the model Anthropic recommends for that category. You can instead [name up to three fallback models of your own](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#naming-your-own-fallback-models). Either way, you get back one response that names the model that answered, so your user gets an answer in one round trip.
233233
234234<Note>
235 Server-side fallback is in beta on the Claude API. The `fallbacks` parameter is not supported on the [Message Batches API](https://platform.claude.com/docs/en/build-with-claude/batch-processing) (a batch item that includes it comes back as an errored result) and is not available on Amazon Bedrock, Google Cloud, or Microsoft Foundry. On those platforms, use [client-side fallback with the SDK middleware](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#client-side-fallback) instead.
235 Server-side fallback is in beta on the Claude API. The `fallbacks` parameter is not supported on the [Message Batches API](https://platform.claude.com/docs/en/build-with-claude/batch-processing) (a batch item that includes it comes back as an errored result) and is not available on Amazon Bedrock, Google Cloud, or Microsoft Foundry. On those platforms, use [client-side fallback with the SDK middleware](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#client-side-fallback) instead. Claude Haiku 5.5 has no server-side fallback: with `fallbacks: "default"`, a declined request stays declined, and a list of fallback models returns a 400 error.
236236</Note>
237237
238238### Making the request
from line 1157
11571157 </Step>
11581158</Steps>
11591159
1160A manual retry writes the fallback model's prompt cache from scratch, which costs more than reading an existing cache. [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) refunds that cost; redeem it on every retry you build yourself.
1160A manual retry writes the fallback model's prompt cache from scratch, which costs more than reading an existing cache. [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit) refunds that cost; redeem it on every retry you build yourself. A Claude Haiku 5.5 refusal carries no fallback credit, so a retry after one writes the fallback model's cache at full price.
11611161
11621162## Refusals in Message Batches
11631163
build-with-claude/thinking Changed · +16 / -8 lines
from line 52
5252| --------------------- | ------------------- | ----------------- | -------------------------------- | ----------------------------------------------- | -------------------------------------- |
5353| Claude Opus 5.5 | Adaptive thinking | Adaptive thinking | 400 error | 400 error | 400 error |
5454| Claude Sonnet 5.5 | Adaptive thinking | Adaptive thinking | 400 error | Up-front thinking off at `high` effort or below | 400 error |
55| Claude Haiku 5.5 | Adaptive thinking | Adaptive thinking | 400 error | 400 error | Thinking off at `high` effort or below |
5556| Claude Fable 5.1 | Adaptive thinking | Adaptive thinking | 400 error | 400 error | 400 error |
5657| Claude Mythos 5.1 | Adaptive thinking | Adaptive thinking | 400 error | 400 error | 400 error |
5758| Claude Fable 5 | Adaptive thinking | Adaptive thinking | 400 error | 400 error | 400 error |
from line 70
6970
7071In the table, "at `high` effort or below" means the request works at `low`, `medium`, and `high` effort and returns a 400 error at `xhigh` or `max`.
7172
72On Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, and Claude Mythos Preview, thinking is already on and needs no configuration. `display` defaults to `"omitted"` on these models, so the thinking text is hidden until you opt in. Opt in with `thinking: {"type": "adaptive", "display": "summarized"}`, which is exactly the following request with the [model string](https://platform.claude.com/docs/en/models/overview) swapped.
73On Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, and Claude Haiku 5.5, thinking is already on and needs no configuration. `display` defaults to `"omitted"` on these models, so the thinking text is hidden until you opt in. Opt in with `thinking: {"type": "adaptive", "display": "summarized"}`, which is exactly the following request with the [model string](https://platform.claude.com/docs/en/models/overview) swapped.
7374
7475On Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6, thinking is off until you set `thinking: {type: "adaptive"}`, which lets Claude determine when and how deeply to think based on the request. The following examples do that, set `display: "summarized"` so the thinking text is visible, and use a roomy `max_tokens`:
7576
from line 469
468469
469470Claude Sonnet 5.5 also has thinking on by default, and it rejects `thinking: {type: "disabled"}` with a 400 error. To turn off up-front thinking, send `thinking: {type: "between_tools"}` instead. It's the lowest thinking setting on Claude Sonnet 5.5, and it's accepted at [effort](https://platform.claude.com/docs/en/build-with-claude/effort) `high` or below. The model still returns its [progress updates between tool calls](https://platform.claude.com/docs/en/build-with-claude/thinking#progress-updates). Without tools, the response contains only text, as with `disabled` on Claude Sonnet 5. See [Running without up-front thinking](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5-5#running-without-up-front-thinking) for prompting guidance.
470471
472Claude Haiku 5.5 also has thinking on by default and accepts `thinking: {type: "disabled"}` at [effort](https://platform.claude.com/docs/en/build-with-claude/effort) `high` or below. At `xhigh` or `max` effort, that combination returns a 400 error. To get less thinking, lower the effort level first. With thinking off, the model can skip a tool call it needs when you also request JSON output. See [Use effort to control thinking](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#use-effort-to-control-thinking) and [Use adaptive thinking with JSON output and your own tools](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#json-output-with-your-own-tools).
473
471474Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, and Claude Mythos Preview reject `thinking: {type: "disabled"}`. Thinking can't be turned off on these models.
472475
473476To check whether a model accepts `"disabled"` before you send a request, read its `capabilities.thinking.types.disabled.supported` value from the Models API. [Using the Models API](https://platform.claude.com/docs/en/models/overview#using-the-models-api) describes the field.
from line 484
481484The `display` field on the thinking configuration controls how thinking content is returned in API responses. `display` works in both modes: set it alongside `type: "adaptive"` or `type: "enabled"`. It accepts these values:
482485
483486* `"summarized"`: thinking blocks contain [summarized thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#summarized-thinking) text, a readable summary of Claude's reasoning. This is the default on Claude Opus 4.6, Claude Sonnet 4.6, and earlier models.
484* `"omitted"`: thinking blocks are returned with an empty `thinking` field. The `signature` field still carries the encrypted full thinking for multi-turn continuity (see [Thinking encryption](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-encryption)). This is the default on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, and [Claude Mythos Preview](https://anthropic.com/glasswing).
487* `"omitted"`: thinking blocks are returned with an empty `thinking` field. The `signature` field still carries the encrypted full thinking for multi-turn continuity (see [Thinking encryption](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-encryption)). This is the default on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, [Claude Mythos Preview](https://anthropic.com/glasswing), and Claude Haiku 5.5.
485488* `"updates"` (beta): reasoning blocks are returned with an empty `thinking` field, as with `"omitted"`, and the short [progress updates](https://platform.claude.com/docs/en/build-with-claude/thinking#progress-updates) some models write between tool calls come back as readable text. Requires the beta header `thinking-display-updates-2026-08-18`.
486489
487490Set `display: "omitted"` when your application doesn't surface thinking content to users. The primary benefit is faster time-to-first-text-token when streaming: the server skips streaming thinking tokens entirely and delivers only the signature, so the final text response begins streaming sooner.
from line 1065
10621065
10631066Whether thinking blocks from previous assistant turns stay in context by default depends on the model:
10641067
1065* **Keep all prior turns:** Claude Opus 4.5 and later Opus models, Claude Sonnet 4.6 and later Sonnet models, Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, and Claude Mythos Preview.
1068* **Keep all prior turns:** Claude Opus 4.5 and later Opus models, Claude Sonnet 4.6 and later Sonnet models, Claude Haiku 5.5, Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, and Claude Mythos Preview.
10661069* **Keep the last turn only:** earlier Opus and Sonnet models, and all Haiku models through Claude Haiku 4.5. When you pass older thinking blocks back, the API strips them automatically. You don't need to remove them yourself.
10671070
10681071Preservation brings two benefits:
from line 1075
10721075
10731076The tradeoff is context usage: long conversations consume more context space on keep-all models, because retained thinking blocks count as input like any other conversation history (see [Thinking and the context window](https://platform.claude.com/docs/en/build-with-claude/thinking#thinking-and-the-context-window)). The behavior is automatic in both regimes. No code changes or beta headers are required, and you should keep passing complete, unmodified thinking blocks back as described in [Preserving thinking blocks](https://platform.claude.com/docs/en/build-with-claude/thinking#preserving-thinking-blocks). To override the default in either direction, use [thinking block clearing](https://platform.claude.com/docs/en/build-with-claude/context-editing#thinking-block-clearing).
10741077
1075**Switching models mid-conversation.** Keep passing thinking blocks back unchanged when you switch models, for example after a [classifier refusal fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback). A thinking block is readable only by the model that produced it and certain other models, and the API ignores or drops the blocks the target model can't read. On Claude Fable 5.1 and Claude Mythos 5.1 the direction matters: they read every earlier model's thinking blocks and no earlier model reads theirs, so switching up to them keeps the conversation's reasoning and switching down drops it (see [how dropped blocks are billed and reported](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models)). Claude Opus 5.5 reads Claude Opus 5's thinking blocks and those of earlier Opus, Sonnet, and Haiku models, but not those of the Claude Fable and Claude Mythos models; on the Claude API, Claude Fable 5.1 and Claude Mythos 5.1 read Claude Opus 5.5's blocks. A switch from Claude Opus 5.5 up to Claude Fable 5.1 on the Claude API keeps the earlier turns' reasoning; a switch from Claude Fable 5.1 to Claude Opus 5.5 drops it. Claude Sonnet 5.5 reads thinking blocks from Claude Sonnet 5, Claude Opus 4.8, Claude Haiku 4.5, and earlier models, but not from Claude Opus 5, Claude Opus 5.5, or any Claude Fable or Claude Mythos model. On the Claude API and Google Cloud, Claude Opus 5.5 reads Claude Sonnet 5.5's blocks and no other model does: a switch from Claude Sonnet 5.5 up to Claude Opus 5.5 on the Claude API and Google Cloud keeps the earlier turns' reasoning, and any other switch away from Claude Sonnet 5.5 drops it. Strip prior `thinking` and `redacted_thinking` blocks yourself only to save input tokens on models that ignore rather than drop them, and never when redeeming a [fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit), which requires the body unchanged.
1078**Switching models mid-conversation.** Keep passing thinking blocks back unchanged when you switch models, for example after a [classifier refusal fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback). A thinking block is readable only by the model that produced it and certain other models, and the API ignores or drops the blocks the target model can't read. On Claude Fable 5.1 and Claude Mythos 5.1 the direction matters: they read every earlier model's thinking blocks and no earlier model reads theirs, so switching up to them keeps the conversation's reasoning and switching down drops it (see [how dropped blocks are billed and reported](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models)). Claude Opus 5.5 reads Claude Opus 5's thinking blocks and those of earlier Opus, Sonnet, and Haiku models, and, on the Claude API and Google Cloud, of Claude Haiku 5.5, but not those of the Claude Fable and Claude Mythos models; on the Claude API, Claude Fable 5.1 and Claude Mythos 5.1 read Claude Opus 5.5's blocks. A switch from Claude Opus 5.5 up to Claude Fable 5.1 on the Claude API keeps the earlier turns' reasoning; a switch from Claude Fable 5.1 to Claude Opus 5.5 drops it. Claude Sonnet 5.5 reads thinking blocks from Claude Sonnet 5, Claude Opus 4.8, Claude Haiku 4.5, and earlier models, and, on the Claude API and Google Cloud, from Claude Haiku 5.5, but not from Claude Opus 5, Claude Opus 5.5, or any Claude Fable or Claude Mythos model. On the Claude API and Google Cloud, Claude Opus 5.5 reads Claude Sonnet 5.5's blocks and no other model does: a switch from Claude Sonnet 5.5 up to Claude Opus 5.5 on the Claude API and Google Cloud keeps the earlier turns' reasoning, and any other switch away from Claude Sonnet 5.5 drops it. Strip prior `thinking` and `redacted_thinking` blocks yourself only to save input tokens on models that ignore rather than drop them, and never when redeeming a [fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit), which requires the body unchanged.
10761079
10771080## Preserved thinking
10781081
from line 1084
10811084* **The model that produced it.** Each model reads its own thinking blocks and those of a fixed set of other models. Claude Fable 5.1 reads blocks from Claude Opus 5 and, on the Claude API, from Claude Opus 5.5; neither Claude Opus 5 nor Claude Opus 5.5 reads blocks from Claude Fable 5.1. The API drops a block the current model can't read, without an error and without billing it. See [Switching models mid-conversation](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#switching-models).
10821085* **Everything sent before it.** A block stays valid only while the top-level `system` prompt, the `tools`, and the messages before it are unchanged. If any of them changes, that block and every later thinking block are invalid, and the API rejects the request with a 400 error or drops the invalid blocks, whichever you choose. See [Keeping the prefix unchanged](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#prefix-check).
10831086
1084Claude Sonnet 5.5's thinking blocks are also tied to the account that produced them. See [Thinking blocks stay with the account that produced them](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#account-bound-thinking) for where the API enforces this.
1087Thinking blocks from Claude Sonnet 5.5 and Claude Haiku 5.5 are also tied to the account that produced them. See [Thinking blocks stay with the account that produced them](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#account-bound-thinking) for where the API enforces this.
10851088
10861089The model check applies to every account. The API enforces the prefix check by default for accounts created on or after August 31, 2026, 00:00 UTC. On older accounts it enforces the check only on requests that set `thinking.block_binding.prefix_mismatch_behavior`. Make your integration append-only regardless of your account's age, so the same code works on every account, including newer accounts enforced by default.
10871090
from line 1147
11441147
11451148The following diagrams illustrate the last-turn-only (stripping) regime. The first shows a multi-turn conversation: each turn's thinking block is generated in the output but not carried into later turns' input.
11461149
1147
1150<Frame>
1151 
1152</Frame>
11481153
11491154The second shows the same regime with tool use: thinking stays in context alongside its tool result for the duration of the assistant turn, then drops out on the next user turn.
11501155
1151
1156<Frame>
1157 
1158</Frame>
11521159
11531160Use the [token counting API](https://platform.claude.com/docs/en/build-with-claude/token-counting) to get accurate counts for your specific use case, especially for multi-turn conversations that include thinking.
11541161
from line 1197
11901197
11911198### Sampling parameters
11921199
1193On Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Sonnet 5.5, and Claude Sonnet 5, non-default `temperature`, `top_p`, or `top_k` values return a 400 error on every request, regardless of whether thinking is used. On older models, the restriction applies only while thinking is on: `temperature` and `top_k` are incompatible with thinking, and `top_p` is allowed at values between 0.95 and 1.
1200On Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Sonnet 5.5, Claude Sonnet 5, and Claude Haiku 5.5, non-default `temperature`, `top_p`, or `top_k` values return a 400 error on every request, regardless of whether thinking is used. On older models, the restriction applies only while thinking is on: `temperature` and `top_k` are incompatible with thinking, and `top_p` is allowed at values between 0.95 and 1.
11941201
11951202### Response prefill and forced tool use
11961203
from line 1224
12171224| Claude Sonnet 5 | 128K | 300K |
12181225| Claude Sonnet 4.6 | 128K | 300K |
12191226| Claude Sonnet 4.5 | 64K | Not available |
1227| Claude Haiku 5.5 | 128K | 300K |
12201228| Claude Haiku 4.5 | 64K | Not available |
12211229
12221230See the [models overview](https://platform.claude.com/docs/en/models/overview) for limits on legacy models.
build-with-claude/thinking-steering-and-cost Changed · +7 / -7 lines
from line 43
4343
4444Effort is the primary steering lever for thinking. Each level sets a different default for how often Claude thinks and how deeply:
4545
46| Effort level | Thinking behavior |
47| ------------------------------------- | ------------------------------------------------------------------------------------------------ |
48| `max` | Claude thinks the most readily and at the greatest depth, with no constraint on thinking length. |
49| `xhigh` | Claude thinks more readily and at greater depth than at `high`, suited to extended exploration. |
50| `high` (default on most models) | Claude thinks on most requests that benefit from it. Provides deep reasoning on complex tasks. |
51| `medium` (default on Claude Opus 5.5) | Claude uses moderate thinking. May skip thinking for simple queries. |
52| `low` | Claude minimizes thinking. Skips thinking for simple tasks where speed matters most. |
46| Effort level | Thinking behavior |
47| ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
48| `max` | Claude thinks the most readily and at the greatest depth, with no constraint on thinking length. |
49| `xhigh` | Claude thinks more readily and at greater depth than at `high`, suited to extended exploration. |
50| `high` (default on most models) | Claude thinks on most requests that benefit from it. Provides deep reasoning on complex tasks. |
51| `medium` (default on Claude Opus 5.5 and Claude Haiku 5.5) | Claude uses moderate thinking. May skip thinking for simple queries. |
52| `low` | Claude minimizes thinking. Skips thinking for simple tasks where speed matters most. |
5353
5454At every level, Claude determines per request whether to think. In a tool-use loop, the first request after new user input typically carries most of the reasoning, and follow-up requests that only process tool results can skip thinking, including at `xhigh` and `max`. Thinking per request also tends to decrease as a conversation grows longer. No level guarantees a thinking block on every request.
5555
build-with-claude/thinking-troubleshooting Changed · +10 / -5 lines
from line 31
3131| Claude Opus 4.7 | Adaptive only | Off | `"enabled"` |
3232| Claude Sonnet 5.5 | Adaptive, `between_tools`3 | On | `"enabled"`, `"disabled"` |
3333| Claude Sonnet 5 | Adaptive only | On | `"enabled"` |
34| Claude Haiku 5.5 | Adaptive only | On | `"enabled"`, `"disabled"`2 |
3435| Claude Opus 4.6 | Adaptive, extended (deprecated)1 | Off | None |
3536| Claude Sonnet 4.6 | Adaptive, extended (deprecated)1 | Off | None |
3637| Claude Opus 4.5 | Extended only | Off | `"adaptive"` |
from line 39
3839| Claude Sonnet 4.5 | Extended only | Off | `"adaptive"` |
3940
4041*1 `enabled` and `budget_tokens` still work on these models but are deprecated; use adaptive thinking instead.*\
41*2 Claude Opus 5 accepts `"disabled"` at [effort](https://platform.claude.com/docs/en/build-with-claude/effort) `high` or below; combining it with effort `xhigh` or `max` returns a 400 error. This restriction is enforced on each request.*\
42*2 Claude Opus 5 and Claude Haiku 5.5 accept `"disabled"` at [effort](https://platform.claude.com/docs/en/build-with-claude/effort) `high` or below; combining it with effort `xhigh` or `max` returns a 400 error. On Claude Opus 5, this restriction is enforced on each request. On Claude Haiku 5.5, a per-message effort that differs from the level in effect also returns a 400 error.*\
4243*3 Claude Sonnet 5.5 accepts `"between_tools"` at effort `high` or below. Combining it with effort `xhigh` or `max` returns a 400 error, and so does a per-message effort that differs from the level in effect.*
4344
44Models marked `Always on` cannot turn thinking off. Models marked `On` default to thinking. Claude Opus 5 and Claude Sonnet 5 accept `thinking: {type: "disabled"}`. On Claude Sonnet 5.5, send `thinking: {type: "between_tools"}` to turn off up-front thinking.
45Models marked `Always on` cannot turn thinking off. Models marked `On` default to thinking. Claude Opus 5, Claude Sonnet 5, and Claude Haiku 5.5 accept `thinking: {type: "disabled"}`. On Claude Sonnet 5.5, send `thinking: {type: "between_tools"}` to turn off up-front thinking.
4546
4647Earlier Claude 4 models (Claude Opus 4.1, Claude Sonnet 4, and Claude Opus 4) support extended thinking only. See [Model deprecations](https://platform.claude.com/docs/en/about-claude/model-deprecations) for their availability. Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, and Claude Mythos 5 are not available under [zero data retention](https://platform.claude.com/docs/en/manage-claude/api-and-data-retention#model-specific-data-retention-requirements) unless expressly authorized by Anthropic.
4748
from line 76
7576
7677Omit the `thinking` parameter; these models think without any configuration. If your goal was to keep thinking text out of responses, use `display: "omitted"` instead of disabling thinking; see [Controlling thinking display](https://platform.claude.com/docs/en/build-with-claude/thinking#controlling-thinking-display).
7778
78A 400 error on `"disabled"` can also occur on Claude Opus 5, which accepts `thinking: {type: "disabled"}` only at [effort](https://platform.claude.com/docs/en/build-with-claude/effort) `high` or below: combining it with effort `xhigh` or `max` is rejected. Lower the effort level, or leave thinking on.
79A 400 error on `"disabled"` can also occur on Claude Opus 5 and Claude Haiku 5.5, which accept `thinking: {type: "disabled"}` only at [effort](https://platform.claude.com/docs/en/build-with-claude/effort) `high` or below: combining it with effort `xhigh` or `max` is rejected. Lower the effort level, or set `thinking` to `{"type": "adaptive"}`.
7980
8081On Claude Sonnet 5.5, `thinking: {type: "disabled"}` returns a 400 error at every effort level. The message reads:
8182
from line 110
109110
110111Lower the effort to `high` or below. To run at `xhigh` or `max`, use adaptive thinking: omit the `thinking` field or send `thinking: {"type": "adaptive"}`. That's what the message means by "enable thinking". Claude Sonnet 5.5 rejects `"enabled"` with a 400 error.
111112
113Claude Haiku 5.5 returns the same message when a request sends `thinking: {type: "disabled"}` at effort `xhigh` or `max`. It accepts `"disabled"` only at `high` or below. Lower the effort, or set `thinking` to `{"type": "adaptive"}`.
114
112115## A 400 error says effort cannot change when thinking is disabled
113116
114117On Claude Sonnet 5.5, a request with `thinking: {type: "between_tools"}` whose [per-message effort](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) changes the level fails with a 400 error whose message reads:
from line 124
121124
122125Remove that per-message effort, or set it to the level in effect. To vary effort per turn, use adaptive thinking, which is what the message means by "enable thinking".
123126
127Claude Haiku 5.5 returns the same message when a request with `thinking: {type: "disabled"}` sets a per-message effort that differs from the level in effect. Remove that per-message effort, or set `thinking` to `{"type": "adaptive"}` to vary effort per turn.
128
124129## A 400 error says adaptive thinking is not supported
125130
126131The request fails with a 400 error whose message reads:
from line 152
147152
148153## A 400 error says a thinking block signature is invalid
149154
150A request to Claude Fable 5.1, Claude Opus 5.5, or Claude Sonnet 5.5 that replays earlier thinking blocks fails with a 400 `invalid_request_error` whose message reads:
155A request to Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, or Claude Haiku 5.5 that replays earlier thinking blocks fails with a 400 `invalid_request_error` whose message reads:
151156
152157```text wrap
153158messages.{i}.content.{j}: Invalid `signature` in `thinking` block. The block is bound to a different conversation. Remove the block, or set `thinking.block_binding.prefix_mismatch_behavior` to "drop_block".
from line 164
159164
160165If the message stops after ``Invalid `signature` in `thinking` block``, the signature itself didn't verify: it was truncated, altered, or sent back empty, and `prefix_mismatch_behavior` doesn't apply. Edited thinking text returns a different error. See [A 400 error says thinking blocks cannot be modified](https://platform.claude.com/docs/en/build-with-claude/thinking-troubleshooting#error-thinking-blocks-modified).
161166
162On Claude Fable 5.1, Claude Opus 5.5, and Claude Sonnet 5.5, the API accepts a replayed thinking block only while the `system` prompt, `tools`, and messages that preceded it are unchanged. See [Keeping the prefix unchanged](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#prefix-check). The error means something earlier in the conversation changed between requests: an edited, reordered, or removed turn, a per-turn reminder that was injected and later removed, a rebuilt `system` prompt or `tools` array, or client-side compaction that kept recent turns and their thinking verbatim. The check is enforced for new accounts created on or after August 31, 2026, and for any request that sets `thinking.block_binding.prefix_mismatch_behavior`. Server-side [compaction](https://platform.claude.com/docs/en/build-with-claude/compaction) and [context editing](https://platform.claude.com/docs/en/build-with-claude/context-editing) never trigger it.
167On Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, and Claude Haiku 5.5, the API accepts a replayed thinking block only while the `system` prompt, `tools`, and messages that preceded it are unchanged. See [Keeping the prefix unchanged](https://platform.claude.com/docs/en/build-with-claude/preserved-thinking#prefix-check). The error means something earlier in the conversation changed between requests: an edited, reordered, or removed turn, a per-turn reminder that was injected and later removed, a rebuilt `system` prompt or `tools` array, or client-side compaction that kept recent turns and their thinking verbatim. The check is enforced for new accounts created on or after August 31, 2026, and for any request that sets `thinking.block_binding.prefix_mismatch_behavior`. Server-side [compaction](https://platform.claude.com/docs/en/build-with-claude/compaction) and [context editing](https://platform.claude.com/docs/en/build-with-claude/context-editing) never trigger it.
163168
164169To fix it, keep the history append-only: pass earlier turns back exactly as sent and received, add instructions with a [mid-conversation system message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages) instead of editing `system` or `tools`, and let server-side [context editing](https://platform.claude.com/docs/en/build-with-claude/context-editing) or [compaction](https://platform.claude.com/docs/en/build-with-claude/compaction) do any trimming. Retrying the same request body doesn't clear the error. To continue this request without the invalidated reasoning, send the `thinking-binding-controls-2026-08-01` beta header and set `thinking.block_binding.prefix_mismatch_behavior` to `"drop_block"`. Alternatively, strip every `thinking` and `redacted_thinking` block from the history (at minimum the named block and every one after it, in that turn and all later turns), leave each turn's other blocks in place, and retry once. On Claude Sonnet 5.5, `block_binding` works only with `thinking: {"type": "adaptive"}`. With `between_tools`, keep the history append-only, or strip the thinking blocks from the edited turn on.
165170
manage-claude/inference-hooks-endpoint Changed · +35 / -17 lines
from line 148
148148
149149Anthropic sends an HTTPS `POST` to the URL your administrator configures. The whole configured URL is the endpoint: there is no fixed path suffix, so choose any path that suits your server.
150150
151Host your AI security server where Anthropic can reach it: an `https://` URL on port 443, on a publicly routable host (private, loopback, and carrier-grade NAT ranges are refused at connect time), with a certificate that validates against the public CA trust store, responding without redirects. The configured URL must be the final destination. Reverse-tunnel hosts (ngrok and similar tunnel services) are not supported: Anthropic's network policy blocks them. Host your server on a domain you control. [Configure Inference hooks](https://platform.claude.com/docs/en/manage-claude/inference-hooks-configuration) covers how your administrator sets and tests the URL.
151Host your AI security server where Anthropic can reach it: an `https://` URL on port 443, on a publicly routable host (private, loopback, and carrier-grade NAT ranges are refused at connect time), with a certificate that validates against the public CA trust store, responding without redirects. The host must have an IPv4 address, which Anthropic uses even when the host also has IPv6 addresses; a URL whose host is `localhost` or an IPv6 address is refused. The configured URL must be the final destination. Reverse-tunnel hosts (ngrok and similar tunnel services) are not supported: Anthropic's network policy blocks them. Host your server on a domain you control. [Configure Inference hooks](https://platform.claude.com/docs/en/manage-claude/inference-hooks-configuration) covers how your administrator sets and tests the URL.
152152
153153Every request carries these fixed headers, along with any [custom request headers](https://platform.claude.com/docs/en/manage-claude/inference-hooks-configuration) your administrator configured and, once your organization has a signing secret, the `webhook-*` signature headers described in [Verify the signature](https://platform.claude.com/docs/en/manage-claude/inference-hooks-endpoint#verify-the-signature):
154154
from line 219
219219
220220Each entry in `messages` has a `role` of `user` or `assistant` (tool results appear under the `user` role, matching the public Messages API content model) and a `content` array of blocks discriminated by `type`:
221221
222| Block `type` | Fields |
223| ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
224| `text` | `text`: the text content. |
225| `tool_use` | `id`: the identifier the matching tool result references. `tool_name`: the tool's name. `input`: the arguments the model passed to the tool. `tool_info`: on a tool call frame only (left out elsewhere, never `null`), an object that says who provides the tool; see [The tool call frame](https://platform.claude.com/docs/en/manage-claude/inference-hooks-endpoint#the-tool-call-frame). |
226| `tool_result` | `content`: the tool's output as text, with parts joined by newlines; binary parts such as images are replaced by placeholder markers, and raw bytes are never sent. `is_error`: whether the tool call failed. `tool_name`: the tool's name, so a policy can condition on tool identity without cross-referencing an earlier block. `tool_use_id`: the `id` of the matching `tool_use` block. |
227| `attachment` | `file_name`: the original file name or path. `media_type`: the attachment's media type. `size_bytes`: the size of the original file. `text`: the text content of the attachment when available, such as extracted document text, an audio transcript, or link metadata. Raw attachment bytes are never sent. |
222| Block `type` | Fields |
223| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
224| `text` | `text`: the text content. |
225| `tool_use` | `id`: the identifier the matching tool result references. `tool_name`: the tool's name. `input`: the arguments the model passed to the tool. `tool_info`: on a tool call frame only (left out elsewhere, never `null`), an object that says who runs or provides the tool; see [The tool call frame](https://platform.claude.com/docs/en/manage-claude/inference-hooks-endpoint#the-tool-call-frame). |
226| `tool_result` | `content`: the tool's output as text, with parts joined by newlines; binary parts such as images are replaced by placeholder markers, and raw bytes are never sent. `is_error`: whether the tool call failed. `tool_name`: the tool's name, so a policy can condition on tool identity without cross-referencing an earlier block. `tool_use_id`: the `id` of the matching `tool_use` block. |
227| `attachment` | `file_name`: the original file name or path. `media_type`: the attachment's media type. `size_bytes`: the size of the original file. `text`: the text content of the attachment when available, such as extracted document text, an audio transcript, or link metadata. Raw attachment bytes are never sent. |
228228
229229Apart from `type`, a `text` block's `text`, and a `tool_result` block's `content` and `is_error`, any of these fields can be `null` when the value isn't known; for example, an image arrives as an `attachment` block with `file_name` and `text` set to `null`.
230230
from line 252
252252
253253* `type` is `"tool_call"`.
254254* `messages` holds only the latest message, the `assistant` message Claude just produced: any `text` blocks and one `tool_use` block per tool call the frame lists, in the order the model produced them. Earlier conversation is left out, because the prompt frame sent before that model call carried it. Read the last entry of `messages`, because the protocol may later add earlier messages before it.
255* Each `tool_use` block carries a `tool_info` object that says who provides the tool.
255* Each `tool_use` block carries a `tool_info` object that says who runs or provides the tool.
256256
257257Where `session_id` is set, it is the same on both frames. The tool call frame has its own `request_id`, which is opaque like the prompt frame's.
258258
259`tool_info` says who provides the tool, not who runs it or what it can reach. It is one of four kinds, told apart by its `type` field, and each kind carries its own fields. Anthropic sends the first of the following kinds that fits the tool. New kinds may appear: accept a `type` you don't recognize, and for such a kind rely only on `type`.
259`tool_info` says who runs or provides the tool, not what the tool can reach. It is one of four kinds, told apart by its `type` field, and each kind carries its own fields. Anthropic sends the first of the following kinds that fits the tool. New kinds may appear: accept a `type` you don't recognize, and for such a kind rely only on `type`.
260260
261261An optional field that doesn't apply is left out, never `null`, so a `tool_info` can be just `{"type": "client"}`. `tool_name` is chosen by whoever defined the tool, and a server's `toolset_name` by whoever wrote the request, so don't treat either as a trust boundary.
262262
263263### Platform tools
264264
265A platform tool is one of the Claude API's predefined tools, such as web search or bash. It is `"platform"` whether Anthropic or the application that calls Claude runs it.
265A platform tool is one that the Claude API itself runs while it serves the request, such as web search or code execution.
266266
267267| Field | Present | Description |
268268| -------------- | -------- | ------------------------------------------------------------------------------------------------------------------- |
269269| `type` | Always | `"platform"` |
270270| `tool_type` | Always | The tool's versioned type, such as `web_search_20250305`. Match it exactly; don't parse a name or a date out of it. |
271| `toolset_name` | Optional | The toolset the tool belongs to, such as `browser`. |
271| `toolset_name` | Optional | The group of tools it belongs to. |
272272
273273### Application tools
274274
275An application tool is one that the Anthropic application making the request provides itself, such as claude.ai's own tools. It is `"application"` whether Anthropic or the application that calls Claude runs it.
275An application tool is one that the Anthropic application making the request provides itself, such as claude.ai's own tools.
276276
277277| Field | Present | Description |
278278| -------------- | -------- | --------------------------------- |
from line 311
311311
312312### Client tools
313313
314A client tool is any other tool. The application that calls Claude declares it and receives its calls. A call to a tool the request doesn't declare is also `"client"`.
314A client tool is any other tool, usually one that the application calling Claude runs. A tool that Anthropic defines but the calling application runs, such as bash or computer use, is a client tool too, and carries its versioned type in `tool_type`. A call to a tool the request doesn't declare is also `"client"`.
315315
316316A tool on an MCP server that Claude Code connects to directly from the user's computer, such as a local MCP server, is a client tool. Its `tool_info` is `{"type": "client"}`, and its `tool_name` is the name Claude Code gives it, in the form `mcp__<server>__<tool>`.
317317
318318A deny stops the call before Claude Code runs it. The exchange between Claude Code and a local server doesn't pass through Anthropic, so your AI security server sees the tool's result only when Claude Code sends it back. Its text is then in a `tool_result` block in the next prompt frame.
319319
320| Field | Present | Description |
321| -------------- | -------- | --------------------------------- |
322| `type` | Always | `"client"` |
323| `toolset_name` | Optional | The group of tools it belongs to. |
320| Field | Present | Description |
321| -------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
322| `type` | Always | `"client"` |
323| `tool_type` | Optional | The tool's versioned type, such as `bash_20250124`, when Anthropic defines the tool. Match it exactly; don't parse a name or a date out of it. |
324| `toolset_name` | Optional | The group of tools it belongs to, such as `browser`. |
324325
325326A `tool_use` block for a tool that the calling application declares:
326327
from line 335
334335 },
335336 "tool_info": {
336337 "type": "client"
338 }
339}
340```
341
342A `tool_use` block for bash, which Anthropic defines and the calling application runs:
343
344```json
345{
346 "type": "tool_use",
347 "id": "toolu_01HiJkLmNoPqRsTuVwXyZaBc",
348 "tool_name": "bash",
349 "input": {
350 "command": "ls -la reports/"
351 },
352 "tool_info": {
353 "type": "client",
354 "tool_type": "bash_20250124"
337355 }
338356}
339357```
managed-agents/reference Changed · +3 / -3 lines
from line 76
7676 </Tab>
7777
7878 <Tab title="System events">
79 | Type | Description |
80 | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
81 | `system.message` | Append privileged system-level context that applies to the accompanying turn and all subsequent turns. Supported on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, and Claude Sonnet 5.5. On an unsupported primary model the event is rejected with `model_does_not_support_mid_conversation_system`. |
79 | Type | Description |
80 | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
81 | `system.message` | Append privileged system-level context that applies to the accompanying turn and all subsequent turns. Supported on Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Mythos 5, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Sonnet 5.5, and Claude Haiku 5.5. On an unsupported primary model the event is rejected with `model_does_not_support_mid_conversation_system`. |
8282 </Tab>
8383
8484 <Tab title="Event deltas">
models/fable-5-1/overview Changed · +7 / -7 lines
from line 28
2828
2929## How it compares
3030
31| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
32| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------- | :------------------- | :------------- | :--------------- |
33| **Claude Fable 5.1** (this model) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
34| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
35| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
36| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Fastest | Extended | — | Feb 2025 |
31| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
32| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------- | :------------------- | :------------- | :--------------- |
33| **Claude Fable 5.1** (this model) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
34| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
35| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
36| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Fastest | Adaptive | `medium` | Jun 2026 |
3737
3838* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
39* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
39* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
4040* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
4141* **Latency:** Comparative latency, relative to the current lineup, as published in the models overview. Actual latency depends on prompt length, output length, and thinking effort.
4242* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
models/fable-5/overview Changed · +8 / -8 lines
from line 20
2020
2121## How it compares to the current lineup
2222
23| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
24| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------------------- | :------------- | :--------------- |
25| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
26| **Claude Fable 5** (this model) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jan 2026 |
27| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
28| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
29| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
23| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
24| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------------------- | :------------- | :--------------- |
25| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
26| **Claude Fable 5** (this model) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jan 2026 |
27| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
28| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
29| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
3030
3131* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
32* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
32* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
3333* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3434* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3535* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
models/haiku-4-5/overview Changed · +23 / -14 lines
from line 1
11---
22title: Claude Haiku 4.5
33url: https://platform.claude.com/docs/en/models/haiku-4-5/overview
4description: "Claude Haiku 4.5 at a glance: what it's for, model IDs on every platform, context window, output limits, pricing, availability, and the guides and resources for building with it."
4description: "Claude Haiku 4.5 reference: lifecycle status, model IDs on every platform, context window, output limits, pricing, and migration resources. Claude Haiku 5.5 is the current Haiku model."
55---
66
7**Latest.** Released October 15, 2025.
7**Legacy.** Released October 15, 2025.
88
99The fastest model with near-frontier intelligence
1010
11Although Claude Haiku 4.5 is still available, you should consider migrating to Claude Haiku 5.5 for improved performance. [See Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) · [Migrate to Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide)
12
1113Model ID: `claude-haiku-4-5-20251001`
1214
1315Context window: 200K tokens · Max output: 64K tokens · Input pricing: $1 / MTok · Output pricing: $5 / MTok
1416
15[Announcement](https://www.anthropic.com/news/claude-haiku-4-5) · [Migration guide](https://platform.claude.com/docs/en/models/haiku-4-5/migration-guide)
17[Announcement](https://www.anthropic.com/news/claude-haiku-4-5)
1618
1719## How it compares
1820
19| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
23| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
24| **Claude Haiku 4.5** (this model) | 200K | 64K | $1 / $5 | Fastest | Extended | — | Feb 2025 |
21| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
22| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------------------- | :------------- | :--------------- |
23| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
24| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
25| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
26| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
27| **Claude Haiku 4.5** (this model) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
2528
2629* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
27* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
30* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
2831* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
29* **Latency:** Comparative latency, relative to the current lineup, as published in the models overview. Actual latency depends on prompt length, output length, and thinking effort.
3032* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3133* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
3234* **Knowledge cutoff:** Reliable knowledge cutoff: the date through which the model’s knowledge is most extensive and reliable.
from line 68
6668| Max output | 64K tokens |
6769| [Thinking](https://platform.claude.com/docs/en/build-with-claude/thinking) | Extended |
6870| [Default effort](https://platform.claude.com/docs/en/build-with-claude/effort) | Not supported |
69| Comparative latency | Fastest |
7071| Input → output | Text and images → text |
7172| Reliable knowledge cutoff | Feb 2025 |
7273| Training data cutoff | Jul 2025 |
from line 76
7576
7677| Feature | Value |
7778| :---------------------------------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
78| [Status](https://platform.claude.com/docs/en/about-claude/model-deprecations) | Active (latest) |
79| [Status](https://platform.claude.com/docs/en/about-claude/model-deprecations) | Active (legacy) |
7980| Released | October 15, 2025 |
8081| Retirement | Not sooner than October 15, 2026 |
8182| Platforms | Claude API, [Amazon Bedrock](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock), [Amazon Bedrock (InvokeModel)](https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy), [Google Cloud](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai), [Microsoft Foundry](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry), [Claude Platform on AWS](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws) |
from line 90
8990## Resources
9091
9192<CardGroup cols={3}>
93 <Card title="Migrate to Claude Haiku 5.5" icon="arrows-left-right" href="https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide">
94 What changes when moving from Claude Haiku 4.5 to Claude Haiku 5.5.
95 </Card>
96
97 <Card title="Claude Haiku 5.5" icon="arrow-right" href="https://platform.claude.com/docs/en/models/haiku-5-5/overview">
98 The current Haiku model: overview, specs, and resources.
99 </Card>
100
92101 <Card title="Extended thinking" icon="brain" href="https://platform.claude.com/docs/en/build-with-claude/extended-thinking">
93102 Claude Haiku 4.5 supports manual extended thinking with `budget_tokens`.
94103 </Card>
from line 107
98107 </Card>
99108
100109 <Card title="Reduce latency" icon="gauge" href="https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latency">
101 Techniques that pair well with the fastest model in the lineup.
110 Techniques that pair well with a fast, low-cost model.
102111 </Card>
103112</CardGroup>
104113
models/mythos-5-1/overview Changed · +8 / -8 lines
from line 18
1818
1919## How it compares
2020
21| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
22| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------- | :------------------- | :------------- | :--------------- |
23| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
24| **Claude Mythos 5.1** (this model) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
25| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
26| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
27| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Fastest | Extended | — | Feb 2025 |
21| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
22| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------- | :------------------- | :------------- | :--------------- |
23| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
24| **Claude Mythos 5.1** (this model) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
25| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
26| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
27| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Fastest | Adaptive | `medium` | Jun 2026 |
2828
2929* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
30* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
30* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
3131* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3232* **Latency:** Comparative latency, relative to the current lineup, as published in the models overview. Actual latency depends on prompt length, output length, and thinking effort.
3333* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
models/mythos-5/overview Changed · +8 / -8 lines
from line 18
1818
1919## How it compares
2020
21| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
22| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------------------- | :------------- | :--------------- |
23| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
24| **Claude Mythos 5** (this model) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jan 2026 |
25| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
26| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
27| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
21| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
22| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------------------- | :------------- | :--------------- |
23| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
24| **Claude Mythos 5** (this model) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jan 2026 |
25| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
26| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
27| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
2828
2929* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
30* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
30* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
3131* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3232* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3333* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
models/opus-4-5/overview Changed · +8 / -8 lines
from line 16
1616
1717## How it compares to the current lineup
1818
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| **Claude Opus 4.5** (this model) | 200K | 64K | $5 / $25 | Extended | `high` | May 2025 |
24| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
25| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| **Claude Opus 4.5** (this model) | 200K | 64K | $5 / $25 | Extended | `high` | May 2025 |
24| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
25| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
2626
2727* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
2929* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3030* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3131* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
models/opus-4-6/overview Changed · +8 / -8 lines
from line 16
1616
1717## How it compares to the current lineup
1818
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :----------------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| **Claude Opus 4.6** (this model) | 1M | 128K | $5 / $25 | Adaptive (extended deprecated) | `high` | May 2025 |
24| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
25| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :----------------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| **Claude Opus 4.6** (this model) | 1M | 128K | $5 / $25 | Adaptive (extended deprecated) | `high` | May 2025 |
24| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
25| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
2626
2727* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
2929* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3030* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3131* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
models/opus-4-7/overview Changed · +8 / -8 lines
from line 16
1616
1717## How it compares to the current lineup
1818
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| **Claude Opus 4.7** (this model) | 1M | 128K | $5 / $25 | Adaptive | `high` | Jan 2026 |
24| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
25| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| **Claude Opus 4.7** (this model) | 1M | 128K | $5 / $25 | Adaptive | `high` | Jan 2026 |
24| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
25| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
2626
2727* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
2929* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3030* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3131* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
models/opus-4-8/overview Changed · +8 / -8 lines
from line 14
1414
1515## How it compares to the current lineup
1616
17| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
18| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------------------- | :------------- | :--------------- |
19| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
20| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
21| **Claude Opus 4.8** (this model) | 1M | 128K | $5 / $25 | Adaptive | `high` | Jan 2026 |
22| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
23| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
17| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
18| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------------------- | :------------- | :--------------- |
19| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
20| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
21| **Claude Opus 4.8** (this model) | 1M | 128K | $5 / $25 | Adaptive | `high` | Jan 2026 |
22| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
23| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
2424
2525* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
26* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
26* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
2727* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
2828* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
2929* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
models/opus-5-5/overview Changed · +7 / -7 lines
from line 22
2222
2323## How it compares
2424
25| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
26| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------- | :------------------- | :------------- | :--------------- |
27| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
28| **Claude Opus 5.5** (this model) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
29| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
30| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Fastest | Extended | — | Feb 2025 |
25| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
26| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------- | :------------------- | :------------- | :--------------- |
27| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
28| **Claude Opus 5.5** (this model) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
29| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
30| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Fastest | Adaptive | `medium` | Jun 2026 |
3131
3232* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
33* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
33* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
3434* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3535* **Latency:** Comparative latency, relative to the current lineup, as published in the models overview. Actual latency depends on prompt length, output length, and thinking effort.
3636* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
models/opus-5/overview Changed · +8 / -8 lines
from line 16
1616
1717## How it compares to the current lineup
1818
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| **Claude Opus 5** (this model) | 1M | 128K | $5 / $25 | Adaptive | `high` | May 2026 |
24| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
25| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| **Claude Opus 5** (this model) | 1M | 128K | $5 / $25 | Adaptive | `high` | May 2026 |
24| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
25| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
2626
2727* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
2929* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3030* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3131* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
models/overview Changed · +25 / -21 lines
from line 24
2424
2525If you're unsure which model to use, start with [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) for most workloads. Use [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5.5 at higher effort still fall short. All current models support text and image input, text output, multilingual capabilities, vision, and tool use. Each model's page lists the platforms it's available on.
2626
27| Feature | Claude Fable 5.1 | Claude Opus 5.5 | Claude Sonnet 5.5 | Claude Haiku 4.5 |
28| :-------------------------------------------------------------------------------------------------------- | :-------------------------------------------------------------------------------- | :------------------------------------------------------------------------------ | :---------------------------------------------------------------------------------- | :-------------------------------------------------------------------------------- |
29| Description | For demanding reasoning and long-horizon agentic work | For long-running agentic coding and knowledge work | The best combination of speed and intelligence | The fastest model with near-frontier intelligence |
30| Model page | [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) |
31| Comparative latency | Slower | Moderate | Fast | Fastest |
32| [Pricing](https://platform.claude.com/docs/en/about-claude/pricing) | $10 / input MTok, $50 / output MTok | $4 / input MTok, $20 / output MTok | $2 / input MTok, $10 / output MTok | $1 / input MTok, $5 / output MTok |
33| Claude API ID | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-4-5-20251001` |
34| [Thinking](https://platform.claude.com/docs/en/build-with-claude/thinking) | Adaptive (always on) | Adaptive (always on) | Adaptive | Extended |
35| [Default effort](https://platform.claude.com/docs/en/build-with-claude/effort) | `high` | `medium` | `high` | Not supported |
36| [Context window](https://platform.claude.com/docs/en/build-with-claude/context-windows) | 1M tokens | 1M tokens | 1M tokens | 200K tokens |
37| Max output | 128K tokens | 128K tokens | 128K tokens | 64K tokens |
38| Reliable knowledge cutoff | Jun 2026 | Jun 2026 | Jun 2026 | Feb 2025 |
39| Training data cutoff | Jun 2026 | Jun 2026 | Jun 2026 | Jul 2025 |
40| [Retirement](https://platform.claude.com/docs/en/about-claude/model-deprecations) | Not sooner than September 1, 2027 | Not sooner than September 22, 2027 | Not sooner than September 28, 2027 | Not sooner than October 15, 2026 |
41| Claude API alias | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-4-5` |
42| [Amazon Bedrock ID](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock) | `anthropic.claude-fable-5-1` | `anthropic.claude-opus-5-5` | `anthropic.claude-sonnet-5-5` | `anthropic.claude-haiku-4-5` |
43| [Google Cloud ID](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai) | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-4-5@20251001` |
44| [Microsoft Foundry ID](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry) | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-4-5` |
45| [Claude Platform on AWS ID](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws) | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-4-5` |
27| Feature | Claude Fable 5.1 | Claude Opus 5.5 | Claude Sonnet 5.5 | Claude Haiku 5.5 |
28| :-------------------------------------------------------------------------------------------------------- | :-------------------------------------------------------------------------------- | :------------------------------------------------------------------------------ | :---------------------------------------------------------------------------------- | :--------------------------------------------------------------------------------------- |
29| Description | For demanding reasoning and long-horizon agentic work | For long-running agentic coding and knowledge work | The best combination of speed and intelligence | For high-volume, latency-sensitive tasks such as classification, extraction, and routing |
30| Model page | [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) |
31| Comparative latency | Slower | Moderate | Fast | Fastest |
32| [Pricing](https://platform.claude.com/docs/en/about-claude/pricing) | $10 / input MTok, $50 / output MTok | $4 / input MTok, $20 / output MTok | $2 / input MTok, $10 / output MTok | From $0.10 / input MTok, From $0.50 / output MTok |
33| Claude API ID | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-5-5` |
34| [Thinking](https://platform.claude.com/docs/en/build-with-claude/thinking) | Adaptive (always on) | Adaptive (always on) | Adaptive | Adaptive |
35| [Default effort](https://platform.claude.com/docs/en/build-with-claude/effort) | `high` | `medium` | `high` | `medium` |
36| [Context window](https://platform.claude.com/docs/en/build-with-claude/context-windows) | 1M tokens | 1M tokens | 1M tokens | 1M tokens |
37| Max output | 128K tokens | 128K tokens | 128K tokens | 128K tokens |
38| Reliable knowledge cutoff | Jun 2026 | Jun 2026 | Jun 2026 | Jun 2026 |
39| Training data cutoff | Jun 2026 | Jun 2026 | Jun 2026 | Jun 2026 |
40| [Retirement](https://platform.claude.com/docs/en/about-claude/model-deprecations) | Not sooner than September 1, 2027 | Not sooner than September 22, 2027 | Not sooner than September 28, 2027 | Not sooner than October 7, 2027 |
41| Claude API alias | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-5-5` |
42| [Amazon Bedrock ID](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock) | `anthropic.claude-fable-5-1` | `anthropic.claude-opus-5-5` | `anthropic.claude-sonnet-5-5` | `anthropic.claude-haiku-5-5` |
43| [Google Cloud ID](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai) | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-5-5` |
44| [Microsoft Foundry ID](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry) | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-5-5` |
45| [Claude Platform on AWS ID](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws) | `claude-fable-5-1` | `claude-opus-5-5` | `claude-sonnet-5-5` | `claude-haiku-5-5` |
4646
4747* **Comparative latency:** Relative to the current lineup. Actual latency depends on prompt length, output length, and thinking effort.
4848* **Pricing:** Base price per million tokens. Batch API requests are 50% off; prompt cache reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for cache writes, long-context, and per-platform pricing.
from line 50
5050* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual thinking.type “enabled” + budget\_tokens mode on earlier models; it is deprecated on Claude Opus 4.6 and Claude Sonnet 4.6 and not accepted on later models.
5151* **Default effort:** The effort parameter’s default on the Claude API. Set effort explicitly to use a different level.
5252* **Context window:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
53* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
53* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
5454* **Reliable knowledge cutoff:** The date through which the model’s knowledge is most extensive and reliable. Training data cutoff (under Additional details) is the broader range of data used. See Anthropic’s Transparency Hub for details.
5555* **Retirement:** Anthropic’s commitment for Anthropic-operated platforms (Claude API, Claude Platform on AWS, Microsoft Foundry). Amazon Bedrock and Google Cloud set their own dates.
5656* **Claude API alias:** For models before the 4.6 generation, the alias is a convenience pointer that resolves to the dated ID. Dateless IDs are their own pinned snapshot; the alias row repeats them.
from line 61
6161
6262See [Model IDs and versioning](https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions) and [Pricing](https://platform.claude.com/docs/en/about-claude/pricing).
6363
64Legacy models (still available): [Claude Fable 5](https://platform.claude.com/docs/en/models/fable-5/overview), [Claude Opus 5](https://platform.claude.com/docs/en/models/opus-5/overview), [Claude Opus 4.8](https://platform.claude.com/docs/en/models/opus-4-8/overview), [Claude Opus 4.7](https://platform.claude.com/docs/en/models/opus-4-7/overview), [Claude Opus 4.6](https://platform.claude.com/docs/en/models/opus-4-6/overview), [Claude Opus 4.5](https://platform.claude.com/docs/en/models/opus-4-5/overview), [Claude Sonnet 5](https://platform.claude.com/docs/en/models/sonnet-5/overview), [Claude Sonnet 4.6](https://platform.claude.com/docs/en/models/sonnet-4-6/overview).
64Legacy models (still available): [Claude Fable 5](https://platform.claude.com/docs/en/models/fable-5/overview), [Claude Opus 5](https://platform.claude.com/docs/en/models/opus-5/overview), [Claude Opus 4.8](https://platform.claude.com/docs/en/models/opus-4-8/overview), [Claude Opus 4.7](https://platform.claude.com/docs/en/models/opus-4-7/overview), [Claude Opus 4.6](https://platform.claude.com/docs/en/models/opus-4-6/overview), [Claude Opus 4.5](https://platform.claude.com/docs/en/models/opus-4-5/overview), [Claude Sonnet 5](https://platform.claude.com/docs/en/models/sonnet-5/overview), [Claude Sonnet 4.6](https://platform.claude.com/docs/en/models/sonnet-4-6/overview), [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview).
6565
6666Once you've picked a model, [learn how to make your first API call](https://platform.claude.com/docs/en/get-started). To understand how model IDs, aliases, and snapshots work, see [Model IDs and versioning](https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions); for the reliable-knowledge and training-data cutoffs behind each model, see [Anthropic's Transparency Hub](https://www.anthropic.com/transparency).
6767
from line 72
7272Each model in the response also has a `line` field, which names the model line it belongs to. Claude Opus 4.5 and Claude Opus 4.6 both report `opus`. Use `line` to group models, for example, in a model picker. `line` is `null` when a model belongs to no line. Read `line` instead of inferring it from the model's `id`. Anthropic might add more lines, so don't treat the set of values as fixed.
7373
7474Each model's `capabilities` object includes `thinking.types.disabled`, which reports whether the model accepts `thinking: {type: "disabled"}`, the setting that [turns thinking off](https://platform.claude.com/docs/en/build-with-claude/thinking#turning-thinking-off). `supported` is `false` when the model rejects `"disabled"` with a 400 error, and `true` on a model that doesn't support thinking. Even when `supported` is `true`, the API can still reject a `"disabled"` request for another reason. One such reason is an [effort](https://platform.claude.com/docs/en/build-with-claude/effort) level that the model doesn't allow with thinking off.
75
76Each model's `capabilities` object also includes `server_tools`, which reports whether the model accepts the [web search](https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool) and [code execution](https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool) tools. It doesn't cover other server tools, such as web fetch. `server_tools.web_search.supported` and `server_tools.code_execution.supported` are `true` when the model accepts at least one version of that tool, not necessarily every version. `server_tools.supported` is `true` when the model accepts at least one of the two tools. Even when a tool is supported, your organization's settings can still cause a request that uses it to fail. For example, an administrator can [disable web search](https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool#how-to-use-web-search).
77
78The top-level `code_execution` capability is a different check. It reports whether code that Claude runs in the code execution tool can call the request's other tools, as in [programmatic tool calling](https://platform.claude.com/docs/en/agents-and-tools/tool-use/programmatic-tool-calling). For Claude Haiku 4.5, for example, the Models API reports `server_tools.code_execution.supported` as `true` and `code_execution.supported` as `false`.
7579
7680## Prompt and output performance
7781
models/sonnet-4-5/overview Changed · +8 / -8 lines
from line 16
1616
1717## How it compares to the current lineup
1818
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
24| **Claude Sonnet 4.5** (this model) | 200K | 64K | $3 / $15 | Extended | — | Jan 2025 |
25| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
24| **Claude Sonnet 4.5** (this model) | 200K | 64K | $3 / $15 | Extended | — | Jan 2025 |
25| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
2626
2727* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
2929* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3030* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3131* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
models/sonnet-4-6/overview Changed · +8 / -8 lines
from line 16
1616
1717## How it compares to the current lineup
1818
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :----------------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
24| **Claude Sonnet 4.6** (this model) | 1M | 128K | $3 / $15 | Adaptive (extended deprecated) | `high` | Aug 2025 |
25| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :----------------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
24| **Claude Sonnet 4.6** (this model) | 1M | 128K | $3 / $15 | Adaptive (extended deprecated) | `high` | Aug 2025 |
25| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
2626
2727* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
2929* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3030* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3131* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
models/sonnet-5-5/overview Changed · +7 / -7 lines
from line 30
3030
3131## How it compares
3232
33| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
34| :-------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------- | :------------------- | :------------- | :--------------- |
35| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
36| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
37| **Claude Sonnet 5.5** (this model) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
38| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Fastest | Extended | — | Feb 2025 |
33| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
34| :-------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------- | :------------------- | :------------- | :--------------- |
35| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | `high` | Jun 2026 |
36| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | `medium` | Jun 2026 |
37| **Claude Sonnet 5.5** (this model) | 1M | 128K | $2 / $10 | Fast | Adaptive | `high` | Jun 2026 |
38| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Fastest | Adaptive | `medium` | Jun 2026 |
3939
4040* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
41* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
41* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
4242* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
4343* **Latency:** Comparative latency, relative to the current lineup, as published in the models overview. Actual latency depends on prompt length, output length, and thinking effort.
4444* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
models/sonnet-5/overview Changed · +8 / -8 lines
from line 16
1616
1717## How it compares to the current lineup
1818
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
24| **Claude Sonnet 5** (this model) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jan 2026 |
25| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | 200K | 64K | $1 / $5 | Extended | — | Feb 2025 |
19| Model | Context | Max output | Price / MTok | Thinking | Default effort | Knowledge cutoff |
20| :---------------------------------------------------------------------------------- | :------ | :--------- | :----------------- | :------------------- | :------------- | :--------------- |
21| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | 1M | 128K | $10 / $50 | Adaptive (always on) | `high` | Jun 2026 |
22| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | 1M | 128K | $4 / $20 | Adaptive (always on) | `medium` | Jun 2026 |
23| [Claude Sonnet 5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jun 2026 |
24| **Claude Sonnet 5** (this model) | 1M | 128K | $2 / $10 | Adaptive | `high` | Jan 2026 |
25| [Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/overview) | 1M | 128K | From $0.10 / $0.50 | Adaptive | `medium` | Jun 2026 |
2626
2727* **Context:** 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Claude Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
28* **Max output:** Synchronous Messages API limit. On the Message Batches API, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Sonnet 5, Claude Haiku 5.5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, and Claude Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
2929* **Price / MTok:** Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5). See Pricing for the full list.
3030* **Thinking:** Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
3131* **Default effort:** The effort parameter’s default on the Claude API. Models without a value don’t support the parameter.
release-notes/overview Changed · +30 / -0 lines
### October 8, 2026 ### October 7, 2026 ### October 7, 2026 ### October 7, 2026 ### October 7, 2026 ### October 7, 2026 ### October 6, 2026
from line 12
1212 For updates to Claude Code, see the [complete CHANGELOG.md](https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md) in the `claude-code` repository.
1313</Tip>
1414
15### October 8, 2026
16
17* The [Compliance API](https://platform.claude.com/docs/en/manage-claude/compliance-api) chat endpoints now also return chats from the unified Claude experience, in beta for Claude Enterprise organizations, with your existing Compliance Access Key. See [Retrieve and delete chats, files, and projects](https://platform.claude.com/docs/en/manage-claude/compliance-content-data).
18
19### October 7, 2026
20
21* We've lowered the price of prompt cache reads on Claude Sonnet 5.5 from $0.20 USD to $0.10 USD per million tokens: 0.05x the base input price instead of 0.1x. Cache writes and all other prices are unchanged. See [Prompt caching pricing](https://platform.claude.com/docs/en/about-claude/pricing#prompt-caching).
22
23### October 7, 2026
24
25* We've launched **Claude Haiku 5.5** (`claude-haiku-5-5`), our most capable model tuned for high-volume and latency-sensitive work. It has a [1M token context window](https://platform.claude.com/docs/en/build-with-claude/context-windows), 128k max output tokens, and [adaptive thinking](https://platform.claude.com/docs/en/build-with-claude/thinking) with the [effort parameter](https://platform.claude.com/docs/en/build-with-claude/effort). It's available on the Claude API, [Claude in Amazon Bedrock](https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock), [Claude Platform on AWS](https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws), [Claude on Google Cloud](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai), and [Claude in Microsoft Foundry](https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry). See [What's new in Claude Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/whats-new-haiku-5-5).
26* Code written for Claude Haiku 4.5 can break on Claude Haiku 5.5. Manual extended thinking (`budget_tokens`) returns a 400 error, and adaptive thinking is on by default, so a response can begin with `thinking` blocks. The same text also counts as more tokens. See the [migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide). For model-specific prompting patterns, see [Prompting Claude Haiku 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5).
27
28### October 7, 2026
29
30* The Python and TypeScript SDKs now include classes, in beta, for the [browser use tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/browser-use-tool) and the [computer use tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool). You subclass one and write one method per tool against your own browser or desktop automation. The SDK runs the tool loop, the URL and file policies you set for the browser, and your approval callback. See [Browser and computer use with the SDK toolsets](https://platform.claude.com/docs/en/agents-and-tools/tool-use/browser-use-sdk).
31
32### October 7, 2026
33
34* Claude Max and Team plans now include monthly API credits. To learn how to claim them, see [API credits for Max and Team plans](https://platform.claude.com/docs/en/about-claude/api-credits-for-subscribers).
35
36### October 7, 2026
37
38* In Claude Managed Agents, a cloud environment with `limited` networking now also applies its `allowed_hosts` to the `web_search` and `web_fetch` tools. A `web_fetch` call for a URL on a host that `allowed_hosts` does not match returns a `url_not_allowed` error result to the agent. `web_search` omits results from such hosts. When `allowed_hosts` lists no hosts, neither tool returns a page or a search result. `allow_package_managers` and `allow_mcp_servers` add no hosts for these tools. To let the tools reach a host, add it to `allowed_hosts`, which also opens it to the sandbox. `unrestricted` networking and self-hosted environments do not limit these tools. See [Environment networking](https://platform.claude.com/docs/en/managed-agents/environments#networking).
39* With `limited` networking, creating a session fails with a 400 error when an enabled web tool's `allowed_domains` has an entry not within `allowed_hosts`. So does a session update that adds such an entry. An `allowed_hosts` entry matches one exact host unless it starts with `*.`, so `docs.example.com` is not within `["example.com"]`. To fix the error, add the host to `allowed_hosts` or remove the entry from `allowed_domains`. See [Restrict web search and web fetch domains](https://platform.claude.com/docs/en/managed-agents/tools-web-restrictions).
40
41### October 6, 2026
42
43* We've added `capabilities.server_tools` to the [Models API](https://platform.claude.com/docs/en/api/models/list). `GET /v1/models` and `GET /v1/models/{model_id}` now report whether each model accepts the web search and code execution tools. To check whether a model accepts the code execution tool, read `capabilities.server_tools.code_execution`. The top-level `capabilities.code_execution` reports whether code can call your request's other tools. See [Using the Models API](https://platform.claude.com/docs/en/models/overview#using-the-models-api).
44
1545### October 5, 2026
1646
1747* We've added `capabilities.thinking.types.disabled` to the [Models API](https://platform.claude.com/docs/en/api/models/list). `GET /v1/models` and `GET /v1/models/{model_id}` now report whether each model accepts `thinking: {type: "disabled"}`, which turns thinking off. See [Using the Models API](https://platform.claude.com/docs/en/models/overview#using-the-models-api).
test-and-evaluate/strengthen-guardrails/mitigate-jailbreaks Changed · +3 / -3 lines
from line 15
1515
1616In this threat model, a user is deliberately crafting inputs to manipulate your application into producing content or taking actions you don't want it to. These mitigations strengthen your application's guardrails:
1717
18* **Harmlessness screens:** Use a lightweight model like Claude Haiku 4.5 to pre-screen user input before it reaches your main conversation. Use [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) to constrain the response to a simple classification.
18* **Harmlessness screens:** Use a lightweight model like Claude Haiku 5.5 to pre-screen user input before it reaches your main conversation. Use [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) to constrain the response to a simple classification. Claude Haiku 5.5 runs [safety classifiers](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback) that can decline the screening request itself, so treat a response with `stop_reason: "refusal"` as a harmful verdict.
1919
2020 <Accordion title="Example: Harmlessness screen for content moderation">
2121 ```text User wrap
from line 115
115115
116116* **Limit Claude's access to sensitive data and actions.** Apply the principle of least privilege so that a successful injection can do minimal damage: don't give Claude access to secrets it doesn't need, run tools in sandboxed environments, and scope permissions as narrowly as possible.
117117
118* **Screen tool outputs before Claude acts on them.** Apply the same lightweight-model screening pattern you use for user input to the content your tools return. Run each tool, pass its raw output to a small classifier call with Claude Haiku 4.5, and only return the content as a `tool_result` block if the screen reports no injection attempt. Use [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) so the classifier's verdict is a parseable value your application can branch on.
118* **Screen tool outputs before Claude acts on them.** Apply the same lightweight-model screening pattern you use for user input to the content your tools return. Run each tool, pass its raw output to a small classifier call with Claude Haiku 5.5, and only return the content as a `tool_result` block if the screen reports no injection attempt. Use [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) so the classifier's verdict is a parseable value your application can branch on.
119119
120120 <Accordion title="Example: Injection screen for tool output">
121121 ```text User wrap
from line 147
147147 }
148148 ```
149149
150 If `injection_suspected` is `true`, return an error or a stripped summary in the `tool_result` block instead of the raw content, and consider surfacing the attempt to the user.
150 If `injection_suspected` is `true`, return an error or a stripped summary in the `tool_result` block instead of the raw content, and consider surfacing the attempt to the user. Treat a `stop_reason: "refusal"` response from Claude Haiku 5.5 the same way: its safety classifiers can decline the screening request itself, which leaves no verdict to read.
151151 </Accordion>
152152
153153 You can also apply the input-validation patterns from the previous section to tool results before passing them to Claude.
test-and-evaluate/strengthen-guardrails/reduce-latency Changed · +70 / -37 lines
from line 1
11---
22title: Reducing latency
33url: https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latency
4description: Reduce Claude's response latency by choosing a faster model like Claude Haiku 4.5, trimming prompt and output tokens, and streaming responses.
4description: Reduce Claude's response latency by choosing a faster model like Claude Haiku 5.5, trimming prompt and output tokens, and streaming responses.
55---
66
77Latency refers to the time it takes for the model to process a prompt and generate an output. Latency can be influenced by various factors, such as the size of the model, the complexity of the prompt, and the underlying infrastructure supporting the model and point of interaction.
from line 29
2929
3030One of the most direct ways to reduce latency is to select the appropriate model for your use case. Anthropic offers a [range of models](https://platform.claude.com/docs/en/models/overview) with different capabilities and performance characteristics. Consider your specific requirements and choose the model that best fits your needs in terms of speed and output quality.
3131
32For speed-critical applications, **Claude Haiku 4.5** offers the fastest response times while maintaining high intelligence:
32For speed-critical applications, **Claude Haiku 5.5** offers the fastest response times while maintaining high intelligence. [Effort](https://platform.claude.com/docs/en/build-with-claude/effort) is its main control for speed and cost. See [Use effort to control thinking](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#use-effort-to-control-thinking). The following example runs it at `low`, the cheapest and fastest level, and leaves room in `max_tokens` for thinking:
3333
3434<CodeGroup>
3535 ```bash cURL
36 # For time-sensitive applications, use Claude Haiku 4.5
36 # For time-sensitive applications, use Claude Haiku 5.5 at low effort
3737 curl https://api.anthropic.com/v1/messages \
3838 -H "x-api-key: $ANTHROPIC_API_KEY" \
3939 -H "anthropic-version: 2023-06-01" \
4040 -H "content-type: application/json" \
4141 -d '{
42 "model": "claude-haiku-4-5",
43 "max_tokens": 100,
42 "model": "claude-haiku-5-5",
43 "max_tokens": 1024,
44 "output_config": {"effort": "low"},
4445 "messages": [{"role": "user", "content": "Summarize this customer feedback in 2 sentences: [feedback text]"}]
4546 }'
4647 ```
4748
4849 ```bash CLI
49 # For time-sensitive applications, use Claude Haiku 4.5
50 # For time-sensitive applications, use Claude Haiku 5.5 at low effort
5051 ant messages create \
51 --model claude-haiku-4-5 \
52 --max-tokens 100 \
52 --model claude-haiku-5-5 \
53 --max-tokens 1024 \
54 --output-config '{effort: low}' \
5355 --message '{"role": "user", "content": "Summarize this customer feedback in 2 sentences: [feedback text]"}'
5456 ```
5557
from line 58
5658 ```python Python
5759 client = anthropic.Anthropic()
5860
59 # For time-sensitive applications, use Claude Haiku 4.5
61 # For time-sensitive applications, use Claude Haiku 5.5 at low effort
6062 message = client.messages.create(
61 model="claude-haiku-4-5",
62 max_tokens=100,
63 model="claude-haiku-5-5",
64 max_tokens=1024,
65 output_config={"effort": "low"},
6366 messages=[
6467 {
6568 "role": "user",
from line 70
6770 }
6871 ],
6972 )
70 print(message.content[0].text)
73 print(next(block.text for block in message.content if block.type == "text"))
7174 ```
7275
7376 ```typescript TypeScript
7477 const client = new Anthropic();
7578
76 // For time-sensitive applications, use Claude Haiku 4.5
79 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
7780 const message = await client.messages.create({
78 model: "claude-haiku-4-5",
79 max_tokens: 100,
81 model: "claude-haiku-5-5",
82 max_tokens: 1024,
83 output_config: { effort: "low" },
8084 messages: [
8185 {
8286 role: "user",
from line 95
9195 ```csharp C#
9296 AnthropicClient client = new();
9397
94 // For time-sensitive applications, use Claude Haiku 4.5
98 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
9599 var parameters = new MessageCreateParams
96100 {
97 Model = Model.ClaudeHaiku4_5,
98 MaxTokens = 100,
101 Model = Model.ClaudeHaiku5_5,
102 MaxTokens = 1024,
103 OutputConfig = new() { Effort = Effort.Low },
99104 Messages = [
100105 new()
101106 {
from line 110
105110 ]
106111 };
107112 var message = await client.Messages.Create(parameters);
108 message.Content[0].TryPickText(out var textBlock);
109 Console.WriteLine(textBlock?.Text);
113 foreach (var block in message.Content)
114 {
115 if (block.TryPickText(out var textBlock))
116 {
117 Console.WriteLine(textBlock.Text);
118 break;
119 }
120 }
110121 ```
111122
112123 ```go Go
113124 client := anthropic.NewClient()
114125
115 // For time-sensitive applications, use Claude Haiku 4.5
126 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
116127 message, err := client.Messages.New(context.TODO(), anthropic.MessageNewParams{
117 Model: anthropic.ModelClaudeHaiku4_5,
118 MaxTokens: 100,
128 Model: anthropic.ModelClaudeHaiku5_5,
129 MaxTokens: 1024,
130 OutputConfig: anthropic.OutputConfigParam{
131 Effort: anthropic.OutputConfigEffortLow,
132 },
119133 Messages: []anthropic.MessageParam{
120134 anthropic.NewUserMessage(anthropic.NewTextBlock("Summarize this customer feedback in 2 sentences: [feedback text]")),
121135 },
from line 137
123137 if err != nil {
124138 log.Fatal(err)
125139 }
126 fmt.Println(message.Content[0].Text)
140 for _, block := range message.Content {
141 if textBlock, ok := block.AsAny().(anthropic.TextBlock); ok {
142 fmt.Println(textBlock.Text)
143 break
144 }
145 }
127146 ```
128147
129148 ```java Java
130149 AnthropicClient client = AnthropicOkHttpClient.fromEnv();
131150
132 // For time-sensitive applications, use Claude Haiku 4.5
151 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
133152 MessageCreateParams params = MessageCreateParams.builder()
134 .model(Model.CLAUDE_HAIKU_4_5)
135 .maxTokens(100L)
153 .model(Model.CLAUDE_HAIKU_5_5)
154 .maxTokens(1024L)
155 .outputConfig(OutputConfig.builder()
156 .effort(OutputConfig.Effort.LOW)
157 .build())
136158 .addUserMessage("Summarize this customer feedback in 2 sentences: [feedback text]")
137159 .build();
138160 Message message = client.messages().create(params);
139 IO.println(message.content().get(0).text().map(TextBlock::text).orElse(""));
161 IO.println(message.content().stream()
162 .flatMap(block -> block.text().stream())
163 .map(TextBlock::text)
164 .findFirst()
165 .orElse(""));
140166 ```
141167
142168 ```php PHP
143169 $client = new Client();
144170
145 // For time-sensitive applications, use Claude Haiku 4.5
171 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
146172 $message = $client->messages->create(
147 maxTokens: 100,
173 maxTokens: 1024,
148174 messages: [['role' => 'user', 'content' => 'Summarize this customer feedback in 2 sentences: [feedback text]']],
149 model: 'claude-haiku-4-5',
175 model: 'claude-haiku-5-5',
176 outputConfig: ['effort' => 'low'],
150177 );
151 echo $message->content[0]->text;
178 foreach ($message->content as $block) {
179 if ($block->type === 'text') {
180 echo $block->text;
181 break;
182 }
183 }
152184 ```
153185
154186 ```ruby Ruby
155187 client = Anthropic::Client.new
156188
157 # For time-sensitive applications, use Claude Haiku 4.5
189 # For time-sensitive applications, use Claude Haiku 5.5 at low effort
158190 message = client.messages.create(
159 model: "claude-haiku-4-5",
160 max_tokens: 100,
191 model: "claude-haiku-5-5",
192 max_tokens: 1024,
193 output_config: { effort: :low },
161194 messages: [{ role: "user", content: "Summarize this customer feedback in 2 sentences: [feedback text]" }]
162195 )
163 puts message.content.first.text
196 puts message.content.find { |block| block.type == :text }&.text
164197 ```
165198</CodeGroup>
166199
from line 222
189222
190223 tokens, the response will be cut off, perhaps mid-sentence or mid-word, so this is a blunt technique that might require post-processing and is usually most appropriate for multiple choice or short answer responses where the answer comes right at the beginning.
191224 </Note>
192* **Experiment with temperature:** The `temperature` [parameter](https://platform.claude.com/docs/en/api/messages/create) controls the randomness of the output. Lower values (for example, 0.2) can sometimes lead to more focused and shorter responses, while higher values (for example, 0.8) might result in more diverse but potentially longer outputs.
225* **Experiment with temperature:** The `temperature` [parameter](https://platform.claude.com/docs/en/api/messages/create) controls the randomness of the output. Lower values (for example, 0.2) can sometimes lead to more focused and shorter responses, while higher values (for example, 0.8) might result in more diverse but potentially longer outputs. Claude Haiku 5.5 accepts only the default `temperature` and returns a 400 error for any other value, so lower its [effort](https://platform.claude.com/docs/en/build-with-claude/effort) instead.
193226
194227Finding the right balance among prompt clarity, output quality, and token count might require some experimentation.
195228