One read of Claude Developer Platformapi-20260923T223709Z
14 pages moved out of 637 read.
Pages moved
14
significant first
Pages read
637
in this capture
Captured
22:37 UTC
Corpus hash
6096ef2a2dda
corpus-hash
What this read moved
1-14 of 14about-claude/models/optimizing-for-cost-and-intelligence Changed · +115 / -113 lines
from line 21
2121
2222Match your situation to a row.
2323
24| Your situation | Do this | Where |
25| ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
26| Any workload, any model | Turn on prompt caching and trim unneeded tokens; both are free | [Cache repeated context](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#cache-repeated-context) · [Trim tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
27| A person waits between turns | Use the 1-hour cache duration once about 1 turn in 20 follows a pause between 5 minutes and an hour and few gaps run over an hour. On Claude Fable 5.1, keep the 5-minute cache warm while pauses run minutes, and buy the 1-hour duration when pauses run toward an hour | [Pick the cache duration](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#pick-the-cache-duration) |
28| Costs are too high; quality is fine | Sweep effort down on your current model | [Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort) |
29| You are not on the latest model | Upgrade; the current model solves more tasks, at a cost per solved task from about 40% lower to about 20% higher | [Upgrade the model](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#upgrade-the-model) |
30| You are choosing or switching models | Compare on cost per completed task, not per token | [Compare models](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#compare-models-on-cost-per-task) |
31| Quality isn't good enough | If you lowered effort, restore it; otherwise try the next tier up at `low` effort | [Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort) · [Compare models](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#compare-models-on-cost-per-task) |
32| Attempts end with `stop_reason: max_tokens` | Raise `max_tokens`; 64,000 covered all but 2 of 14,000 turns measured at the default effort, and 128,000 cost nothing extra per solved task | [Set budgets](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
33| You can check outputs (tests, a verifier) | Run everything at low effort and re-run failures at `high`; on the coding benchmark measured, the pass rate held at about half the cost | [Re-run failures](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#re-run-failures-at-higher-effort) |
34| Agent loops with a few very costly runs | Set a task budget (beta; check the support table for which models), a Claude Managed Agents session budget, and a workspace spend limit | [Set budgets](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
35| A lower-cost model stalls only on hard decisions | Add a frontier advisor. It pays off when priced well above the executor and actually consulted, so first price the advisor's model alone at low effort and measure the consult rate | [Advisor strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#advisor-strategy-escalate-hard-decisions) |
36| The work exceeds one context window | Delegate partitions to cheaper workers | [Orchestrator strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#orchestrator-strategy-delegate-bulk-work) |
24| Your situation | Do this | Where |
25| ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
26| Any workload, any model | Turn on prompt caching and trim unneeded tokens; both are free | [Cache repeated context](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#cache-repeated-context) · [Trim tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
27| A person waits between turns | Use the 1-hour cache duration once about 1 turn in 20 follows a pause between 5 minutes and an hour and few gaps run over an hour. On Claude Fable 5.1, keep the 5-minute cache warm while pauses run minutes, and buy the 1-hour duration when pauses run toward an hour. On Claude Opus 5.5, keep the 5-minute cache warm instead when only a turn or two in 20 follow a pause of up to about half an hour | [Pick the cache duration](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#pick-the-cache-duration) |
28| Costs are too high; quality is fine | Sweep effort down on your current model | [Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort) |
29| You are not on the latest model | Upgrade; in Anthropic's measurements each newer model solved at least as many tasks as the one before it, usually for less per solved task | [Upgrade the model](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#upgrade-the-model) |
30| You are choosing or switching models | Compare on cost per completed task, not per token | [Compare models](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#compare-models-on-cost-per-task) |
31| Quality isn't good enough | If you lowered effort, restore it; otherwise try the next tier up at `low` effort | [Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort) · [Compare models](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#compare-models-on-cost-per-task) |
32| Attempts end with `stop_reason: max_tokens` | Raise `max_tokens`; 64,000 covered all but 2 of 14,000 turns measured at the default effort, and 128,000 cost nothing extra per solved task | [Set budgets](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
33| You can check outputs (tests, a verifier) | Run everything at low effort and re-run failures at `high`; on the coding benchmark measured, the pass rate held at about half the cost | [Re-run failures](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#re-run-failures-at-higher-effort) |
34| Agent loops with a few very costly runs | Set a task budget (beta; check the support table for which models), a Claude Managed Agents session budget, and a workspace spend limit | [Set budgets](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
35| A lower-cost model stalls only on hard decisions | Add a frontier advisor. It pays off when priced well above the executor and actually consulted, so first price the advisor's model alone at low effort and measure the consult rate | [Advisor strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#advisor-strategy-escalate-hard-decisions) |
36| The work exceeds one context window | Delegate partitions to cheaper workers | [Orchestrator strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#orchestrator-strategy-delegate-bulk-work) |
3737
3838These results are Anthropic-internal ([Benchmarks referenced](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs)) and directional, not guarantees, so measure on your own workload with the [four-step method](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#measure-on-your-own-workload).
3939
from line 61
6161
6262To decide, count the gaps between consecutive requests in a conversation:
6363
64* More than about 1 gap in 20 falls between 5 minutes and an hour, and gaps over an hour are rare: use the 1-hour duration.
65* Turns arrive seconds apart: stay on the 5-minute default. When nothing paused, it cost 15% less than the 1-hour setting on Claude Sonnet 5 and 11% less on Claude Opus 5.
64* More than about 1 gap in 20 falls between 5 minutes and an hour, and gaps over an hour are rare: use the 1-hour duration. On Claude Opus 5.5, when only 1 or 2 gaps in 20 fall in that range and none lasts more than about half an hour, keep the 5-minute cache warm instead, with the keep-alive requests described below.
65* Turns arrive seconds apart: stay on the 5-minute default. When nothing paused, it cost 15% less than the 1-hour setting on Claude Sonnet 5 and about 15% to 18% less on Claude Opus 5.5.
6666* Gaps over an hour are common: stay on the default. A gap over an hour expires both durations, and the 1-hour setting then re-writes the prefix at its higher write price, so it loses on each of those gaps. Of your pauses longer than 5 minutes, if about 60% or more also run past an hour, stay on the default; the 1-hour duration pays only when at least about 40% of long pauses end within the hour.
6767
68Anthropic measured the triage job from [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) with pauses inserted before some turns to simulate a person's delay[16](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs). On both models measured, the 1-hour cache became the cheaper setting once about 1 turn in 30 followed a pause, so the 1-in-20 rule leaves a margin, and the gap widens quickly past the crossover because every paused turn on the 5-minute setting re-writes the whole prefix. Every current model uses the same cache-write multipliers, and every model but Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5.5 the same read price, so the crossover is in the same range on the other models; Fable 5.1 is the case covered next. Accuracy stayed within run-to-run noise in every cell. The turn after a pause kept its warm-cache latency on the 1-hour setting (measured on Claude Sonnet 5 and Claude Opus 5, not on Claude Opus 5.5). The following chart plots cost per session against the share of paused turns on Claude Sonnet 5:
68Anthropic measured the triage job from [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) with pauses inserted before some turns to simulate a person's delay[16](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs). On Claude Sonnet 5 and Claude Opus 5.5, the 1-hour cache became the cheaper setting once about 1 turn in 30 followed a pause, so the 1-in-20 rule leaves a margin, and the gap widens quickly past the crossover because every paused turn on the 5-minute setting re-writes the whole prefix. Every current model uses the same cache-write multipliers, and every model but Claude Fable 5.1, Claude Mythos 5.1, and Claude Opus 5.5 the same read price, so the crossover is in the same range on the other models; Fable 5.1 is the case covered next. Accuracy stayed within run-to-run noise in every cell. The turn after a pause kept its warm-cache latency on the 1-hour setting (measured on Claude Sonnet 5 and Claude Opus 5, not on Claude Opus 5.5). The following chart plots cost per session against the share of paused turns on Claude Sonnet 5:
6969
7070
7171
72Anthropic also measured extra requests that keep the 5-minute cache warm. On Claude Sonnet 5 and Claude Opus 5 they saved nothing measurable over the 1-hour duration at any share of paused turns and cost more with a pause before every turn, so use the duration instead.
72Anthropic also measured extra requests that keep the 5-minute cache warm. On Claude Sonnet 5 they cost about 8% less than the 1-hour duration when 1 turn in 20 followed a pause, but about the same at 2 in 20; on Claude Opus 5, the previous Opus model, they saved nothing measurable. With a pause of 6 minutes or more before every turn they cost more on both models. Because the Claude Sonnet 5 saving was gone by 2 turns in 20, use the 1-hour duration instead on Claude Sonnet 5 and Claude Opus 5.
7373
7474On Claude Fable 5.1 the cheapest setting is a different one. Its [cache read](https://platform.claude.com/docs/en/about-claude/pricing#prompt-caching) costs 0.025x the input price ($0.25 per million tokens) while its cache writes keep the standard multipliers, so a keep-alive request that re-reads the prefix is cheap and the 1-hour duration's write premium is the larger bill. Anthropic measured the triage job on Claude Fable 5.1 with the same three settings[19](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs). Keeping the 5-minute cache warm cost 13% to 20% less per session than the 1-hour cache whenever pauses ran for minutes; only with pauses near 45 minutes did the 1-hour cache win, by about 12 cents a session. On Claude Fable 5.1, keep the 5-minute cache warm while a person is away for minutes, and buy the 1-hour duration when pauses run toward an hour:
7575
76
76
7777
78To keep the cache warm, send the previous request again with `max_tokens` set to 0 within 4 minutes of the previous request's start, and every 4 minutes after that, dropping `stream` if it was set. Count from the request's start, not its response's end: the [cache's 5-minute lifetime](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#how-prompt-caching-works) runs from the start of the request that wrote or refreshed the entry, so time the response spent generating counts against it. That is the [pre-warming request](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#pre-warming-the-cache): it refreshes the cache's lifetime, generates nothing, and bills only the cache read. Do not change a byte of the prefix, and do not use `max_tokens: 1`, which samples a token for no reason. Re-send the request's headers as well as its body: if your requests carry an `anthropic-beta` header (for a [task budget](https://platform.claude.com/docs/en/build-with-claude/task-budgets), say), the keep-alive request needs the same header, or the beta-gated fields in the replayed body are rejected. A `max_tokens: 0` request is rejected when the request sets `thinking.type: "enabled"` (the default adaptive thinking on Claude Fable 5.1 is fine), structured outputs, or a forced tool choice ([its limitations](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#limitations)); on those workloads, buy the 1-hour duration instead.
78On Claude Opus 5.5, whose cache read costs 0.05x the input price, keep-alive requests cost 8% to 13% less than the 1-hour duration when 5% or 10% of turns followed a pause of 6 to 32 minutes (at the default effort, `medium`; 10% to 18% less at `high`), but more with a pause before every turn: about 4% to 6% more at 6-minute pauses, rising to over 50% more at 45-minute pauses. So on Claude Opus 5.5, keep the 5-minute cache warm when only a turn or two in 20 follow a pause of up to about half an hour, and otherwise follow the list at the start of this section. These measurements sent keep-alive requests with `max_tokens: 1`. For the `max_tokens: 0` request described next, Anthropic's pre-launch API tests on Claude Opus 5.5 show that it writes the cache and that the next request reads it; whether it refreshes an existing entry was not measured on Opus 5.5.
7979
80To keep the cache warm, send the previous request again with `max_tokens` set to 0 within 4 minutes of the previous request's start, and every 4 minutes after that, dropping `stream` if it was set. Count from the request's start, not its response's end: the [cache's 5-minute lifetime](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#how-prompt-caching-works) runs from the start of the request that wrote or refreshed the entry, so time the response spent generating counts against it. That is the [pre-warming request](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#pre-warming-the-cache): it refreshes the cache's lifetime, generates nothing, and bills only the cache read. Do not change a byte of the prefix, and do not use `max_tokens: 1`, which samples a token for no reason. Re-send the request's headers as well as its body: if your requests carry an `anthropic-beta` header (for a [task budget](https://platform.claude.com/docs/en/build-with-claude/task-budgets), say), the keep-alive request needs the same header, or the beta-gated fields in the replayed body are rejected. A `max_tokens: 0` request is rejected when the request sets `thinking.type: "enabled"` (the default adaptive thinking on Claude Fable 5.1 is fine), structured outputs, or a forced tool choice ([its limitations](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#limitations)); on those workloads, buy the 1-hour duration instead. A `max_tokens: 0` request is also rejected when it carries the top-level `compaction` parameter, so don't re-send a compaction request from [compaction on demand](https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand) as a keep-alive request.
81
8082<CodeGroup exclude="shell:CLI, python, typescript, csharp, go, java, php, ruby">
8183 ```bash cURL
8284 # Within 4 minutes of the last request's start (time spent generating counts
from line 123
121123
122124Several things can break your cache during a task. Anything that changes per request, such as a timestamp or a queue position, placed ahead of the stable prefix turns every request into a full cache write: on the triage run in [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens), a 25-token status line at the front of the system prompt cost $4.24 per run instead of $0.59, more than running with caching off. Keep per-request text in the newest user turn.
123125
124The cache is a byte-exact prefix match over the request in order (tools, then system prompt, then messages), so a change anywhere invalidates everything after it. Changing [`effort`](https://platform.claude.com/docs/en/build-with-claude/effort) or the thinking configuration between requests invalidates the cache from that point onward, and on some models the tools and system prompt ahead of it as well; any edit to the system prompt invalidates the cache from that point onward; setting or changing an output format invalidates the cache for the whole conversation; adding, removing, or reordering a tool definition invalidates all of it. The [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#what-invalidates-the-cache) page lists these cases, apart from the output format, which [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs#prompt-modification-and-token-costs) covers. On the most recent models, change instructions with a [mid-conversation system message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages), a `{"role": "system"}` message appended to `messages`, instead of editing the top-level `system` field: the cached prefix stays intact. Check that page for which models support it. On models that support it, a [per-message effort change](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) leaves the cached prefix intact too. The stakes are highest on Claude Fable 5.1 and Claude Mythos 5.1: a break re-writes the prefix at 1.25x the input price instead of reading it at 0.025x, so on a 100,000-token prefix one broken turn costs $1.25 instead of $0.03, 50 times the read, against 12.5 times ($0.63 instead of $0.05) on Claude Opus 5.
126The cache is a byte-exact prefix match over the request in order (tools, then system prompt, then messages), so a change anywhere invalidates everything after it. Changing [`effort`](https://platform.claude.com/docs/en/build-with-claude/effort) or the thinking configuration between requests invalidates the cache from that point onward, and on some models the tools and system prompt ahead of it as well; any edit to the system prompt invalidates the cache from that point onward; setting or changing an output format invalidates the cache for the whole conversation; adding, removing, or reordering a tool definition invalidates all of it. The [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#what-invalidates-the-cache) page lists these cases, apart from the output format, which [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs#prompt-modification-and-token-costs) covers. On the most recent models, change instructions with a [mid-conversation system message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages), a `{"role": "system"}` message appended to `messages`, instead of editing the top-level `system` field: the cached prefix stays intact. Check that page for which models support it. On models that support it, a [per-message effort change](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) leaves the cached prefix intact too. The stakes are highest on Claude Fable 5.1 and Claude Mythos 5.1, where a break re-writes the prefix at 1.25x the input price instead of reading it at 0.025x. On a 100,000-token prefix, one broken turn there costs $1.25 instead of $0.03, 50 times the read; on Claude Opus 5.5 it costs $0.50 instead of $0.02, 25 times, and on the other current models 12.5 times.
125127
126128Anthropic measured this on the triage agent's long sessions[18](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs). An effort change and an added tool made mid-session rewrote 39,000 and 60,000 cached tokens, and those sessions cost $0.95 per session. The same two changes on the first request after compaction cost $0.75, and on the request that triggered the compaction $0.92, because the compaction's summarization pass then re-processed the 81,000-token context at the cache-write price: that summarization pass cost $0.21, against $0.04 when the same changes came one request later, with accuracy within run-to-run noise in every arm:
127129
from line 251
249251
250252Anthropic measured this on the SWE-bench Pro[3](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) subset, priced as a customer is billed:
251253
252
254
253255
254Claude Fable 5.1 at `low` effort solved 88.6% of tasks for $0.54 per solved task, against 77.4% for $0.84 from Claude Sonnet 5 at its default: 11 more points for 35% less per solved task, despite a per-token price five times higher. It does not always win, though. On the same subset, which both models largely saturate and whose scores are not comparable to the public leaderboard, Claude Opus 5 alone matched Claude Fable 5.1 alone at the default (91.7% compared with 92.1%, inside run-to-run noise) at about 15% less per solved task ($1.01 against $1.19), and Opus 5 at `low` solved 84.0% for $0.25. And on long research loops the frontier model does more work, not less: on DeepResearch Bench II[7](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), Fable 5.1 at `low` scored 10 points above Sonnet 5 (66% against 56%) at about four times the cost per task ($4.66 against $1.20), because it runs a longer research loop over a larger context. Claude Opus 5 at its default scored 71% on the same basis for $6.71 per task, above Fable 5.1 at its default (65% for $7.12), so on research too Fable 5.1 earns its price only at `low`.
256Claude Fable 5.1 at `low` effort solved 88.6% of tasks for $0.54 per solved task, against 77.4% for $0.84 from Claude Sonnet 5 at its default: 11 more points for 35% less per solved task, despite a per-token price five times higher. It does not always win, though. On the same subset, which Claude Opus 5.5 and Claude Fable 5.1 both largely saturate and whose scores are not comparable to the public leaderboard, Opus 5.5 at its default, `medium`, matched Fable 5.1 at its default (92.8% against 92.3%, inside run-to-run noise) for about a fifth of the cost per solved task ($0.22 against $1.19). At `low`, Opus 5.5 solved 87.4% for $0.12. These figures use the 478 problems described in reference 3. And on long research loops the frontier model does more work, not less: on DeepResearch Bench II[7](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), Fable 5.1 at `low` scored 10 points above Sonnet 5 (66% against 56%) at about four times the cost per task ($4.66 against $1.20), because it runs a longer research loop over a larger context. Claude Opus 5 at its default scored 71% on the same basis for $6.71 per task, above Fable 5.1 at its default (65% for $7.12), so on research too Fable 5.1 earns its price only at `low`.
255257
256For most agent workloads, start with Claude Fable 5.1 at `low` effort and raise effort where it misses. Per token it costs twice what Claude Opus 5 does on uncached input, but half as much on cached input ($0.25 against $0.50 per million), and in an agent loop cached input is the largest term. On the coding benchmark in [Advisor strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#advisor-strategy-escalate-hard-decisions), Fable 5.1 at `medium` matched Opus 5 at its default for about a third of the cost per attempt ($2.91 against $8.50). On Chartography[13](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), a chart-reading benchmark, Fable 5.1 at `low` scored 62.5 for $0.15 a chart, compared with 49 for $0.38 from Opus 5 at `low`. On the SWE-bench Pro subset, Claude Opus 5 at its default remains the cheaper way to the top score, as noted earlier. At the other end, Claude Haiku 4.5 answered GPQA Diamond[9](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) questions at about a tenth of Opus 5's cost per question, with 63% accuracy compared with 92% for Opus, and fell much further behind on long coding tasks. It fits high-volume work with checkable outputs, not long agentic loops.
258For most agent workloads, start with Claude Opus 5.5 at its default effort (`medium`), and use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5.5 at higher effort still fall short. On the SWE-bench Pro subset, Opus 5.5 at its default matched Fable 5.1 at its default for about a fifth of the cost per solved task, as noted earlier. On the coding benchmark in [Advisor strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#advisor-strategy-escalate-hard-decisions), it scored 86.6% against 84.2% for Fable 5.1 at `medium` (a single Fable 5.1 run), for under a third of the cost per attempt ($0.84 against $2.68). On Chartography[13](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), a chart-reading benchmark, Opus 5.5 at `low` scored 68.7 for about $0.03 a chart, against 62.5 for $0.15 from Fable 5.1 at `low` and 49 for $0.16 from Claude Opus 5 at `low`. At the other end, Claude Haiku 4.5 answered GPQA Diamond[9](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) questions at about a fifth of Claude Opus 5.5's cost per question, with 63% accuracy compared with 92% for Opus 5.5, and fell much further behind on long coding tasks. It fits high-volume work with checkable outputs, not long agentic loops.
257259
258The ranking flips by workload, and no price list tells you which way. Price every candidate in cost per completed task on your own traffic, including Claude Opus 5 and the frontier model at reduced effort.
260The ranking flips by workload, and no price list tells you which way. Price every candidate in cost per completed task on your own traffic, including Claude Opus 5.5 at its default effort and the frontier model at reduced effort.
259261
260262Price the tail of your workload, not the median: compare models on the hardest tenth of your tasks, not the typical one. On the typical task every model looks similar and the cheapest looks best, but the bill is decided by the tasks the cheaper model fails, because a failed task still bills its tokens, then the retry, then whatever the failure costs downstream. The tail is also where the money goes even when nothing fails. On a 20-problem WideSearch[1](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) run, two problems carried 43% of the spend:
261263
from line 285
283285
284286Lower effort settings are often faster, which matters when latency is the constraint. In these runs, `low` took 4.5 minutes per problem on DeepWideSearch, compared with 7.9 minutes at the default. On the [corpus benchmark](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#orchestrator-strategy-delegate-bulk-work), whose input does not fit in any single context window, Fable 5.1 took 15.2, 17.5, and 19.9 hours per episode at `low`, `medium`, and `high`.
285287
286Long-horizon coding is where effort genuinely buys accuracy. On SWE-bench Pro[3](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), Claude Opus 5 gave up about 2 points at `medium` for half the cost and about 8 points at `low` for a quarter of it: a real tradeoff, which [re-running failures at higher effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#re-run-failures-at-higher-effort) turns back into a saving. This chart plots accuracy against cost for the research and knowledge-work benchmarks and for SWE-bench Pro:
288Long-horizon coding is where effort genuinely buys accuracy. On SWE-bench Pro[3](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), measured against `high`, Claude Opus 5.5 scored about 2.5 points lower at its default, `medium`, for about 70% of the cost, and about 8 points lower at `low` for about a third of the cost; `xhigh` scored about 1.4 points higher for 2.5 times the cost of `high`: a real tradeoff, which [re-running failures at higher effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#re-run-failures-at-higher-effort) turns back into a saving. This chart plots accuracy against cost for the research and knowledge-work benchmarks and for SWE-bench Pro:
287289
288
290
289291
290292Two consequences follow. First, draw this curve for your own workload before you add a second model: in these internal measurements, a multi-model configuration that looked cheaper than the default single model cost more than that same model at lower effort. Second, this curve is the single-model baseline any multi-model strategy must beat, so [step 2 of measuring on your own workload](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#measure-on-your-own-workload) baselines across effort levels.
291293
from line 301
299301
300302When a task's outcome is checkable, the cheapest policy on the effort curve is not a fixed setting: run every task at a low setting and re-run only the failures at a higher one.
301303
302Anthropic computed this policy task by task from the effort runs on the SWE-bench Pro[3](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) subset in [Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort). With Claude Opus 5 at `low`, 16% of tasks failed; with those re-run at the default, about 93% passed for about $0.45 each, against 91.7% for $0.93 running everything at the default: the same pass rate for half the cost, counting the failed cheap attempts. Starting at `medium` instead solved about 94% for about $0.61. Most of the small lift is the second attempt (re-running the default's own failures at the default scores about the same, for more money), so use this policy for the saving, not the lift:
304Anthropic computed this policy task by task from the effort runs on the SWE-bench Pro[3](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) subset in [Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort). With Claude Opus 5.5 at `low`, 13% of tasks failed; with those re-run at `high`, about 97% passed for about $0.17 each, against 95.3% for $0.29 running everything at `high`: a slightly higher pass rate for a little over half the cost, counting the failed cheap attempts. Starting at `medium` instead solved about 97% for about $0.24. Most of the small lift is the second attempt (re-running the failures of one `high` run at `high` scores about the same, for more money), so use this policy for the saving, not the lift:
303305
304
306
305307
306308Two conditions apply. First, you need a failure signal (here, the benchmark's own tests); a checker that passes bad work lets those failures through. Second, every first-pass failure takes two runs' worth of wall-clock time, so the saving is paid for in latency on the failures.
307309
from line 320
318320Three controls do three different jobs. A task budget saves money, because the model sees it. `max_tokens` is a safety cap: lowering it cut cost per attempt without lowering cost per solved task. On Claude Managed Agents, a session budget is the hard dollar stop behind both. Set all three: a task budget, a high `max_tokens`, and a session cap for the run you never want on a bill, with a [workspace spend limit](https://platform.claude.com/docs/en/api/rate-limits#setting-lower-limits-for-workspaces) as the final backstop.
319321
320322* **Task budgets** are in beta (beta header `task-budgets-2026-03-13`) on the most recent models; check the [support table](https://platform.claude.com/docs/en/build-with-claude/task-budgets#feature-support) for which. Start near your loop's 90th-percentile token usage, then tighten ([Choosing a budget](https://platform.claude.com/docs/en/build-with-claude/task-budgets#choosing-a-budget) shows how to collect that distribution). Budgets below the current 20,000-token floor are rejected, and very tight budgets can produce refusal-like behavior. Set the budget once, on the first request, because a mid-task change [invalidates the cache](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#cache-repeated-context). The budget is advisory, steering the model rather than stopping it, so verify adherence on your workload.
321* **`max_tokens`** caps a single response, invisibly to the model, so lowering it does not make the model economize. The turns that needed the room are discarded and still billed. On an internal repository-task benchmark[12](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), a 16,384-token cap ended 15% of Claude Opus 5's attempts and 43% of Claude Fable 5.1's at the default effort, and only 9 of the 117 capped Fable attempts still passed. Capped runs spent less per attempt but bought proportionally fewer solves, so cost per solved task was about the same as at 64,000 ($21 against $22). At 64,000, 2 of about 14,000 turns at the default effort were still cut off, and Fable 5.1 solved 58.5% of tasks instead of 36.3% (on a separate cut of the SWE-bench Pro[3](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) subset, described in reference 12, no difference: 94 of 100 at either cap). Retrying capped attempts rarely helps: at the same cap most of them fail again, and at a higher one you also pay for the wasted attempt. Set `max_tokens` to 64,000 for agentic work, or to 128,000, the maximum, when a single cut-off attempt is costly; at 128,000 Fable 5.1 solved 60.0% for the same cost per solved task. [Stream responses](https://platform.claude.com/docs/en/build-with-claude/streaming) that large, treat [`stop_reason: max_tokens`](https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons#max-tokens) as a failure, and save money with effort and task budgets, which the model can see.
323* **`max_tokens`** caps a single response, invisibly to the model, so lowering it does not make the model economize. The turns that needed the room are discarded and still billed. On an internal repository-task benchmark[12](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), a 16,384-token cap ended about a quarter of Claude Opus 5.5's attempts and 43% of Claude Fable 5.1's, each at its default effort. Only 1 of the 66 capped Opus 5.5 attempts and 9 of the 117 capped Fable attempts still passed. Capped runs spent less per attempt but bought proportionally fewer solves, so cost per solved task was about the same as at 64,000 (on Fable 5.1, $21 against $22; on Opus 5.5, within 1%). At 64,000, 2 of about 14,000 Claude Fable 5.1 turns at its default effort were still cut off (no Claude Opus 5.5 turn was), and Fable 5.1 solved 58.5% of tasks instead of 36.3% (on a separate cut of the SWE-bench Pro[3](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) subset, described in reference 12, no difference: 94 of 100 at either cap). Retrying capped attempts rarely helps: at the same cap most of them fail again, and at a higher one you also pay for the wasted attempt. Set `max_tokens` to 64,000 for agentic work, or to 128,000, the maximum, when a single cut-off attempt is costly; at 128,000 Fable 5.1 solved 60.0% for the same cost per solved task. [Stream responses](https://platform.claude.com/docs/en/build-with-claude/streaming) that large, treat [`stop_reason: max_tokens`](https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons#max-tokens) as a failure, and save money with effort and task budgets, which the model can see.
322324* **Session budgets on Claude Managed Agents** are the hard stop. A [session budget](https://platform.claude.com/docs/en/managed-agents/budgets) is a dollar cap on one session at list rates for tokens, searches, and session time. At the cap, the session pauses with `stop_reason: budget_reached`; raising the budget resumes it. It is platform-enforced, works on any model with a list price, including models where task budgets are not yet available, and combines with the advisory task budget. Deployments apply the same field to every run.
323325
324326Ask for shorter answers. Output tokens cost five times input tokens on Claude Sonnet 5, and in an agent loop every token the model writes comes back as input on every later turn, so you pay for a long answer again and again. Anthropic ran the triage job under three final-answer instructions, three runs each, with the same model and tools. The original asked for two lines:
from line 356
354356
355357At the lower `max_tokens` cap both models spend less per attempt but solve proportionally fewer tasks, so cost per solved task barely moves:
356358
357
359
358360
359361Almost every turn finishes far below either cap. The rare long turn is what the higher cap buys:
360362
361
363
362364
363365## Combine models
364366
from line 389
387389
388390
389391
390The consult rate responds to prompting. With only the tool's built-in description, executors under-call, especially on coding work, so the [advisor tool documentation](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool#prompting-for-coding-and-agent-tasks) gives a system prompt that asks for one call before substantive work and one before finishing, about two to three calls per task. The coding pairing measured next ran at that cadence, about two consultations on every task. That page also covers nudging an under-calling executor and capping calls client-side to bound cost. So watch the consult rate: prompt for it, measure it, and restore the executor's effort if it collapses.
392The consult rate responds to prompting. With only the tool's built-in description, executors under-call, especially on coding work, so the [advisor tool documentation](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool#prompting-for-coding-and-agent-tasks) gives a system prompt that asks for one call before substantive work and one before finishing, about two to three calls per task. The coding pairing measured next ran at that cadence with Claude Opus 5 as the executor, about two consultations on every task; with Claude Opus 5.5 as the executor it asked for advice about 1.4 times per attempt, and 4% of its attempts received no advice. That page also covers nudging an under-calling executor and capping calls client-side to bound cost. So watch the consult rate: prompt for it, measure it, and restore the executor's effort if it collapses.
391393
392**When it pays on cost.** An advisor saves money when a few short consultations, billed at the advisor's rate, replace running the advisor's model for the whole task. That works best when the advisor's model is priced well above the executor's, so the most cost-effective configuration is a frontier advisor over a mid-tier executor. A pairing can hold its own even at the top of the range, because advice also saves executor tokens: an executor told the right approach explores fewer dead ends, which can cover the consultations.
394**When it pays on cost.** An advisor saves money when a few short consultations, billed at the advisor's rate, replace running the advisor's model for the whole task. That works best when the advisor's model is priced well above the executor's, so the most cost-effective configuration is a frontier advisor over a mid-tier executor. A pairing at the top of the range can recover part of the advice's cost, because advice also saves executor tokens: an executor told the right approach explores fewer dead ends. In the coding pairing below with a Claude Opus 5 executor, that saving paid for about half of the advice: the executor spent $1.26 less per attempt than Opus 5 alone at its default, and the consultations cost $2.47. With a Claude Opus 5.5 executor, the advice saved almost no executor cost: $1.36 per attempt against $1.38 for Opus 5.5 alone at `high`, while the consultations cost $1.55.
393395
394On an internal agentic-coding benchmark[11](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), run with a plain API agent, a Claude Opus 5 executor with a Claude Fable 5.1 advisor was the most accurate configuration measured, at $7.69 per attempt. It sits above the line through each model's own effort settings: 3.5 points over Opus 5 alone at the default setting for slightly less money, a gap that five attempts per task do separate from noise, and about 2.5 points over the advisor's model alone for about half again the money:
396On an internal agentic-coding benchmark[11](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), run with a plain API agent, a Claude Opus 5.5 executor at `high` with a Claude Fable 5.1 advisor scored 90.1% at $2.92 per attempt. That is 1.7 points over Opus 5.5 alone at `high`, the executor's own setting, a gap at the edge of run-to-run noise with five attempts per task, for about 2.1 times the money; against Opus 5.5 at its default, `medium`, it is 3.5 points for about 3.5 times the money. It lands about on Opus 5.5's own effort curve, so the advisor buys about what more effort does: Opus 5.5 alone at `xhigh` scored 91.1% for $4.11 per attempt (one attempt per task). In August, a Claude Fable 5.1 advisor over a Claude Opus 5 executor was the most accurate configuration measured, at $6.21 per attempt, a little over twice what the Opus 5.5 pairing costs. The chart plots the Opus 5.5 pairing against Opus 5.5's own effort curve and Claude Fable 5.1's from August:
395397
396
398
397399
398An earlier measurement through [Claude Code's advisor mode](https://code.claude.com/docs/en/advisor) produced the same ordering. Read this result as a shape to test on your workload: the advisor buys a few points at about the executor's own price. A wider capability gap does not guarantee a better deal. The latency cost is the consultations themselves: about two extra frontier-model calls per task on this benchmark, each on the task's critical path.
400An earlier measurement through [Claude Code's advisor mode](https://code.claude.com/docs/en/advisor) also ranked its advisor pairing above both of its models alone. Read the Claude Opus 5.5 result as a shape to test on your workload: the advisor buys a few points for about twice what the executor costs alone. A wider capability gap does not guarantee a better deal. The latency cost is the consultations themselves: about one or two extra frontier-model calls per task on this benchmark, each on the task's critical path.
399401
400**When the stronger model alone is the better step.** Where a workload's accuracy responds to effort, compare the pairing with the advisor's model alone at a reduced setting before building it: the advisor is paid for only on tasks that need it, but a consult that fires on most tasks costs more than running the stronger model itself. On Chartography[13](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) the same pairing matched Claude Fable 5.1 alone at `medium` within run-to-run noise (65.0 against 67.5) at about 2.6 times the cost per task, because the advisor was consulted on nearly every task. Measure your own consult rate first: if the executor asks on most of its tasks, you are paying advisor rates across the whole workload, and running the advisor's model itself is the cheaper way to the same score.
402**When the stronger model alone is the better step.** Where a workload's accuracy responds to effort, compare the pairing with the advisor's model alone at a reduced setting before building it: the advisor is paid for only on tasks that need it, but a consult that fires on most tasks costs more than running the stronger model itself. On Chartography[13](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), a Claude Opus 5.5 executor at `low`, given a Claude Fable 5.1 advisor, consulted it on 1 of 300 tasks and scored 61.7, 7 points below Opus 5.5 alone and beyond run-to-run noise, at about the same cost; in August, a Claude Opus 5 executor consulted on nearly every task, and that pairing matched Fable 5.1 alone at `medium` within run-to-run noise (65.0 against 67.5) at about 1.8 times the cost per task. Measure your own consult rate first: if the executor asks on most of its tasks, you are paying advisor rates across the whole workload, and running the advisor's model itself is the cheaper way to the same score.
401403
402404Whatever the pairing, first price the advisor's model alone at low effort; that is the baseline to beat. Recheck at every model release, because releases move both the capability gap and the price ratio.
403405
from line 457
4554573. If the curve shows a gap effort can't close, add the multi-model strategy that fits and re-run the suite.
4564584. Run the winner in shadow on a traffic slice before cutover, then keep the suite running.
457459
458The following example computes one request's step 1 cost at Claude Opus 5's list prices:
460The following example computes one request's step 1 cost at Claude Opus 5.5's list prices:
459461
460462<CodeGroup>
461463 ```bash cURL
462464 # Per-million-token prices from the pricing page; change these three for another model.
463 INPUT_PER_MTOK=5.00 # Claude Opus 5
464 CACHE_READ_PER_MTOK=0.50 # 0.1x the input price on Claude Opus 5; some models use a different multiplier
465 OUTPUT_PER_MTOK=25.00
465 INPUT_PER_MTOK=4.00 # Claude Opus 5.5
466 CACHE_READ_PER_MTOK=0.20 # 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
467 OUTPUT_PER_MTOK=20.00
466468
467469 response=$(curl --fail-with-body -sS https://api.anthropic.com/v1/messages \
468470 -H "x-api-key: $ANTHROPIC_API_KEY" \
from line 471
469471 -H "anthropic-version: 2023-06-01" \
470472 -H "content-type: application/json" \
471473 -d '{
472 "model": "claude-opus-5",
474 "model": "claude-opus-5-5",
473475 "max_tokens": 1024,
474476 "messages": [{"role": "user", "content": "Hello, Claude"}]
475477 }')
from line 489
487489
488490 ```bash CLI
489491 # Per-million-token prices from the pricing page; change these three for another model.
490 INPUT_PER_MTOK=5.00 # Claude Opus 5
491 CACHE_READ_PER_MTOK=0.50 # 0.1x the input price on Claude Opus 5; some models use a different multiplier
492 OUTPUT_PER_MTOK=25.00
492 INPUT_PER_MTOK=4.00 # Claude Opus 5.5
493 CACHE_READ_PER_MTOK=0.20 # 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
494 OUTPUT_PER_MTOK=20.00
493495
494496 USAGE=$(ant messages create \
495 --model claude-opus-5 \
497 --model claude-opus-5-5 \
496498 --max-tokens 1024 \
497499 --message '{role: user, content: "Hello, Claude"}' \
498500 --transform usage)
from line 511
509511
510512 ```python Python
511513 # Per-million-token prices from the pricing page; change these three for another model.
512 INPUT_PER_MTOK = 5.00 # Claude Opus 5
513 # 0.1x the input price on Claude Opus 5; some models use a different multiplier
514 CACHE_READ_PER_MTOK = 0.50
515 OUTPUT_PER_MTOK = 25.00
514 INPUT_PER_MTOK = 4.00 # Claude Opus 5.5
515 # 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
516 CACHE_READ_PER_MTOK = 0.20
517 OUTPUT_PER_MTOK = 20.00
516518
517519 client = anthropic.Anthropic()
518520 response = client.messages.create(
519 model="claude-opus-5",
521 model="claude-opus-5-5",
520522 max_tokens=1024,
521523 messages=[{"role": "user", "content": "Hello, Claude"}],
522524 )
from line 539
537539
538540 ```typescript TypeScript
539541 // Per-million-token prices from the pricing page; change these three for another model.
540 const INPUT_PER_MTOK = 5.0; // Claude Opus 5
541 const CACHE_READ_PER_MTOK = 0.5; // 0.1x the input price on Claude Opus 5; some models use a different multiplier
542 const OUTPUT_PER_MTOK = 25.0;
542 const INPUT_PER_MTOK = 4.0; // Claude Opus 5.5
543 const CACHE_READ_PER_MTOK = 0.2; // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
544 const OUTPUT_PER_MTOK = 20.0;
543545
544546 const client = new Anthropic();
545547 const response = await client.messages.create({
546 model: "claude-opus-5",
548 model: "claude-opus-5-5",
547549 max_tokens: 1024,
548550 messages: [{ role: "user", content: "Hello, Claude" }]
549551 });
from line 562
560562
561563 ```csharp C#
562564 // Per-million-token prices from the pricing page; change these three for another model.
563 const double InputPerMtok = 5.00; // Claude Opus 5
564 const double CacheReadPerMtok = 0.50; // 0.1x the input price on Claude Opus 5; some models use a different multiplier
565 const double OutputPerMtok = 25.00;
565 const double InputPerMtok = 4.00; // Claude Opus 5.5
566 const double CacheReadPerMtok = 0.20; // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
567 const double OutputPerMtok = 20.00;
566568
567569 AnthropicClient client = new();
568570 var response = await client.Messages.Create(
569571 new MessageCreateParams
570572 {
571 Model = Model.ClaudeOpus5,
573 Model = Model.ClaudeOpus5_5,
572574 MaxTokens = 1024,
573575 Messages = [new() { Role = Role.User, Content = "Hello, Claude" }],
574576 }
from line 590
588590 ```go Go
589591 // Per-million-token prices from the pricing page; change these three for another model.
590592 const (
591 inputPerMTok = 5.00 // Claude Opus 5
592 cacheReadPerMTok = 0.50 // 0.1x the input price on Claude Opus 5; some models use a different multiplier
593 outputPerMTok = 25.00
593 inputPerMTok = 4.00 // Claude Opus 5.5
594 cacheReadPerMTok = 0.20 // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
595 outputPerMTok = 20.00
594596 )
595597
596598 // ...
from line 599
597599 client := anthropic.NewClient()
598600
599601 response, err := client.Messages.New(context.TODO(), anthropic.MessageNewParams{
600 Model: anthropic.ModelClaudeOpus5,
602 Model: anthropic.ModelClaudeOpus5_5,
601603 MaxTokens: 1024,
602604 Messages: []anthropic.MessageParam{
603605 anthropic.NewUserMessage(anthropic.NewTextBlock("Hello, Claude")),
from line 620
618620
619621 ```java Java
620622 // Per-million-token prices from the pricing page; change these three for another model.
621 static final double INPUT_PER_MTOK = 5.00; // Claude Opus 5
622 static final double CACHE_READ_PER_MTOK = 0.50; // 0.1x the input price on Claude Opus 5; some models use a different multiplier
623 static final double OUTPUT_PER_MTOK = 25.00;
623 static final double INPUT_PER_MTOK = 4.00; // Claude Opus 5.5
624 static final double CACHE_READ_PER_MTOK = 0.20; // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
625 static final double OUTPUT_PER_MTOK = 20.00;
624626
625627 void main() {
626628 AnthropicClient client = AnthropicOkHttpClient.fromEnv();
627629
628630 Message response = client.messages().create(MessageCreateParams.builder()
629 .model(Model.CLAUDE_OPUS_5)
631 .model(Model.CLAUDE_OPUS_5_5)
630632 .maxTokens(1024)
631633 .addUserMessage("Hello, Claude")
632634 .build());
from line 647
645647
646648 ```php PHP
647649 // Per-million-token prices from the pricing page; change these three for another model.
648 const INPUT_PER_MTOK = 5.00; // Claude Opus 5
649 const CACHE_READ_PER_MTOK = 0.50; // 0.1x the input price on Claude Opus 5; some models use a different multiplier
650 const OUTPUT_PER_MTOK = 25.00;
650 const INPUT_PER_MTOK = 4.00; // Claude Opus 5.5
651 const CACHE_READ_PER_MTOK = 0.20; // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
652 const OUTPUT_PER_MTOK = 20.00;
651653
652654 $client = new Client();
653655 $response = $client->messages->create(
654 model: 'claude-opus-5',
656 model: 'claude-opus-5-5',
655657 maxTokens: 1024,
656658 messages: [['role' => 'user', 'content' => 'Hello, Claude']],
657659 );
from line 670
668670
669671 ```ruby Ruby
670672 # Per-million-token prices from the pricing page; change these three for another model.
671 INPUT_PER_MTOK = 5.00 # Claude Opus 5
672 CACHE_READ_PER_MTOK = 0.50 # 0.1x the input price on Claude Opus 5; some models use a different multiplier
673 OUTPUT_PER_MTOK = 25.00
673 INPUT_PER_MTOK = 4.00 # Claude Opus 5.5
674 CACHE_READ_PER_MTOK = 0.20 # 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
675 OUTPUT_PER_MTOK = 20.00
674676
675677 client = Anthropic::Client.new
676678 response = client.messages.create(
677 model: "claude-opus-5",
679 model: "claude-opus-5-5",
678680 max_tokens: 1024,
679681 messages: [{ role: "user", content: "Hello, Claude" }]
680682 )
from line 692
690692 ```
691693</CodeGroup>
692694
693In agent loops the cache-read term is usually the largest of the five; if not, check that caching is engaged. When the [advisor tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool#usage-and-billing) or [compaction](https://platform.claude.com/docs/en/build-with-claude/compaction-threshold#understanding-usage) is enabled, some tokens are reported only in `usage.iterations` and not in the top-level totals, so sum over `usage.iterations` instead, pricing `advisor_message` entries at the advisor model's rates.
695In agent loops most input tokens should be cache reads; if `cache_read_input_tokens` is small next to `input_tokens` plus `cache_creation_input_tokens`, check that caching is engaged and that the prefix stays the same between requests. When the [advisor tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool#usage-and-billing) or [compaction](https://platform.claude.com/docs/en/build-with-claude/compaction-threshold#understanding-usage) is enabled, some tokens are reported only in `usage.iterations` and not in the top-level totals, so sum over `usage.iterations` instead, pricing `advisor_message` entries at the advisor model's rates.
694696
695697The following table lists the levers in the order to try them:
696698
697| Lever | Saving in these runs | Quality cost | Latency | Where |
698| ------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
699| Prompt caching | Cost cut by a factor of 2.7 to 5.3 on agent loops; 83% on the triage run | None | Faster | [Cache repeated context](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#cache-repeated-context) |
700| 1-hour cache duration | Cheaper than the 5-minute default once about 1 turn in 20 follows a pause between 5 minutes and an hour and few gaps run over an hour, except on Claude Fable 5.1, where keeping the 5-minute cache warm is cheaper while pauses run minutes and the 1-hour duration wins when pauses run toward an hour; with no pauses the default cost 15% less on Claude Sonnet 5 and 11% less on Claude Opus 5 | None | Stays warm after a pause | [Pick the cache duration](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#pick-the-cache-duration) |
701| Input trimming | A further 5 percentage points on the triage run | None | Neutral | [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
702| Prune stale tool results at task boundaries | 39% on the long triage run (compaction 32%); nothing on short loops | None measured | Neutral | [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
703| Tool search | 45% with 500 tool definitions attached; 20% with a GitHub MCP server | None | Neutral | [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
704| Data files through code execution | 92% on a 25-question data task | A gain, 25 of 25 instead of 6 of 25 | Faster | [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
705| Batch API | 50% | None | Results within 24 hours | [Batch work that can wait](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#batch-work-that-can-wait) |
706| Prompt audit against the current model | 14% on both migrations measured | None; a gain on one | Faster (fewer tool rounds) | [Audit prompts against the current model](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#audit-prompts-against-the-current-model) |
707| Upgrade the model | Opus 4.8 to Opus 5: 12 more points at 21% more per solved task (Opus 5 at `low` beats Opus 4.8 for about 30% of the cost); Sonnet 4.6 to Sonnet 5: 15% less per solved task, 5 more points; Fable 5 to Fable 5.1: 43% less per solved task at about the same score | A gain | Neutral | [Upgrade the model](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#upgrade-the-model) |
708| Lower effort | Knowledge work: `medium` 13% to 31%, `low` a third to a half; long coding: `medium` about half, `low` about three quarters | 1 to 3 points on knowledge work, 2 to 8 on long coding | Faster | [Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort) |
709| Re-run failures | About half, at the same pass rate | None | Two runs on the tasks that fail | [Re-run failures at higher effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#re-run-failures-at-higher-effort) |
710| Task budget | 44% to 58% | 3 to 6 points | Faster | [Set budgets and output caps](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
711| Ask for shorter answers | 39% of output tokens, 14% of cost on the triage run | None | Faster | [Set budgets and output caps](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
712| Raising `max_tokens` | None per solved task, but more tasks solved | Gains of up to 22 points on the internal set; none on the public pair | Neutral | [Set budgets and output caps](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
713| Advisor | Depends on the capability gap and the consult rate; the coding pairing scored 3.5 points over Opus 5 alone and about 2.5 over Fable 5.1 alone, the chart-reading pairing matched the advisor's model alone at `medium` for about 2.6 times the price | Small gains | About two extra calls per task | [Advisor strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#advisor-strategy-escalate-hard-decisions) |
714| Orchestrator | About half against the frontier model, both beyond one context window and on routine tails (the latter measured on Claude Fable 5) | 10 to 12 points below the frontier model | Much faster on large inputs | [Orchestrator strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#orchestrator-strategy-delegate-bulk-work) |
699| Lever | Saving in these runs | Quality cost | Latency | Where |
700| ------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
701| Prompt caching | Cost cut by a factor of 2.7 to 5.3 on agent loops; 83% on the triage run | None | Faster | [Cache repeated context](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#cache-repeated-context) |
702| 1-hour cache duration | Cheaper than the 5-minute default once about 1 turn in 20 follows a pause between 5 minutes and an hour and few gaps run over an hour, except on Claude Fable 5.1, where keeping the 5-minute cache warm is cheaper while pauses run minutes and the 1-hour duration wins when pauses run toward an hour, and on Claude Opus 5.5, where keeping the 5-minute cache warm is cheaper when only a turn or two in 20 follow a pause of up to about half an hour; with no pauses the default cost 15% less on Claude Sonnet 5 and about 15% to 18% less on Claude Opus 5.5 | None | Stays warm after a pause | [Pick the cache duration](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#pick-the-cache-duration) |
703| Input trimming | A further 5 percentage points on the triage run | None | Neutral | [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
704| Prune stale tool results at task boundaries | 39% on the long triage run (compaction 32%); nothing on short loops | None measured | Neutral | [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
705| Tool search | 45% with 500 tool definitions attached; 20% with a GitHub MCP server | None | Neutral | [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
706| Data files through code execution | 92% on a 25-question data task | A gain, 25 of 25 instead of 6 of 25 | Faster | [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens) |
707| Batch API | 50% | None | Results within 24 hours | [Batch work that can wait](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#batch-work-that-can-wait) |
708| Prompt audit against the current model | 14% on both migrations measured | None; a gain on one | Faster (fewer tool rounds) | [Audit prompts against the current model](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#audit-prompts-against-the-current-model) |
709| Upgrade the model | Opus 4.8 to Opus 5: 12 more points at 21% more per solved task (Opus 5 at `low` beats Opus 4.8 for about 30% of the cost); Sonnet 4.6 to Sonnet 5: 15% less per solved task, 5 more points; Fable 5 to Fable 5.1: 43% less per solved task at about the same score | A gain | Neutral | [Upgrade the model](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#upgrade-the-model) |
710| Lower effort | Knowledge work: `medium` 13% to 31%, `low` a third to a half; long coding: `medium` about 30% and `low` about two thirds, both against `high` | 1 to 3 points on knowledge work, 2 to 8 on long coding | Faster | [Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort) |
711| Re-run failures | About 40% against running everything at `high`, at the same pass rate or slightly better | None | Two runs on the tasks that fail | [Re-run failures at higher effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#re-run-failures-at-higher-effort) |
712| Task budget | 44% to 58% | 3 to 6 points | Faster | [Set budgets and output caps](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
713| Ask for shorter answers | 39% of output tokens, 14% of cost on the triage run | None | Faster | [Set budgets and output caps](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
714| Raising `max_tokens` | None per solved task, but more tasks solved | Gains of up to 22 points on the internal set; none on the public pair | Neutral | [Set budgets and output caps](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
715| Advisor | Depends on the capability gap and the consult rate; the coding pairing scored 1.7 points over Claude Opus 5.5 alone at `high` for about 2.1 times the price, about what more effort buys; with Claude Opus 5.5, the chart-reading pairing almost never consulted the advisor and scored 7 points below Opus 5.5 alone | A gain on coding, a loss on chart reading | About one or two extra calls per task | [Advisor strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#advisor-strategy-escalate-hard-decisions) |
716| Orchestrator | About half against the frontier model, both beyond one context window and on routine tails (the latter measured on Claude Fable 5) | 10 to 12 points below the frontier model | Much faster on large inputs | [Orchestrator strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#orchestrator-strategy-delegate-bulk-work) |
715717
716718## Benchmarks referenced
717719
from line 721
719721
7207221. **WideSearch:** Wong et al., "WideSearch: Benchmarking Agentic Broad Info-Seeking," arXiv:2508.07999, 2025. Broad web-research tasks graded on a many-row table's completeness and accuracy; 200 problems, 3 runs per configuration, run August 1 to 2, 2026. The cost-concentration chart is a separate 20-problem run, 3 runs per problem, run August 3 to 4, 2026, costed from per-request billing records.
7217232. **GDPval:** OpenAI, "GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks," 2025. Knowledge-work deliverables graded against task rubrics; a 210-task run of the released gold set, one attempt per task, run August 2, 2026. A Claude model grades, so absolute scores may differ from published results.
7223. **SWE-bench Pro:** Scale AI, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?", 2025. A 482-problem subset selected for compatibility with Anthropic's evaluation harness; scores are not comparable to the public leaderboard. Claude Opus 5 at the default effort averages two runs; reduced-effort settings are single runs; all ran August 4, 2026. Escalation figures come task by task from those runs: `low` first, then the default on its failures, solved 92.5% to 93.6% across run pairings for about $0.45; `medium` first, 93.8% to 94.2% for about $0.61; the default re-run on its own failures, 94.0% for $1.06; everything at the default, 90.9% to 92.5% for $0.93. Costs on this subset are priced as a customer's organization is metered: each request's prior prompt as a cache read and its new tokens as a 5-minute cache write, from the runs' own usage records, checked against a customer ledger; the evaluation organization's own metering, which bills cache in 8,192-token pages, gave figures 1.4 to 1.8 times higher. The Claude Sonnet 5 executor pairings on the advisor chart come from the same measurement series on this subset: the Sonnet-plus-Opus pairing was run twice (August 7 and August 8, 2026, a run and an exact replication), the low-effort pairing once (August 8, 2026), and Claude Sonnet 5 alone twice (77.4%, the baseline for both Pro rows). The Claude Fable 5 point in Upgrade the model is the mean of three runs at the default effort, run August 26, 2026, priced the same way. The Claude Fable 5.1 task-budget figures are one run per budget (two at 35,000 tokens) on the same subset at the default effort, run August 26, 2026, with an unbudgeted run the same day (92.1%, $1.10 per task) as the baseline; an earlier set at `low` effort, run August 21, 2026, scored 88.6% unbudgeted at $0.48 per task. The [Compare models](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#compare-models-on-cost-per-task) comparison pairs that single run with the two pooled Claude Sonnet 5 runs from the same subset; at Fable 5.1's default effort the pair reads the other way, 41% more per solved task than Sonnet 5. The upgrade ladder is one run per model at its shipped defaults (two each for Opus 5 and Sonnet 5, and the Fable 5 point as described above), the Opus and Sonnet runs the same week in one harness and organization.
7243. **SWE-bench Pro:** Scale AI, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?", 2025. A 482-problem subset selected for compatibility with Anthropic's evaluation harness; scores are not comparable to the public leaderboard. They are also not comparable to the SWE-bench Pro results in the Claude Opus 5.5 system card, which come from runs at `max` effort on a different problem set. Claude Opus 5.5 figures average two runs at `low`, `medium` (its default) and `high`, and use one run at `xhigh`, all run September 19 to 20, 2026, with the same 16,384-token cap per turn as the August Claude Opus 5 runs; the cap cut off 2 `xhigh` attempts and none at other settings. The Opus 5.5 runs used a version of the benchmark whose grading containers can reach only internal package mirrors. That version drops one problem whose test needs a live website, and on three more the reference solution fails there, so Opus 5.5 comparisons, and the Claude Fable 5.1 figures set beside them, use the remaining 478 problems. The Claude Opus 5 SWE-bench Pro figures in [Upgrade the model](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#upgrade-the-model) and the [advisor pairings chart](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#advisor-strategy-escalate-hard-decisions) average two runs at its default effort and use one run at `low`, all run August 4, 2026. Escalation figures come task by task from the Opus 5.5 runs: `low` first, then `high` on its failures, solved 96.4% to 97.5% across run pairings for about $0.17; `medium` first, 96.0% to 97.1% for about $0.24; `high` re-run on its own failures, 96.9% for $0.31; everything at `high`, 94.8% to 95.8% for $0.29. Costs on this subset are priced as a customer's organization is metered: each request's prior prompt as a cache read and its new tokens as a 5-minute cache write, from the runs' own usage records, checked against a customer ledger; the evaluation organization's own metering, which until September 10, 2026, billed cache reads in 8,192-token blocks for Claude Opus 5, Claude Fable 5, Claude Opus 4.7, and Claude Opus 4.8, gave figures 1.4 to 1.8 times higher for those models' runs; for Claude Fable 5.1, Claude Sonnet 5, and Claude Sonnet 4.6 the two differ by at most about 9%, and for the Claude Opus 5.5 figures they agree within 3% at each effort setting. The Claude Sonnet 5 executor pairings on the advisor chart come from the same measurement series on this subset: the Sonnet-plus-Opus 5 pairing was run twice (August 7 and August 8, 2026, a run and an exact replication), the low-effort pairing once (August 8, 2026), and Claude Sonnet 5 alone twice (77.4%, the baseline for both Pro rows). The Claude Fable 5 point in Upgrade the model is the mean of three runs at the default effort, run August 26, 2026, priced the same way. The Claude Fable 5.1 task-budget figures are one run per budget (two at 35,000 tokens) on the same subset at the default effort, run August 26, 2026, with an unbudgeted run the same day (92.1%, $1.10 per task) as the baseline; an earlier set at `low` effort, run August 21, 2026, scored 88.6% unbudgeted at $0.48 per task. The [Compare models](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#compare-models-on-cost-per-task) comparison pairs that single run with the two pooled Claude Sonnet 5 runs from the same subset; at Fable 5.1's default effort the pair reads the other way, 41% more per solved task than Sonnet 5. The upgrade ladder is one run per model at its shipped defaults (two each for Opus 5 and Sonnet 5, and the Fable 5 point as described above), the Opus and Sonnet runs the same week in one harness and organization.
7237254. **BrowseComp:** Wei et al., "BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents," OpenAI, 2025. Effort figures use a 500-problem cut, one to three runs per setting, run August 3, 2026, with the default point pooling two runs from July 26 to 27, 2026. The cost-insurance chart uses 10 reliably solved problems from a 26-problem slice, 50 delegated runs (August 1 to 2, 2026) and 70 solo runs (50 from August 2 to 3, 2026; 20 archived from July 12 to 13 and August 1, 2026), $6.45 compared with $11.99 per run in expectation; delegated figures carry a measurement band of about 20%.
7247265. **Agent-architecture scaling:** Kim et al., "Towards a Science of Scaling Agent Systems," arXiv:2512.08296, 2025. Independent external study, cited only for the direction of the finding on when delegation does not pay, not for any figure.
7257276. **DeepWideSearch:** "DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking," arXiv:2510.20168, 2025. The 220 questions span 15 domains, each combining many-row collection with multi-hop retrieval; measured on the benchmark's standing row set, 3 runs per configuration, run August 2, 2026 (the single-worker team point ran July 26 to 27, 2026).
7267287. **DeepResearch Bench II:** Li et al., "DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Report," arXiv:2601.08536, 2026. Its 132 research tasks across 22 domains are graded against expert-derived binary rubrics; measured on a 50-task subset stratified across all themes, one attempt per task, 3 runs per setting, on [Claude Managed Agents](https://platform.claude.com/docs/en/managed-agents/overview) with the platform's own web search and fetch tools (August 26 to 27, 2026); scored on the 33 tasks no configuration refused, with attempts the production safety classifiers cut short removed; costs are what a customer is billed, the platform's requests plus web-search fees. Scores are each model's mean on the 33-task basis with its own pre-empted tasks removed; on the 21 tasks clean in every arm, Claude Fable 5.1 holds a 2-to-3-point lead over Claude Fable 5 at every effort level and both models are flat across effort. The caching chart re-prices the same requests with every input token at the uncached rate. Claude Opus 4.6 judges under the benchmark's rubric protocol; the original uses a different judge, and an Anthropic judge may favor the house style. Claude Opus 5 at its default effort ran on the same surface and subset, three runs, on August 28, 2026: 68.8% on the raw 50 tasks, 70.8% on the 33-task basis, and 71.1% on the 21-task set, at $6.71 per task ($23.72 without caching); none of its attempts was cut short by the safety classifiers, under a safeguards deployment newer than the one the other models ran under.
7277298. **Corpus defect sweep:** Anthropic-internal, for work larger than one context window: a 21.6-million-token corpus from 14 public Python package sources with 130 planted defects and deterministic grading; protocol fixed before the runs and internally reviewed; three runs per configuration. Every configuration ran on Claude Managed Agents. The charted team configuration is a run in which the Claude Fable 5.1 coordinator ran the whole sweep inside the platform at its documented limit of 25 concurrent Claude Sonnet 5 workers, run August 30, 2026; its three episodes scored F1 0.764, 0.825, and 0.791 after the extras audit (raw 0.751, 0.821, and 0.781) for $225, $234, and $283. The Claude Sonnet 5 solo configuration ran August 3 to 4, 2026; the Claude Fable 5.1 solo configurations ran August 24 to 25, 2026, under the platform's launch serving settings, three seeds per effort setting, on the same corpus build. The sandbox image carried installed copies of part of the corpus, and Claude Fable 5.1's final assembly step compared against them in 7 of 9 episodes; re-grading without those additions moved the affected seeds by up to 3 points. Absolute F1 is specific to this corpus build, not comparable across benchmarks; configuration comparisons are like for like.
7289. **GPQA Diamond:** Rein et al., "GPQA: A Graduate-Level Google-Proof Q\&A Benchmark," 2023. The 198-question Diamond subset, two runs per configuration, run August 7, 2026, model-graded against reference answers, advisor tokens metered per request. A platform safety check refused two biology questions on the Sonnet and Opus executors; excluding them changes no comparison by more than one point.
7309. **GPQA Diamond:** Rein et al., "GPQA: A Graduate-Level Google-Proof Q\&A Benchmark," 2023. The 198-question Diamond subset, two runs per configuration, run August 7, 2026 (Claude Opus 5.5: September 19, 2026), model-graded against reference answers, advisor tokens metered per request. A platform safety check refused two biology questions on the Claude Sonnet 5 executors, and one of them also on Claude Opus 5; excluding them changes no comparison by more than one point. Claude Opus 5.5's 92% comes from two runs that set `fallbacks: "default"` to opt into [server-side fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#server-side-fallback), with any attempt that still ended in a refusal counted as wrong. In each run the safety check flagged six biology questions, Claude Opus 5 answered five of them through the fallback, and the sixth still ended in a refusal. Opus 5.5's cost per question includes those fallback answers. Without counting refusals as wrong, these runs score 93%, because the grader still assigns an answer option to a refused attempt, usually the correct one. With refusals counted as wrong, Claude Opus 5's runs score 91% (one refusal per run), as do two Claude Opus 5.5 runs with fallback off, in which Opus 5.5 refused five or six biology questions per run.
72973110. **DeepSWE:** Datacurve, "DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks," arXiv:2607.07946, 2026. The set has 113 original tasks across five languages with program-based verifiers. Pairings are two runs each, run August 7, 2026, with advisor tokens metered per request, and used a client-side advisor loop rather than the advisor tool, with identical accounting. Single-model effort sweeps are single runs priced from token counts, a cache-aware approximation. Costs per task are run totals divided by 113.
73011. **Internal agentic-coding benchmark:** Anthropic-internal: 370 repository tasks graded by the repositories' own tests. The API figures were measured with a 128,000-token output cap, one run per configuration: Opus 5 alone at the default effort August 9 to 10, 2026, Claude Fable 5.1 alone at five explicitly set effort values August 20, 2026, and the pairing August 24 to 25, 2026. Attempts per task: five for the pairing and the Opus-alone control, one for the single-model points; the pairing averaged about two advisor consultations per attempt; costs are per attempt. The Claude Code figures are runs of the same tasks from July 8 to 23, 2026, one run per configuration, costs approximate.
73112. **Internal repository-task benchmark (cap measurement):** A separate Anthropic-internal set of about 130 repository tasks, run August 8 to 10, 2026 (Claude Opus 5) and August 20, 2026 (Claude Fable 5.1), with a plain API agent loop, one attempt per task. The Claude Fable 5.1 runs are 135 tasks per cap at the default effort set explicitly: the 16,384-token figure averages two runs (36.3% on both); the 64,000 and 128,000 figures are single runs (58.5% and 60.0%). Six problems drew a safety refusal in every run and count as failures. The Opus 5 16,384-token figure averages two runs and its 64,000 figure is a single run (124 tasks scored). The SWE-bench Pro cap figures are one Claude Fable 5.1 run per cap at the default effort, run August 26, 2026, on a 100-problem subset stratified from reference 3's 482-problem set, not comparable to its scores; the two caps scored the same at the default. The chart's per-turn distributions come from the Opus run at 64,000 and the Claude Fable 5.1 run at 128,000; no Opus turn reached its cap, and one Fable 5.1 turn reached 128,000 (0.46% of its turns exceeded 16,384).
73213. **Chartography:** Surge AI, "Chartography," 2026. The complete released 100-question set, measured August 8 to 10, 2026, with Anthropic's implementation on Claude Managed Agents (standard cloud sandbox; advisor configurations use the Managed Agents advisor). Claude Sonnet 4.6 grades instead of the reference judge and the benchmark runs with tools, so scores compare across configurations here but not to the published leaderboard. Two runs per configuration, pooled; run-to-run spreads were 4 to 10 points. Costs exclude sandbox time, which added under 1%. The Claude Fable 5.1 solo runs are from August 24, 2026, under the platform's launch serving settings, two runs per setting; six attempts hit the 15-minute session cap and score 0, and two charts per run were answered by Claude Opus 5 after a safety refusal. The Claude Opus 5 low-effort executor with a Claude Fable 5.1 advisor ran twice on August 30, 2026, under the same settings (63.0 and 67.0, mean 65.0, at $0.72 a chart; the advisor was consulted on 88% of tasks in each run, and 4 of its 219 replies came from Claude Opus 5 instead, each after a production safety filter stopped the advisor's own reply). The consult-rate comparison for the earlier pairings comes from rerunning the same configurations on the Messages API with a container tool set, August 10 to 11, 2026.
73211. **Internal agentic-coding benchmark:** Anthropic-internal: 370 repository tasks graded by the repositories' own tests. The API figures were measured with a 128,000-token output cap, one run per configuration: Opus 5 alone at the default effort August 9 to 10, 2026, and at `low` and `medium` August 10, 2026; Claude Fable 5.1 alone at five explicitly set effort values August 20, 2026 (the chart shows three of them); and the pairing August 24 to 25, 2026. Claude Opus 5.5 alone ran on all 370 tasks, September 19 to 20, 2026: at its default effort (`medium`) and at `high` with five attempts per task, and at `low` and `xhigh` with one (369 of 370 scored at each, after a setup-check failure). The Claude Opus 5.5 executor at `high` with the released Claude Fable 5.1 as advisor (the August runs used a pre-release snapshot) ran five attempts per task on the same dates; one task failed its setup check, so 1,845 attempts were scored. The 279 attempts in which the advisor was turned away under load were re-run, and attempts whose consults timed out were kept, as in August. The August runs had five attempts per task for the pairing and the Claude Opus 5 control and one for the other points. The August pairing averaged about two advisor consultations per attempt; the Claude Opus 5.5 pairing requested 1.39 and received 1.35. Costs are per attempt. Costs are priced as a customer's organization is metered: each agent-loop request's prior prompt as a cache read and its new tokens as a 5-minute cache write, from the runs' own usage records, and each advisor call, which uses no cache, from its recorded tokens, all at list prices. The Claude Code figures are runs of the same tasks from July 8 to 23, 2026, one run per configuration, costs approximate.
73312. **Internal repository-task benchmark (cap measurement):** A separate Anthropic-internal set of about 130 repository tasks, run August 20, 2026 (Claude Fable 5.1) and September 19, 2026 (Claude Opus 5.5, at its default effort, `medium`), with a plain API agent loop, one attempt per task. The Claude Fable 5.1 runs are 135 tasks per cap at the default effort set explicitly: the 16,384-token figure averages two runs (36.3% on both); the 64,000 and 128,000 figures are single runs (58.5% and 60.0%). Six problems drew a safety refusal in every run and count as failures. The Claude Opus 5.5 16,384-token figure averages two runs (134 and 135 tasks scored), and its 64,000 and 128,000 figures are single runs (135 tasks each); two attempts in each 16,384-token run ended in a safety refusal and count as failures. The SWE-bench Pro cap figures are one Claude Fable 5.1 run per cap at the default effort, run August 26, 2026, on a 100-problem subset stratified from reference 3's 482-problem set, not comparable to its scores; the two caps scored the same at the default. The chart's per-turn distributions come from the Claude Opus 5.5 and Claude Fable 5.1 runs at 128,000: no Opus 5.5 turn reached the cap (the longest was about 61,000 tokens, and 0.56% of its turns exceeded 16,384), and one Fable 5.1 turn reached 128,000 (0.46% of its turns exceeded 16,384).
73413. **Chartography:** Surge AI, "Chartography," 2026. The complete released 100-question set, measured August 6 and 9, 2026 (Claude Opus 5 alone) and September 20, 2026 (Claude Opus 5.5), with Anthropic's implementation on Claude Managed Agents (standard cloud sandbox; advisor configurations use the Managed Agents advisor). Claude Sonnet 4.6 grades instead of the reference judge and the benchmark runs with tools, so scores compare across configurations here but not to the published leaderboard. They are also not comparable to the Chartography results in the Claude Opus 5.5 system card, which use a different grader and run at `max` effort. Two runs per configuration (three for Claude Opus 5.5), pooled; run-to-run spreads were up to 10 points. Costs are what a customer running the agent routinely is billed: each chart's first request reads the agent's shared system prompt and tools from the cache, as it does when another session of the same agent ran in the previous 5 minutes. A chart run on its own costs about $0.03 more with Claude Opus 5 or Claude Opus 5.5 and about $0.12 more with Claude Fable 5.1. The August figures are re-priced this way from the runs' usage records; the evaluation organization's own metering, which until September 10, 2026, billed Claude Opus 5's cache reads in 8,192-token blocks, overstated Claude Opus 5's costs. Costs exclude sandbox time, which added under 1% to the August runs. The Claude Fable 5.1 solo runs are from August 24, 2026, under the platform's launch serving settings, two runs per setting; six attempts hit the 15-minute session cap and score 0, and two charts per run were answered by Claude Opus 5 after a safety refusal. The Claude Opus 5 low-effort executor with a Claude Fable 5.1 advisor ran twice on August 30, 2026, under the same settings (63.0 and 67.0, mean 65.0, at $0.47 a chart; the advisor was consulted on 88% of tasks in each run, and 4 of its 219 replies came from Claude Opus 5 instead, each after a production safety filter stopped the advisor's own reply). Claude Opus 5.5 ran at `low`, with server-side fallback off and a safety classifier judging every tool call: three runs alone (70, 68, and 68) and three with a Claude Fable 5.1 advisor configured (59, 63, and 63), in which it consulted the advisor on 1 of 300 tasks. The consult-rate comparison for the earlier pairings comes from rerunning the same configurations on the Messages API with a container tool set, August 10 to 11, 2026.
73373514. **Support-desk prompt-audit evaluation:** An Anthropic-constructed set of 44 support tickets with deterministic grading, run in early August 2026 and reported on August 8, 2026, under six system prompts, each adding to the same clean prompt one pattern common in prompts written for Claude Opus 4.8 and Claude Sonnet 4.6. Each chart point is one of three cases (older model, newer model on the same prompt, newer model after the audit) averaged over the six prompts and 44 tickets. The Opus 5 accuracy gain has a 95% confidence interval of 3 to 8 points; the Sonnet accuracy differences are within noise.
73473615. **Data-file question set:** An Anthropic-constructed set of 25 aggregate questions over a 1,862-row slice of a public liquor-sales CSV, with ground truth computed by pandas and exact-match grading, run on Claude Sonnet 5 and Claude Opus 5 with thinking disabled (the in-context arm cannot complete at the default), a 4,000-token output cap, and no prompt caching, three runs per configuration, run August 19, 2026. The file arm uploads the CSV through the Files API and uses the `code_execution_20260120` tool.
73516. **Cache duration measurement:** The 20-issue triage job from [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens), run August 23, 2026, on Claude Sonnet 5 and Claude Opus 5 on the Messages API with the same harness, the Claude Opus 5 cells with `max_tokens` raised to 4,096, with pauses inserted before a randomly chosen share of turns (none, 5%, 10%, and every turn at 6 minutes on all 20 issues on both models, plus every turn at 2 minutes on Claude Sonnet 5; 20-minute pauses on a 5-issue subset on both models; 45-minute pauses on a 5-issue subset on Claude Sonnet 5 only). Three runs per cell, cost computed from each response's `usage` fields on a customer-billed organization at list prices, accuracy against the same gold labels. The crossover is about 3.3% of turns on both models: the median of each session's break-even share, computed by the cost model from that session's turn-by-turn context sizes, over all 45 Claude Sonnet 5 and 36 Claude Opus 5 twenty-issue sessions in the analysis (every pause schedule run on the full job, under all three cache settings, three runs each; the 5-issue cells are not in it). The 5% cell tied on Claude Sonnet 5 because that draw's pauses fell on small prefixes. The page's 1-in-20 rule sits above the measured crossover. Anthropic measured keep-alive requests that refresh the 5-minute cache as a comparator only. They matched the 1-hour setting at best and cost more with a pause before every turn, so do not use them on these two models; on Claude Fable 5.1 the arithmetic reverses (reference 19).
73716. **Cache duration measurement:** The 20-issue triage job from [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens), run on Claude Sonnet 5 on August 23, 2026, and on Claude Opus 5.5 on September 19 and 20, 2026, at its default effort (`medium`) and at `high`, on the Messages API with the same harness (for Claude Opus 5.5, a port of it that sends the same request bodies), the Claude Opus 5.5 cells with `max_tokens` raised to 4,096, with pauses inserted before a randomly chosen share of turns (none, 5%, 10%, and every turn at 6 minutes on all 20 issues on both models, plus every turn at 2 minutes on Claude Sonnet 5; 20-minute and 45-minute pauses on a 5-issue subset on both models). Claude Opus 5's keep-alive figures below come from the same job on August 23, 2026, with `max_tokens` raised to 4,096, on the same schedules except the 2-minute and 45-minute pauses. Three runs per cell, cost computed from each response's `usage` fields at list prices (for Claude Opus 5.5, $4 input, $5 5-minute write, $8 1-hour write, $0.20 cache read, and $20 output per million tokens; Claude Sonnet 5 ran on an Anthropic-internal organization whose usage is metered the same way as a customer organization's), accuracy against the same gold labels. The Claude Opus 5.5 figures on this page cover both effort levels. The crossover is about 3.3% of turns on Claude Sonnet 5 and 3.1% to 3.2% on Claude Opus 5.5: the median of each session's break-even share, computed by the cost model from that session's turn-by-turn context sizes, over all 45 Claude Sonnet 5 twenty-issue sessions and the 36 Claude Opus 5.5 twenty-issue sessions at each effort level (every pause schedule run on the full job, under all three cache settings, three runs each; the 5-issue cells are not in it). In the 5% cell the 5-minute and 1-hour settings tied on Claude Sonnet 5, because that draw's pauses fell on small prefixes; on Claude Opus 5.5 they nearly tied. The page's 1-in-20 rule sits above the measured crossover. Claude Opus 5.5's time to first token after a pause was not measured. Anthropic measured keep-alive requests that refresh the 5-minute cache on Claude Sonnet 5 and Claude Opus 5 on August 23, 2026, and on Claude Opus 5.5 in the runs above, always sent with `max_tokens: 1`. On Claude Sonnet 5 they cost 7.7% less than the 1-hour setting with 5% of turns paused and about the same with 10%; on Claude Opus 5 no difference was measurable at either share; on both they cost more with a pause of 6 minutes or more before every turn. On Claude Opus 5.5 they cost 8% to 18% less than the 1-hour setting with 5% and 10% of turns paused (about 10% to 15% once between-session noise is removed by re-billing each keep-alive session's own tokens at 1-hour cache prices), and more with a pause before every turn: 4% to 6% more at 6 minutes, 9% to 10% at 20 minutes, and 56% to 58% at 45 minutes. Keep-alive saved more on Claude Opus 5.5 because each keep-alive request re-reads the prefix at the cache-read price: 0.05x the input price, against 0.1x on Claude Sonnet 5 and Claude Opus 5; Claude Opus 5's sessions, re-billed at Claude Opus 5.5's prices, show nearly the same savings as Claude Opus 5.5. Anthropic's pre-launch API tests on Claude Opus 5.5 show that a `max_tokens: 0` request writes the cache and that the next request reads it; whether such a request refreshes an existing entry was not measured on Opus 5.5. On Claude Fable 5.1, at 0.025x, keep-alive was cheaper even with a pause before every turn, except at 45-minute pauses (reference 19).
73673817. **Cache-read share in production:** Aggregated first-party Claude API usage for the 14 days ending August 23, 2026, direct API product only, Anthropic-internal organizations excluded, no organization identified. An organization-day counts as an agent loop when its requests carry tool definitions and tool results, its prompts hold 9 or more prior tool calls on average, caching was used, and it made at least 10 such requests (the API has no conversation identifier, so this stands in for conversation length): 303,003 organization-days across 106,487 organizations, median cache-read share 84.2% of all input tokens, upper quartile 91.7%. Use-case labels (the organization's declared use case, or otherwise its classified one) cover 74% of those organization-days and 99% of their tokens; coding organizations supply 87% of agentic input tokens and read a median 88.5% (90.9% at 25 or more prior tool calls), upper quartile 93.4%, with about 72% of coding organization-days at 80% or more; support, research, and data agents read 84% to 85%. The top decile of organization-days reads 95.9% or more for coding and 94.2% to 94.8% for support, research, data, and other agents. The request-level split at 25 or more prior tool calls comes from a six-hour sample: coding 92% read, 7% write, under 1% uncached. Unlabeled organizations, mostly small, read a median 11%. Organization-days with no tool definitions read a median 34.6%. An independent query over the same window that reconstructs conversations of 10 or more requests, rather than scoring organization-days, puts the median at 90.2%; the difference is scope, not data.
73773918. **Compaction timing measurement:** The triage agent's long variant from [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens), run August 24, 2026, on Claude Sonnet 5 with the 5-minute cache, cost from the usage fields at list prices, five sessions per arm: a no-change arm at the default effort throughout ($0.81 per session), and two arms that start at low effort and make the same two cache-breaking changes, a switch to the default effort and one added tool, either mid-session at requests 12 and 17 ($0.95) or together on the first request after the first compaction ($0.75). A fourth arm of six sessions, run August 25, 2026, made the same two changes on the request that triggered the first compaction ($0.92 per session): that request's summarization pass wrote the 81,000-token context to the cache instead of reading it, so that pass cost $0.21 against $0.04 for the same pass in the boundary arm. Sessions first compacted at request 21 to 25 (16 of the 21 sessions at request 22), once the prompt passed the 80,000-token compaction trigger, and two no-change sessions compacted a second time near the end. The boundary arm's lower total than the no-change arm reflects its low-effort requests before the change and those second compactions rather than caching: the two arms' re-write costs differ by under a cent. The mid-session arm paid $0.23 per session in cache re-writes; the difference between the mid-session and boundary arms was $0.20 with a 95% confidence interval of $0.11 to $0.29. One mid-session session ran cheap ($0.82) after its model mis-called the search tool following compaction and got empty results; it is included, and without it the arm averages $0.98. Accuracy averaged 14.2 of 20 labels in each August 24 arm and 14.7 in the August 25 arm; cache reads were 91% of prompt tokens with no changes, 85% mid-session, 91% at the boundary, and 86% with the changes on the triggering request.
73819. **Cache duration measurement on Claude Fable 5.1:** The same 20-issue triage job and harness as reference 16, run August 23 and August 26, 2026, on the Claude Fable 5.1 launch snapshot at its launch prices ($10 input, $12.50 5-minute write, $20 1-hour write, $0.25 cache read, $50 output per million tokens), three settings per schedule: the 5-minute cache, the 1-hour cache, and the 5-minute cache kept warm by a `max_tokens: 0` request on the unchanged prefix every 4 minutes of idle time (the August 23 runs pinged with `max_tokens: 1`; every August 26 ping refreshed the cache and billed no output). Schedules: no pauses, 10% of turns, and every turn at 6 minutes on all 20 issues, and 45-minute pauses on the 5-issue subset; three runs per cell, cost computed from each response's `usage` fields at list prices, accuracy against the same gold labels (12 to 17 exact labels of 20). Per-session means on August 26 for the 5-minute, 1-hour, and keep-alive settings: no pauses $2.42, $3.09, $2.29; 10% paused $4.50, $2.96, $2.36; every turn $22.89, $3.01, $2.62; the August 23 cells agree within 6%. The 45-minute figures ($1.68, $0.59, and $0.71 per 5-issue session) are from a clean re-run on August 26 after a cache-billing incident spoiled that day's first cells; the August 23 runs gave $1.67, $0.58, and $0.70. The crossover between the 5-minute and 1-hour settings is 3.1% of turns, the same measure as reference 16.
73920. **Terminal-Bench 3:** the public terminal-agent benchmark's 74 tasks, run on [Claude Managed Agents](https://platform.claude.com/docs/en/managed-agents/overview) with two custom tools, a shell and a file editor that the evaluation harness runs in each task's own container, in place of the platform's built-in tools, and otherwise at the platform's default settings for external accounts, two runs per model at `high` effort, August 27 to 28, 2026. Scores are raw pass rates over the 148 attempts per model; single runs swing by 5 to 11 points. Costs are what a customer would be billed at list prices, re-priced request by request from the runs' usage records with the 5-minute cache lifetime. Claude Opus 4.7 ended 11 of its 148 attempts at its output cap.
74019. **Cache duration measurement on Claude Fable 5.1:** The same 20-issue triage job and harness as reference 16, run August 23 and August 26, 2026, on the Claude Fable 5.1 launch snapshot at its launch prices ($10 input, $12.50 5-minute write, $20 1-hour write, $0.25 cache read, $50 output per million tokens), three settings per schedule: the 5-minute cache, the 1-hour cache, and the 5-minute cache kept warm by a `max_tokens: 0` request on the unchanged prefix every 4 minutes, timed from the previous request's start (the August 23 runs sent keep-alive requests with `max_tokens: 1`; in the August 26 cells reported here, every keep-alive request refreshed the cache and billed no output). Schedules: no pauses, 10% of turns, and every turn at 6 minutes on all 20 issues, and 45-minute pauses on the 5-issue subset; three runs per cell (six for the August 26 keep-alive cell with 45-minute pauses), cost computed from each response's `usage` fields at list prices, accuracy against the same gold labels (12 to 17 exact labels of 20). Per-session means on August 26 for the 5-minute, 1-hour, and keep-alive settings: no pauses $2.42, $3.09, $2.29; 10% paused $4.50, $2.96, $2.36; every turn $22.89, $3.01, $2.62; the August 23 cells agree within 6%. The 45-minute figures ($1.68, $0.59, and $0.71 per 5-issue session) are from a clean re-run on August 26 after a cache-billing incident spoiled that day's first cells; the August 23 runs gave $1.67, $0.58, and $0.70. The crossover between the 5-minute and 1-hour settings is 3.1% of turns, the same measure as reference 16.
74120. **Terminal-Bench 3:** the public terminal-agent benchmark's 74 tasks, run on [Claude Managed Agents](https://platform.claude.com/docs/en/managed-agents/overview) with two custom tools, a shell and a file editor that the evaluation harness runs in each task's own container, in place of the platform's built-in tools, and otherwise at the platform's default settings for external accounts, two runs per model at `high` effort, August 27 to 28, 2026. These runs used Terminal-Bench version 3.0, and their scores are not comparable with the public Terminal-Bench leaderboard or with the Terminal-Bench 4.0 results in the Claude Opus 5.5 system card, which come from runs in Claude Code at `max` effort. Each task's time limits are 2.5 times the benchmark's own, which gives the agent between 75 minutes and 20 hours per task (5 hours for the median task), and each task gets three times the memory it specifies, from 6 GiB to 96 GiB, with extra memory for the 12 tasks that run helper services. The agent had no general internet access: its containers could reach an internal package mirror, a short list of download sites including GitHub and the Python Package Index, and a few sites specific to some tasks, and eight of the tasks had no network access at all. Scores are raw pass rates over the 148 attempts per model; single runs swing by 5 to 11 points. Costs are what a customer would be billed at list prices, re-priced request by request from the runs' usage records with the 5-minute cache lifetime. Claude Opus 4.7 ended 11 of its 148 attempts at its output cap.
740742
741743## Next steps
742744
agents-and-tools/tool-use/tool-runner Changed · +54 / -42 lines
from line 14
1414* Provides type safety and validation
1515
1616<Note>
17 The tool runner is in beta and available in the [Python SDK](https://github.com/anthropics/anthropic-sdk-python/blob/main/tools.md), [TypeScript SDK](https://github.com/anthropics/anthropic-sdk-typescript/blob/main/helpers.md#tool-helpers), [C# SDK](https://github.com/anthropics/anthropic-sdk-csharp/blob/main/examples/ToolRunnerExample/Program.cs), [Go SDK](https://github.com/anthropics/anthropic-sdk-go/blob/main/tools.md), [Java SDK](https://github.com/anthropics/anthropic-sdk-java/blob/main/anthropic-java-example/src/main/java/com/anthropic/example/BetaToolRunnerExample.java), [PHP SDK](https://github.com/anthropics/anthropic-sdk-php/blob/main/examples/beta/beta_tool_runner.php), and [Ruby SDK](https://github.com/anthropics/anthropic-sdk-ruby/blob/main/helpers.md#3-auto-looping-tool-runner-beta).
17 The tool runner is in beta and available in the [Python SDK](https://github.com/anthropics/anthropic-sdk-python/blob/main/tools.md), [TypeScript SDK](https://github.com/anthropics/anthropic-sdk-typescript/blob/main/helpers.md#tool-helpers), [C# SDK](https://github.com/anthropics/anthropic-sdk-csharp/blob/main/examples/ToolRunnerExample/Program.cs), [Go SDK](https://github.com/anthropics/anthropic-sdk-go/blob/main/tools.md), [Java SDK](https://github.com/anthropics/anthropic-sdk-java/blob/main/anthropic-java-example/src/main/java/com/anthropic/example/BetaToolRunnerRunnableToolExample.java), [PHP SDK](https://github.com/anthropics/anthropic-sdk-php/blob/main/examples/beta/beta_tool_runner.php), and [Ruby SDK](https://github.com/anthropics/anthropic-sdk-ruby/blob/main/helpers.md#3-auto-looping-tool-runner-beta).
1818</Note>
1919
2020## Basic usage
from line 350
350350 </Tab>
351351
352352 <Tab title="Java">
353 Define each tool as a class implementing `Supplier<String>`. Annotate the class with `@JsonClassDescription` for the tool description, and each public field with `@JsonPropertyDescription` for parameter descriptions. The SDK derives the JSON schema, tool name (snake-cased class name), and input parsing from the class, and marks the tool with `strict: true` ([strict tool use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/strict-tool-use)).
353 Define each tool as a `BetaRunnableTool` that pairs an input class with a function that runs when Claude calls the tool. Annotate the input class with `@JsonClassDescription` for the tool description, and each public field with `@JsonPropertyDescription` for parameter descriptions. The SDK derives the JSON schema, tool name (snake-cased class name), and input parsing from the class, and marks the tool with `strict: true` ([strict tool use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/strict-tool-use)).
354354
355 The function receives the parsed input as an instance of that class and returns a `BetaToolResultBlockParam.Content`. To return text, wrap it with `BetaToolResultBlockParam.Content.ofString()`. Because the function is a lambda, it can use objects from your application, such as the `WeatherService` in the following example.
356
355357 ```java
356358 import com.anthropic.client.AnthropicClient;
357359 import com.anthropic.client.okhttp.AnthropicOkHttpClient;
360 import com.anthropic.helpers.BetaRunnableTool;
358361 import com.anthropic.helpers.BetaToolRunner;
359362 import com.anthropic.models.beta.messages.BetaMessage;
363 import com.anthropic.models.beta.messages.BetaToolResultBlockParam;
360364 import com.anthropic.models.beta.messages.MessageCreateParams;
361365 import com.anthropic.models.messages.Model;
362366 import com.fasterxml.jackson.annotation.JsonClassDescription;
363367 import com.fasterxml.jackson.annotation.JsonPropertyDescription;
364 import java.util.function.Supplier;
365368
366369 @JsonClassDescription("Get the current weather in a given location")
367 static class GetWeather implements Supplier<String> {
370 static class GetWeather {
368371 @JsonPropertyDescription("The city and state, e.g. San Francisco, CA")
369372 public String location;
370373
371374 @JsonPropertyDescription("Temperature unit, either 'celsius' or 'fahrenheit'")
372375 public String unit;
373
374 @Override
375 public String get() {
376 return "{\"temperature\": \"20°C\", \"condition\": \"Sunny\"}";
377 }
378376 }
379377
380378 @JsonClassDescription("Add two numbers together")
381 static class CalculateSum implements Supplier<String> {
379 static class CalculateSum {
382380 @JsonPropertyDescription("First number")
383381 public double a;
384382
385383 @JsonPropertyDescription("Second number")
386384 public double b;
385 }
387386
388 @Override
389 public String get() {
390 return String.valueOf(a + b);
387 // Stands in for a class your application already has,
388 // such as a database client or an API wrapper.
389 static class WeatherService {
390 String currentWeather(String location, String unit) {
391 return "{\"temperature\": \"20°C\", \"condition\": \"Sunny\"}";
391392 }
392393 }
393394
394395 void main() {
395396 AnthropicClient client = AnthropicOkHttpClient.fromEnv();
397 WeatherService weatherService = new WeatherService();
396398
399 // The lambda can use weatherService.
400 BetaRunnableTool getWeather = BetaRunnableTool.of(
401 GetWeather.class,
402 input -> BetaToolResultBlockParam.Content.ofString(
403 weatherService.currentWeather(input.location, input.unit)));
404
405 BetaRunnableTool calculateSum = BetaRunnableTool.of(
406 CalculateSum.class,
407 input -> BetaToolResultBlockParam.Content.ofString(
408 String.valueOf(input.a + input.b)));
409
397410 BetaToolRunner runner = client.beta()
398411 .messages()
399412 .toolRunner(MessageCreateParams.builder()
from line 414
401414 .maxTokens(1024)
402415 .addBeta("structured-outputs-2025-11-13")
403416 .addUserMessage("What's the weather like in Paris? Also, what's 15 + 27?")
404 .addTool(GetWeather.class)
405 .addTool(CalculateSum.class)
417 .addTool(getWeather)
418 .addTool(calculateSum)
406419 .build());
407420
408421 for (BetaMessage message : runner) {
from line 721
708721 .maxTokens(1024)
709722 .addBeta("structured-outputs-2025-11-13")
710723 .addUserMessage("What's the weather like in Paris? Also, what's 15 + 27?")
711 .addTool(GetWeather.class)
712 .addTool(CalculateSum.class)
724 .addTool(getWeather)
725 .addTool(calculateSum)
713726 .build());
714727
715728 BetaMessage finalMessage = null;
from line 1004
9911004 .maxTokens(1024)
9921005 .addBeta("structured-outputs-2025-11-13")
9931006 .addUserMessage("Give me a detailed weather report for every major US city.")
994 .addTool(GetWeather.class)
1007 .addTool(getWeather)
9951008 .build())
9961009 .maxIterations(10L)
9971010 .build());
from line 1255
12421255 </Tab>
12431256
12441257 <Tab title="Java">
1245 Intercepting tool errors before they're sent to Claude is not currently supported in the Java SDK. The runner catches any exception thrown from a tool's `get()` method and converts it into a tool result with `is_error: true` automatically. To control the error content, catch the exception inside your tool and return a custom string.
1258 Intercepting tool errors before they're sent to Claude is not currently supported in the Java SDK. The runner catches any exception thrown from a tool's function and converts it into a tool result with `is_error: true` automatically. To control the error content, catch the exception inside the function and return your own content.
12461259 </Tab>
12471260
12481261 <Tab title="PHP">
from line 1429
14161429 </Tab>
14171430
14181431 <Tab title="Java">
1419 To set `cache_control` on a tool result, return `BetaToolResultBlockParam.Content` from the tool instead of `String` and set `cacheControl` on the inner text block. The runner does not currently support setting `cache_control` on the outer `tool_result` block.
1432 To set `cache_control` on a tool result, build the returned `BetaToolResultBlockParam.Content` with `ofBlocks()` instead of `ofString()` and set `cacheControl` on the inner text block. The runner does not currently support setting `cache_control` on the outer `tool_result` block.
14201433
14211434 ```java
14221435 @JsonClassDescription("Look up reference documentation for a topic")
1423 static class SearchDocuments implements Supplier<BetaToolResultBlockParam.Content> {
1436 static class SearchDocuments {
14241437 @JsonPropertyDescription("The search query")
14251438 public String query;
1426
1427 @Override
1428 public BetaToolResultBlockParam.Content get() {
1429 String largeResult = "..."; // a long document worth caching
1430 return BetaToolResultBlockParam.Content.ofBlocks(List.of(
1431 BetaToolResultBlockParam.Content.Block.ofText(
1432 BetaTextBlockParam.builder()
1433 .text(largeResult)
1434 .cacheControl(BetaCacheControlEphemeral.builder().build())
1435 .build())));
1436 }
14371439 }
1440
1441 BetaRunnableTool searchDocuments = BetaRunnableTool.of(SearchDocuments.class, input -> {
1442 String largeResult = "..."; // a long document worth caching
1443 return BetaToolResultBlockParam.Content.ofBlocks(List.of(
1444 BetaToolResultBlockParam.Content.Block.ofText(
1445 BetaTextBlockParam.builder()
1446 .text(largeResult)
1447 .cacheControl(BetaCacheControlEphemeral.builder().build())
1448 .build())));
1449 });
14381450 ```
14391451 </Tab>
14401452
from line 1680
16681680 .maxTokens(1024)
16691681 .addBeta("structured-outputs-2025-11-13")
16701682 .addUserMessage("What is 15 + 27?")
1671 .addTool(CalculateSum.class)
1683 .addTool(calculateSum)
16721684 .build());
16731685
16741686 for (StreamResponse<BetaRawMessageStreamEvent> stream : runner.streaming()) {
build-with-claude/mid-conversation-effort-example Changed · +8 / -8 lines
from line 20
2020
2121The example is a single file. The constants control the effort level, the fan-out shape, and how often the mode refresher is re-sent. `MAX_CONCURRENT` caps how many subagents run at the same time (the PHP port is sequential and ignores it); `MAX_TOTAL_SUBTASKS` caps how many the model may queue in a single Workflow call. Splitting the two lets the model plan a large backlog without launching it all at once. The `DOC_TEST_MODE` check caps the loops to a single turn when that environment variable is set, so the automated docs harness can validate that the file compiles and finishes quickly without running the full orchestration; leave it unset when running the example yourself.
2222
23<CodeGroup>
23<CodeGroup exclude="shell:cURL, shell:CLI">
2424 ```python Python
2525 import atexit
2626 import concurrent.futures
from line 300
300300
301301The reminders are short on purpose. They flip the mode and point at the tool description, where the heavyweight instructions live. The full text is sent once when the mode turns on, the refresher is re-sent only after several user turns, and the exit notice is sent once when the mode turns off.
302302
303<CodeGroup>
303<CodeGroup exclude="shell:cURL, shell:CLI">
304304 ```python Python
305305 MODE_ENTER = (
306306 "Orchestration mode is on: optimize for the most exhaustive, correct answer rather than "
from line 400
400400
401401The Workflow tool carries the real behavioral contract: the opt-in rule, the standing consent that applies while the mode is on, granularity guidance for sizing the fan-out, and the quality patterns the model can reach for (a verification wave, a completeness critic, multiphase sequencing). Subagents also get a `report_findings` tool so their results come back as structured JSON instead of prose, and the bash tool is the Anthropic-defined `bash_20250124` tool run locally.
402402
403<CodeGroup>
403<CodeGroup exclude="shell:cURL, shell:CLI">
404404 ```python Python
405405 WORKFLOW_TOOL = {
406406 "name": "Workflow",
from line 886
886886
887887The bash handler runs the requested command with a timeout, captures combined stdout and stderr, and truncates the result so a runaway command can't flood the context window. Commands run in the directory you launch the example from, so pointing it at a project means starting it there; when `DOC_TEST_MODE` is set, the harness instead gives bash a small throwaway fixture directory that is removed on exit. There is no sandbox here: the command runs with the permissions of the process that launched the example. For clarity this example runs each call in a fresh subshell rather than maintaining the persistent session the `bash_20250124` contract describes; a production agent should back the tool with a long-lived shell so that working directory, environment, and the `restart` action behave as documented.
888888
889<CodeGroup>
889<CodeGroup exclude="shell:cURL, shell:CLI">
890890 ```python Python
891891 # Run bash where the example was launched. In DOC_TEST_MODE the docs harness
892892 # points it at a throwaway fixture directory instead, removed on exit.
from line 1433
14331433
14341434Each workflow subtask becomes its own small agent loop with the bash tool, running at the same effort as the main loop. A per-request timeout bounds each API call so a dropped connection degrades one subagent instead of stalling the whole run.
14351435
1436<CodeGroup>
1436<CodeGroup exclude="shell:cURL, shell:CLI">
14371437 ```python Python
14381438 def run_subagent(model: str, prompt: str) -> str:
14391439 """One subagent: a small nested agent loop with the bash tool plus report_findings.
from line 1977
19771977
19781978A fan-out that spawns dozens of subagents is expensive to restart from scratch. A small content-addressed journal makes it idempotent: before dispatching a subagent, look up the SHA-256 of its prompt in a local JSON file, and return the recorded result if one exists. Interrupt the run, rerun it, and only the subtasks that never finished are recomputed. The journal deduplicates across runs, not within a single fan-out wave; delete the journal file to start fresh.
19791979
1980<CodeGroup>
1980<CodeGroup exclude="shell:cURL, shell:CLI">
19811981 ```python Python
19821982 _journal_lock = threading.Lock()
19831983
from line 2268
22682268
22692269The fan-out accepts up to `MAX_TOTAL_SUBTASKS` prompts, runs them through the journal with at most `MAX_CONCURRENT` in flight (sequential in the PHP port), and isolates failures so one broken subagent degrades to an error string instead of ending the run. Once the first wave finishes, a second wave reuses the same subagent path to try to refute each result: every verifier re-derives the claims from the source, defaulting to refuted when uncertain. Both the original result and its verdict are returned to the orchestrator so it can weigh them together.
22702270
2271<CodeGroup>
2271<CodeGroup exclude="shell:cURL, shell:CLI">
22722272 ```python Python
22732273 def normalize_subtasks(raw) -> list[str]:
22742274 """Accept the subtasks input in whatever shape the model emits: an array, the array
from line 3828
38283828 The bash tool in this example runs model-written commands directly on your machine with no sandbox, and the fan-out runs several of those agents in parallel. Run it in a directory and environment you are comfortable exposing, and add sandboxing before adapting it for anything beyond local experimentation.
38293829</Warning>
38303830
3831<CodeGroup>
3831<CodeGroup exclude="shell:cURL, shell:CLI">
38323832 ```python Python
38333833 if __name__ == "__main__":
38343834 task = (
build-with-claude/refusals-and-fallback Changed · +3 / -3 lines
## How refusals are billed
from line 190
190190
191191A refusal can arrive before any output, or mid-stream after partial output. In either case, treat any partial output as incomplete and discard it.
192192
193<Note>
194 **How refusals are billed:** You are not billed for a refusal that arrives before any output. `content` is empty, and token counts appear in `usage` but are not charged. The request still counts against your rate limits. A mid-stream refusal bills the input tokens and the output already streamed at normal rates.
195</Note>
193## How refusals are billed
194
195You are not billed for a refusal that arrives before any output. `content` is empty, and token counts appear in `usage` but are not charged. The request still counts against your rate limits. A mid-stream refusal bills the input tokens and the output already streamed at normal rates.
196196
197197## Picking a fallback approach
198198
build-with-claude/thinking-tool-workflows Changed · +31 / -0 lines
from line 30
3030 Send a request with adaptive thinking enabled and the tool defined. Apart from the `thinking` parameter, this is a standard [tool use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview) request:
3131
3232 <CodeGroup>
33 ```bash cURL
34 curl https://api.anthropic.com/v1/messages \
35 -H "anthropic-version: 2023-06-01" \
36 -H "content-type: application/json" \
37 -H "x-api-key: $ANTHROPIC_API_KEY" \
38 -d @- <<'EOF'
39 {
40 "model": "claude-opus-4-8",
41 "max_tokens": 16000,
42 "thinking": {"type": "adaptive"},
43 "tools": [{
44 "name": "get_weather",
45 "description": "Get current weather for a location",
46 "input_schema": {
47 "type": "object",
48 "properties": {
49 "location": {"type": "string", "description": "City name"}
50 },
51 "required": ["location"]
52 }
53 }],
54 "messages": [{"role": "user", "content": "What's the weather in Paris?"}]
55 }
56 EOF
57 ```
58
3359 ```bash CLI
3460 ant messages create --transform content <<'YAML'
3561 model: claude-opus-4-8
from line 320
294320 Each sample is a self-contained script: it repeats the first request, then immediately sends the follow-up using the response it just received.
295321
296322 <CodeGroup>
323 ```bash cURL
324 # This workflow does not translate well to a one-off shell command.
325 # Use one of the SDK examples in this code group instead.
326 ```
327
297328 ```bash CLI
298329 # First turn: write the assistant content array (thinking and tool_use
299330 # blocks, signatures intact) to a file. Routing model-generated text
test-and-evaluate/strengthen-guardrails/handle-streaming-refusals Changed · +15 / -3 lines
from line 60
6060
6161<CodeGroup>
6262 ```bash cURL
63 # Stream request and check for refusal
6463 response=$(curl -N https://api.anthropic.com/v1/messages \
6564 -H "anthropic-version: 2023-06-01" \
6665 -H "content-type: application/json" \
from line 71
7271 "stream": true
7372 }')
7473
75 # Check for refusal in the stream
76 if echo "$response" | grep -q '"stop_reason":"refusal"'; then
74 if echo "$response" | jq -R -e 'select(startswith("data: "))
75 | sub("^data: "; "") | fromjson
76 | select(.delta.stop_reason == "refusal")' >/dev/null; then
77 echo "Response refused - resetting conversation context"
78 # Reset your conversation state here
79 fi
80 ```
81
82 ```bash CLI
83 response=$(ant messages create --stream --format jsonl \
84 --model claude-opus-5-5 \
85 --max-tokens 1024 \
86 --message '{role: user, content: Hello}')
87
88 if echo "$response" | jq -e 'select(.delta.stop_reason == "refusal")' >/dev/null; then
7789 echo "Response refused - resetting conversation context"
7890 # Reset your conversation state here
7991 fi
build-with-claude/prompt-caching Changed · +2 / -2 lines
from line 3318
33183318 <Accordion title="Why am I seeing the error `AttributeError: 'Beta' object has no attribute 'prompt_caching'` in Python?">
33193319 This error typically appears when you have upgraded your SDK or you are using outdated code examples. Prompt caching no longer requires the beta prefix. Instead of:
33203320
3321 <CodeGroup>
3321 <CodeGroup exclude="shell:cURL, shell:CLI, typescript, csharp, go, java, php, ruby">
33223322 ```python Python
33233323 client.beta.prompt_caching.messages.create(**params)
33243324 ```
from line 3326
33263326
33273327 Use:
33283328
3329 <CodeGroup>
3329 <CodeGroup exclude="shell:cURL, shell:CLI, typescript, csharp, go, java, php, ruby">
33303330 ```python Python
33313331 client.messages.create(**params)
33323332 ```
build-with-claude/streaming Changed · +2 / -2 lines
from line 10
1010
1111The [Python SDK](https://github.com/anthropics/anthropic-sdk-python) and [TypeScript SDK](https://github.com/anthropics/anthropic-sdk-typescript) offer multiple ways of streaming. The [PHP SDK](https://github.com/anthropics/anthropic-sdk-php) provides streaming through `createStream()`. The Python SDK allows both sync and async streams. See the documentation in each SDK for details.
1212
13<CodeGroup>
13<CodeGroup exclude="shell:cURL">
1414 ```bash CLI
1515 ant messages create --stream --format jsonl \
1616 --model claude-opus-5-5 \
from line 140
140140
141141If you don't need to process text as it arrives, the SDKs provide a way to use streaming internally while returning the complete `Message` object, identical to what `.create()` returns. This is especially useful for requests with large `max_tokens` values, where the SDKs require streaming to avoid HTTP timeouts.
142142
143<CodeGroup>
143<CodeGroup exclude="shell:cURL">
144144 ```bash CLI
145145 # The ant CLI's --stream flag emits one event per line and does not
146146 # accumulate into a final Message. For long generations, stream the
build-with-claude/task-budgets Changed · +1 / -1 lines
from line 512
512512
513513Run a representative sample of tasks **without** `task_budget` set and record the total tokens Claude spends per task. For an agentic loop, sum `usage.output_tokens` across every request in the loop, plus the tokens of the tool results you append between requests:
514514
515<CodeGroup>
515<CodeGroup exclude="shell:cURL">
516516 ```bash CLI
517517 ant messages create --transform 'usage.output_tokens' <<'YAML'
518518 model: claude-opus-5-5
manage-claude/spend-limits-api Changed · +1 / -1 lines
from line 74
7474| ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
7575| `pending` | Awaiting admin action. The request normally carries a live `spend_summary` so you can see the member's current effective spend limit and period-to-date spend while deciding; `spend_summary` may be `null` if it could not be computed. |
7676| `approved` | The request was resolved with approval: either an admin approved it explicitly, another admin action raised the member's spend limit, or Anthropic support raised a spend limit on the organization's behalf. `spend_summary` is `null`. |
77| `denied` | An admin declined. `spend_summary` is `null`. claude.ai hides that member's request button for 30 days from `resolved_at`; an admin can still raise the member's spend limit directly at any time. |
77| `denied` | An admin declined. `spend_summary` is `null`. The member can send a new request right away; only a `pending` request blocks a new one. An admin can still raise the member's spend limit directly at any time. |
7878
7979Both `approved` and `denied` are terminal. A member has at most one `pending` request at a time.
8080
manage-claude/workload-identity-federation Changed · +1 / -1 lines
from line 89
8989
9090You can construct the client with explicit credentials or with no arguments. With no arguments, the SDK resolves credentials from environment variables or the active profile, as described under [Credential precedence](https://platform.claude.com/docs/en/manage-claude/workload-identity-federation#credential-precedence). The zero-argument form is the recommended pattern for production workloads: ship the same container image everywhere and inject `ANTHROPIC_FEDERATION_RULE_ID`, `ANTHROPIC_ORGANIZATION_ID`, `ANTHROPIC_SERVICE_ACCOUNT_ID`, `ANTHROPIC_WORKSPACE_ID`, and `ANTHROPIC_IDENTITY_TOKEN_FILE` per environment.
9191
92<CodeGroup>
92<CodeGroup exclude="shell:CLI">
9393 ```bash cURL
9494 # 1. Acquire your IdP's JWT (platform-specific; see the per-provider guides).
9595 JWT=$(cat /var/run/secrets/anthropic.com/token)
managed-agents/agent-setup Changed · +1 / -1 lines
from line 471
471471
472472The preceding example supplies `version` from the create response, so the update only applies if nothing else has changed the agent since you read it. To apply an update unconditionally, omit `version` from the request:
473473
474<CodeGroup>
474<CodeGroup exclude="shell:CLI, python, typescript, csharp, go, java, php, ruby">
475475 ```bash cURL
476476 updated_agent=$(curl -fsSL "https://api.anthropic.com/v1/agents/$AGENT_ID" \
477477 -H "x-api-key: $ANTHROPIC_API_KEY" \
managed-agents/migration Changed · +1 / -1 lines
from line 29
2929
3030**Before** (Messages API loop, simplified):
3131
32<CodeGroup>
32<CodeGroup exclude="shell:cURL, shell:CLI">
3333 ```python Python
3434 messages = [{"role": "user", "content": task}]
3535 while True:
managed-agents/webhooks Changed · +2 / -2 lines
from line 115
115115
116116Set `ANTHROPIC_WEBHOOK_SIGNING_KEY` to the `whsec_`-prefixed secret shown at endpoint creation.
117117
118<CodeGroup>
118<CodeGroup exclude="shell:cURL, shell:CLI">
119119 ```python Python
120120 from flask import Flask, request
121121 import anthropic
from line 358
358358}
359359```
360360
361<CodeGroup>
361<CodeGroup exclude="shell:cURL, shell:CLI">
362362 ```python Python
363363 if event.data.type == "session.status_idled":
364364 session = client.beta.sessions.retrieve(event.data.id)