Follow Discord
Sweep 09 Oct 2026 · 17:27Z Build v2.1.296 517 read Stable v2.1.287 Latest v2.1.296 Next v2.1.296 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · api

optimizing-for-cost-and-intelligence changedabout-claude/models/optimizing-for-cost-and-intelligence

Nearest release: v2.1.296, published an hour after this site recorded the change. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Recorded here
Lines+21added
Lines−12removed
From line 132 where the diff opens
First seen 14 Aug 2026 this site's first read of the page
Recorded edits18to this page, all time

The whole hunk

from line 132, old and new numbered
/
lines
from line 132
132132 
133133Several things can break your cache during a task. Anything that changes per request, such as a timestamp or a queue position, placed ahead of the stable prefix turns every request into a full cache write: on the triage run in [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens), a 25-token status line at the front of the system prompt cost $4.24 USD per run instead of $0.59 USD, more than running with caching off. Keep per-request text in the newest user turn.
134134 
135The cache is a byte-exact prefix match over the request in order (tools, then system prompt, then messages), so a change anywhere invalidates everything after it. Changing [`effort`](https://platform.claude.com/docs/en/build-with-claude/effort) or the thinking configuration between requests invalidates the cache from that point onward, and on some models the tools and system prompt ahead of it as well; any edit to the system prompt invalidates the cache from that point onward; setting or changing an output format invalidates the cache for the whole conversation; adding, removing, or reordering a tool definition invalidates all of it. The [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#what-invalidates-the-cache) page lists these cases, apart from the output format, which [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs#prompt-modification-and-token-costs) covers. On the most recent models, change instructions with a [mid-conversation system message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages), a `{"role": "system"}` message appended to `messages`, instead of editing the top-level `system` field: the cached prefix stays intact. Check that page for which models support it. On models that support it, a [per-message effort change](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) leaves the cached prefix intact too. The stakes are highest on Claude Fable 5.1 and Claude Mythos 5.1, where a break re-writes the prefix at 1.25x the input price instead of reading it at 0.025x. On a 100,000-token prefix, one broken turn there costs $1.25 USD instead of $0.03 USD, 50 times the read; on Claude Opus 5.5 it costs $0.50 USD instead of $0.02 USD, and on Claude Sonnet 5.5 $0.25 USD instead of $0.01 USD, 25 times, and on the other current models 12.5 times.
135The cache is a byte-exact prefix match over the request in order (tools, then system prompt, then messages), so a change anywhere invalidates everything after it. Changing [`effort`](https://platform.claude.com/docs/en/build-with-claude/effort) or the thinking configuration between requests invalidates the cache from that point onward, and on some models the tools and system prompt ahead of it as well; any edit to the system prompt invalidates the cache from that point onward; setting or changing an output format invalidates the cache for the whole conversation; adding, removing, or reordering a tool definition invalidates all of it. The [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching#what-invalidates-the-cache) page lists these cases, apart from the output format, which [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs#prompt-modification-and-token-costs) covers. On the most recent models, change instructions with a [mid-conversation system message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages), a `{"role": "system"}` message appended to `messages`, instead of editing the top-level `system` field: the cached prefix stays intact. Check that page for which models support it. A per-message effort change leaves the cached prefix intact too, but on fewer models: [Change effort mid-conversation](https://platform.claude.com/docs/en/build-with-claude/effort#changing-effort-mid-conversation) lists them. The stakes are highest on Claude Fable 5.1 and Claude Mythos 5.1, where a break re-writes the prefix at 1.25x the input price instead of reading it at 0.025x. On a 100,000-token prefix, one broken turn there costs $1.25 USD instead of $0.03 USD, 50 times the read; on Claude Opus 5.5 it costs $0.50 USD instead of $0.02 USD, and on Claude Sonnet 5.5 $0.25 USD instead of $0.01 USD, 25 times, and on the other current models 12.5 times.
136136 
137137Anthropic measured this on the triage agent's long sessions[18](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs). An effort change and an added tool made mid-session rewrote 39,000 and 60,000 cached tokens, and those sessions cost $0.95 USD per session. The same two changes on the first request after compaction cost $0.75 USD, and on the request that triggered the compaction $0.92 USD, because the compaction's summarization pass then re-processed the 81,000-token context at the cache-write price: that summarization pass cost $0.21 USD, against $0.04 USD when the same changes came one request later, with accuracy within run-to-run noise in every arm:
138138 
from line 302
302302 
303303On harder work the gap widens. On Terminal-Bench 3[20](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), where the tasks are hard enough that pass rate rather than tokens sets the bill, Claude Opus 4.7, Opus 4.8, and Opus 5 each spend $8 USD to $15 USD per task but solve 7%, 15%, and 41% of tasks, so cost per solved task falls from $183 USD to $63 USD to $28 USD up the ladder. The 21% premium Claude Opus 5 carries over Opus 4.8 on the saturated coding subset becomes a 56% saving on Terminal-Bench 3, where the older model mostly fails: the more your workload defeats the old model, the more the upgrade saves per result.
304304 
305Compare on cost per solved task, not per token: the same text costs about 30% more tokens on Claude Opus 4.7 and later, so a per-token comparison makes the newer models look more expensive by construction.
305Compare on cost per solved task, not per token: the same text costs about 30% more tokens on Claude 4.7 and later models, so a per-token comparison makes the newer models look more expensive by construction.
306306 
307307### Tune effort
308308 
from line 602
602602 
603603The numbers on this page reflect list prices at the time of measurement and will drift as models and prices change. Your escalation rate, how cleanly tasks split, and transcript length move them too. The method stays the same:
604604 
6051. Pull a few tasks from production logs, weighted like real traffic, and [write outcome checks](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests) for each: tests pass, ticket closed, row count correct. Record cost per task beside the score: price the five priced token counts in each response's `usage` at their own rates (uncached input, 5-minute and 1-hour cache writes at 1.25x and 2x the input price, cache reads, and output), summed across the task's requests (the [Usage and Cost API](https://platform.claude.com/docs/en/manage-claude/usage-cost-api) reports the aggregate).
6051. Pull a few tasks from production logs, weighted like real traffic, and [write outcome checks](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests) for each: tests pass, ticket closed, row count correct. Record cost per task beside the score: price the five priced token counts in each response's `usage` at their own rates (uncached input, 5-minute and 1-hour cache writes at 1.25x and 2x the input price, cache reads, and output), summed across the task's requests (the [Usage and Cost API](https://platform.claude.com/docs/en/manage-claude/usage-cost-api) reports the aggregate). On a model priced by prompt length, such as Claude Haiku 5.5, price each request at the rates for its prompt length: a request whose prompt, cache reads and writes included, is over the model's threshold pays the higher prices for all of its tokens (see [Long context pricing](https://platform.claude.com/docs/en/about-claude/pricing#long-context-pricing)).
6066062. Baseline the model tiers across effort levels, not only the default, and plot score against spend. A multi-model configuration must beat the single model's whole curve.
6076073. If the curve shows a gap effort can't close, add the multi-model strategy that fits and re-run the suite.
6086084. Run the winner in shadow on a traffic slice before cutover, then keep the suite running.
from line 611
611611 
612612<CodeGroup>
613613 ```bash cURL
614 # Per-million-token prices from the pricing page; change these three for another model.
614 # Per-million-token prices from the pricing page; change these for another model.
615 # A model priced by prompt length, such as Claude Haiku 5.5, charges its long-prompt prices over its threshold.
615616 INPUT_PER_MTOK=4.00 # Claude Opus 5.5
616617 CACHE_READ_PER_MTOK=0.20 # 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
617618 OUTPUT_PER_MTOK=20.00
from line 639
638639 ```
639640 
640641 ```bash CLI
641 # Per-million-token prices from the pricing page; change these three for another model.
642 # Per-million-token prices from the pricing page; change these for another model.
643 # A model priced by prompt length, such as Claude Haiku 5.5, charges its long-prompt prices over its threshold.
642644 INPUT_PER_MTOK=4.00 # Claude Opus 5.5
643645 CACHE_READ_PER_MTOK=0.20 # 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
644646 OUTPUT_PER_MTOK=20.00
from line 662
660662 ```
661663 
662664 ```python Python
663 # Per-million-token prices from the pricing page; change these three for another model.
665 # Per-million-token prices from the pricing page; change these for another model.
666 # A model priced by prompt length, such as Claude Haiku 5.5, charges its long-prompt prices over its threshold.
664667 INPUT_PER_MTOK = 4.00 # Claude Opus 5.5
665668 # 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
666669 CACHE_READ_PER_MTOK = 0.20
from line 691
688691 ```
689692 
690693 ```typescript TypeScript
691 // Per-million-token prices from the pricing page; change these three for another model.
694 // Per-million-token prices from the pricing page; change these for another model.
695 // A model priced by prompt length, such as Claude Haiku 5.5, charges its long-prompt prices over its threshold.
692696 const INPUT_PER_MTOK = 4.0; // Claude Opus 5.5
693697 const CACHE_READ_PER_MTOK = 0.2; // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
694698 const OUTPUT_PER_MTOK = 20.0;
from line 715
711715 ```
712716 
713717 ```csharp C#
714 // Per-million-token prices from the pricing page; change these three for another model.
718 // Per-million-token prices from the pricing page; change these for another model.
719 // A model priced by prompt length, such as Claude Haiku 5.5, charges its long-prompt prices over its threshold.
715720 const double InputPerMtok = 4.00; // Claude Opus 5.5
716721 const double CacheReadPerMtok = 0.20; // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
717722 const double OutputPerMtok = 20.00;
from line 743
738743 ```
739744 
740745 ```go Go
741 // Per-million-token prices from the pricing page; change these three for another model.
746 // Per-million-token prices from the pricing page; change these for another model.
747 // A model priced by prompt length, such as Claude Haiku 5.5, charges its long-prompt prices over its threshold.
742748 const (
743749 inputPerMTok = 4.00 // Claude Opus 5.5
744750 cacheReadPerMTok = 0.20 // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
from line 775
769775 ```
770776 
771777 ```java Java
772 // Per-million-token prices from the pricing page; change these three for another model.
778 // Per-million-token prices from the pricing page; change these for another model.
779 // A model priced by prompt length, such as Claude Haiku 5.5, charges its long-prompt prices over its threshold.
773780 static final double INPUT_PER_MTOK = 4.00; // Claude Opus 5.5
774781 static final double CACHE_READ_PER_MTOK = 0.20; // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
775782 static final double OUTPUT_PER_MTOK = 20.00;
from line 803
796803 ```
797804 
798805 ```php PHP
799 // Per-million-token prices from the pricing page; change these three for another model.
806 // Per-million-token prices from the pricing page; change these for another model.
807 // A model priced by prompt length, such as Claude Haiku 5.5, charges its long-prompt prices over its threshold.
800808 const INPUT_PER_MTOK = 4.00; // Claude Opus 5.5
801809 const CACHE_READ_PER_MTOK = 0.20; // 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
802810 const OUTPUT_PER_MTOK = 20.00;
from line 827
819827 ```
820828 
821829 ```ruby Ruby
822 # Per-million-token prices from the pricing page; change these three for another model.
830 # Per-million-token prices from the pricing page; change these for another model.
831 # A model priced by prompt length, such as Claude Haiku 5.5, charges its long-prompt prices over its threshold.
823832 INPUT_PER_MTOK = 4.00 # Claude Opus 5.5
824833 CACHE_READ_PER_MTOK = 0.20 # 0.05x the input price on Claude Opus 5.5; the multiplier differs by model
825834 OUTPUT_PER_MTOK = 20.00
Feedback