Follow Discord
Sweep 25 Sep 2026 · 19:33Z Build v2.1.283 504 read Stable v2.1.274 Latest v2.1.283 Next v2.1.283 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · api

optimizing-for-cost-and-intelligence changedabout-claude/models/optimizing-for-cost-and-intelligence

Nearest release: v2.1.282, published an hour before this site recorded the change. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Recorded here
Lines+103added
Lines−2removed
From line 13 where the diff opens
First seen 14 Aug 2026 this site's first read of the page
Recorded edits12to this page, all time

### Show the model elapsed time

The whole hunk

from line 13, old and new numbered
/
lines
from line 13
1313The levers come in two kinds:
1414 
1515* **Free wins** cut spend without touching quality: prompt caching, token hygiene, a prompt audit against the model you are running, [batch processing](https://platform.claude.com/docs/en/build-with-claude/batch-processing) at 50% off for work that can wait up to 24 hours, and [workspace spend limits](https://platform.claude.com/docs/en/api/rate-limits#setting-lower-limits-for-workspaces) as the backstop.
16* **Tradeoffs** exchange cost for intelligence: model choice, effort, output caps and task budgets, and multi-model architectures.
16* **Tradeoffs** exchange cost for intelligence: model choice, effort, output caps and task budgets, an elapsed-time clock, and multi-model architectures.
1717 
1818Each lever comes with measured results and the rule for when it pays. In Anthropic's measurements, prompt caching was the largest lever by a wide margin: it cut agent-loop cost by a factor of 2.7 to 5.3 on this guide's benchmarks and cut a small triage agent's bill by 83%, or 88% with input trimming added. The multi-model levers are narrower; a second model paid off in two shapes, an advisor and an orchestrator.
1919 
from line 32
3232| Attempts end with `stop_reason: max_tokens` | Raise `max_tokens`; 64,000 covered all but 2 of 14,000 turns measured at the default effort, and 128,000 cost nothing extra per solved task | [Set budgets](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
3333| You can check outputs (tests, a verifier) | Run everything at low effort and re-run failures at `high`; on the coding benchmark measured, the pass rate held at about half the cost | [Re-run failures](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#re-run-failures-at-higher-effort) |
3434| Agent loops with a few very costly runs | Set a task budget (beta; check the support table for which models), a Claude Managed Agents session budget, and a workspace spend limit | [Set budgets](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#set-budgets-and-output-caps) |
35| You want agent runs to finish sooner | Tell the model that time matters, and show it the elapsed time; on DRACO, HLE, and an internal physics set, runs took 33% to 69% less time at a 28% to 54% lower cost per task, with scores up to 1.9 points lower | [Show the model elapsed time](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#show-the-model-elapsed-time) |
3536| A lower-cost model stalls only on hard decisions | Add a frontier advisor. It pays off when priced well above the executor and actually consulted, so first price the advisor's model alone at low effort and measure the consult rate | [Advisor strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#advisor-strategy-escalate-hard-decisions) |
3637| The work exceeds one context window | Delegate partitions to cheaper workers | [Orchestrator strategy](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#orchestrator-strategy-delegate-bulk-work) |
3738 
from line 244
243244 
244245## Trade cost against intelligence
245246 
246These levers set where a single model sits between cost and intelligence: model choice, effort, re-running failures at a higher setting, and the budgets and caps it works within. Start with an effort sweep on your current model ([Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort)). From lowest to highest cost and capability, the current models are Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 5.5, and Claude Fable 5.1 (the frontier model); [Models overview](https://platform.claude.com/docs/en/models/overview) has the full lineup and prices.
247These levers set where a single model sits between cost and intelligence: model choice, effort, re-running failures at a higher setting, the budgets and caps it works within, and whether it can see how much time has passed. Start with an effort sweep on your current model ([Tune effort](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#tune-effort)). From lowest to highest cost and capability, the current models are Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 5.5, and Claude Fable 5.1 (the frontier model); [Models overview](https://platform.claude.com/docs/en/models/overview) has the full lineup and prices.
247248 
248249### Compare models on cost per task
249250 
from line 363
362363 
363364![Dot plot of per-turn output for Opus 5.5 and Fable 5.1: medians a few hundred tokens, longest 61k and 128k, against the caps](https://platform.claude.com/docs/images/cost-intel-max-tokens-ladder-opus-5-5.png)
364365 
366### Show the model elapsed time
367 
368A model in an agent loop can't see a clock. A [task budget](https://platform.claude.com/docs/en/build-with-claude/task-budgets) shows it how many tokens are left, but by default nothing in the request shows it how long the work has taken. Two small changes give it that signal. Add a two-sentence instruction to the system prompt that says time matters, and from the second request on, send the elapsed time before each of the model's turns.
369 
370Anthropic measured both changes together with Claude Fable 5.1 at `high` effort, on two public benchmarks, DRACO[21](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) and HLE[22](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), and on an internal set of 70 research-level physics problems, adapted from the public CritPt benchmark[23](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs). This page calls that set the physics set. Each of the three ran in two shapes: a single agent, and a team in which a lead agent starts helper agents of the same model that work in parallel. A score change counts as inside the margin when its 95% interval stays within a limit that Anthropic set before the runs: 1.5 points on DRACO and 2.5 points on HLE. The following chart plots score against cost per task for each configuration. A second row of bars gives each configuration's time as a ratio to the single agent at `high` effort, without retry waits. A third row gives the score change that both changes make, with its 95% interval:
371 
372![Charts, DRACO, HLE, and the physics set: both changes cut each setup's cost and time, and its score moves by under 2 points](https://platform.claude.com/docs/images/cost-intel-time-awareness.png)
373 
374**With a team of agents.** A team does more work than a single agent, so by default it costs more. On DRACO, the team cost 4.0 times as much as the single agent and took about as long (95% interval 12% less to 13% more). With the instruction and the clock on every agent, the team finished in 33% less time at a 54% lower cost per task. Its score was 1.5 points lower (95% interval 0.9 to 2.1 lower), and the far end of that interval, 2.1 points lower, is past the 1.5-point margin. On HLE, the team finished in 51% less time at a 54% lower cost per task. Its score was 1.7 points lower (95% interval 0.3 to 3.1 lower), and the far end of that interval, 3.1 points lower, is past the 2.5-point margin. On the physics set[23](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), the team finished in 39% less time. Its cost per task was 28% lower, and that saving depends on how often the prompt cache expired between requests. With no expiry, it would be 23%. Its score was 0.2 points higher (95% interval 1.5 lower to 2.0 higher).
375 
376On DRACO, the lead started a median of 4 helpers per attempt, so the DRACO result shows a team working in parallel. On HLE and the physics set, the lead started a median of 0 helpers, so at least half of those team runs had only the lead agent. Those team results mostly show the lead agent's own behavior, not the effect of parallel helpers.
377 
378**With a single agent.** On the physics set[23](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), the same changes cut a single agent's time by 34% and its cost per task by 34%. Its score was 0.2 points lower (95% interval 2.5 lower to 2.1 higher). On the physics set, a lower effort level saved cost but not clearly time. At `medium` effort, the single agent cost 37% less per task than at `high`, and its time was 9% less (95% interval 30% less to 16% more). It scored 3.4 points lower (95% interval 0.4 to 6.8 lower), and the interval reaches close to zero. With both changes at `high` effort, the single agent took 27% less time than at `medium` effort (95% interval 5% less to 44% less). Its cost per task was 6% more (95% interval 12% less to 27% more), and its score was 3.2 points higher (95% interval 0.1 lower to 6.5 higher).
379 
380On HLE, the same changes cut a single agent's time by 54% and its cost per task by 48%. Its score was 1.1 points lower (95% interval 2.6 lower to 0.3 higher), and the far end of that interval, 2.6 points lower, is just past the 2.5-point margin. At `medium` effort, the single agent cost 43% less per task than at `high`, took 39% less time, and scored 1.3 points lower (95% interval 2.8 lower to 0.1 higher). With both changes at `high` effort, the single agent took 25% less time than at `medium` effort (95% interval 12% less to 35% less). Its cost per task was 9% less (95% interval 21% less to 6% more), and its score was 0.2 points higher (95% interval 1.3 lower to 1.7 higher).
381 
382On DRACO, the same changes cut a single agent's time by 69% and its cost per task by 49%. Its score was 1.9 points lower (95% interval 1.1 to 2.8 lower), and the far end of that interval, 2.8 points lower, is past the 1.5-point margin. At `medium` effort, the single agent cost 25% less per task than at `high`, took 30% less time, and scored 0.7 points lower (95% interval 0.1 to 1.3 lower). With both changes at `high` effort, the single agent took 53% less time than at `medium` effort (95% interval 42% less to 63% less), and its cost per task was 31% less (95% interval 28% less to 35% less). Its score was 1.2 points lower (95% interval 0.5 to 1.9 lower), and the far end of that interval, 1.9 points lower, is past the 1.5-point margin.
383 
384On all three sets, both changes at `high` effort saved more time than `medium` effort did. On HLE and the physics set, there was no clear difference in cost, and on DRACO the cost was lower. The score was about the same on HLE. On the physics set it was 3.2 points higher, but that interval includes zero, so the difference isn't clear. So for a single agent, the clock saves more time than a lower effort level does. On DRACO, though, the single agent with both changes scored 1.2 points lower than at `medium` effort (95% interval 0.5 to 1.9 lower).
385 
386**When to use it.**
387 
388* Use both changes when an agent's time matters and a small score change is acceptable. On every configuration measured, they cut the time and cost per task, for teams and for single agents.
389* Check score on your own tasks before you adopt them. On DRACO, the score was 1.5 points lower for a team and 1.9 points lower for a single agent. On HLE, it was 1.7 points lower for a team and 1.1 points lower for a single agent. On the physics set, neither score change was clearly different from zero.
390* If you're already thinking about a lower effort level to save time, compare it with the clock. A single agent with both changes at `high` effort took less time than at `medium` effort: 53% less on DRACO, 25% less on HLE, and 27% less on the physics set. Its cost per task was 31% lower on DRACO, with no clear difference on HLE and the physics set.
391 
392**How to add it.** Put this instruction at the start of every agent's system prompt:
393 
394```text wrap
395Time matters here: do not spend time that can be avoided, and the earlier a correct result is obtained, the better. The elapsed time so far is shown before each of your turns.
396```
397 
398The second sentence tells the model that the clock messages exist. The first request carries no clock, and the measured runs used exactly this wording.
399 
400Then, before each request after an agent's first, append a [mid-conversation system message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages) that gives the elapsed time in whole seconds, such as `Elapsed time: 412 seconds`. Count from the start of the task, not from the start of the agent. In a team, every agent reads the same clock, so the first clock a helper sees already counts the time the team spent before the helper started. In a tool loop, put the message right after the `user` message that carries the tool results, as [Placement after tool results](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#placement-after-tool-results) shows. If you send the agent a new `user` message instead, put the clock after that message.
401 
402Leave earlier clock messages where they are. Each one becomes part of the conversation history, so the cached prefix still matches on the next request (see [Combining with prompt caching](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#combining-with-prompt-caching)). Anthropic measured these plain system messages, which stay visible to the model. A [turn-scoped system message](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#turn-scoped-system-messages) would show the model only the newest clock, and Anthropic didn't measure that form.
403 
404The following example runs one agent's tool loop with both changes. It adds the clock after tool results, and it handles client tools only:
405 
406```python
407import time
408 
409import anthropic
410 
411client = anthropic.Anthropic()
412 
413TIME_MATTERS = (
414 "Time matters here: do not spend time that can be avoided, and the earlier a "
415 "correct result is obtained, the better. The elapsed time so far is shown before "
416 "each of your turns."
417)
418 
419 
420def run_agent(task, system, tools, run_tool, started_at=None):
421 """Run one agent's tool loop. In a team, pass the lead's started_at to every helper."""
422 if started_at is None:
423 # Wall-clock seconds, so helpers in other processes can share the lead's start time.
424 started_at = time.time()
425 messages = [{"role": "user", "content": task}]
426 while True:
427 # Stream because a 128,000-token cap is too large for a non-streaming request.
428 with client.messages.stream(
429 model="claude-fable-5-1",
430 max_tokens=128000,
431 cache_control={"type": "ephemeral"},
432 system=TIME_MATTERS + "\n\n" + system,
433 tools=tools,
434 messages=messages,
435 ) as stream:
436 response = stream.get_final_message()
437 messages.append({"role": "assistant", "content": response.content})
438 if response.stop_reason != "tool_use":
439 return response
440 results = [
441 {
442 "type": "tool_result",
443 "tool_use_id": block.id,
444 "content": run_tool(block.name, block.input),
445 }
446 for block in response.content
447 if block.type == "tool_use"
448 ]
449 messages.append({"role": "user", "content": results})
450 # A system message must follow a user turn, so the clock goes after the tool results.
451 elapsed = int(time.time() - started_at)
452 messages.append(
453 {"role": "system", "content": f"Elapsed time: {elapsed} seconds"}
454 )
455```
456 
457Claude Fable 5.1 supports mid-conversation system messages. The [list of supported models](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages) covers the others. On a model without them, such as Claude Sonnet 5, you can put the same line in a text block after the last `tool_result` block in the `user` turn. Anthropic measured only the system-message form.
458 
459On Claude Managed Agents, you can send a [`system.message` event](https://platform.claude.com/docs/en/managed-agents/events-and-streaming#sending-system-messages) with a tool result or a user message. The message applies to that turn and every later turn. So turns that follow the platform's built-in tools, such as web search, see the last clock you sent, not the current time. A `system.message` also reaches only the session's primary thread. In a [multiagent session](https://platform.claude.com/docs/en/managed-agents/multiagent-orchestration), that is the coordinator's thread, so the worker agents never see a clock that you send this way. To show the current time before every turn, and to every agent in a team, run the agent loop yourself on the Messages API.
460 
365461## Combine models
366462 
367463Multi-model architectures fit workloads whose task complexity varies enough that different steps are best served by different models. When your traffic mixes routine work that a smaller model handles reliably with harder steps that need frontier capability, splitting the work keeps frontier intelligence where it matters while most tokens bill at smaller-model rates. When a workload lacks that mix, because its difficulty is uniform or it is one dependent chain, a single well-tuned model is usually the better choice. Each strategy section gives the rule for telling the two cases apart.
from line 511
415511 
416512This pattern saves wall-clock time when workers can run in parallel: on the corpus benchmark[8](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), an episode took about 2.3 hours with the coordinator running the platform's documented limit of 25 concurrent workers, compared with 15 to 20 hours solo. It saved money in only two measured situations. On work a single model could handle alone, the same model at lower effort was cheaper every time.
417513 
514When workers run in parallel, a time instruction and an elapsed-time clock can shorten the run. On DRACO[21](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs), a team of same-model agents with the instruction and the clock finished in 33% less time at a 54% lower cost per task, and scored 1.5 points lower. Every agent in that team had the instruction and the clock. Anthropic didn't measure the clock with lower-cost workers. On Claude Managed Agents, the clock reaches only the coordinator, so the workers never see it. Anthropic didn't measure a team in which only the coordinator has the clock. The coordinator's clock is also current only on turns that follow your own tool results or messages. [Show the model elapsed time](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#show-the-model-elapsed-time) has the recipe for an agent loop that you run on the Messages API.
515 
418516**Case 1: insurance against the cost tail on routine work.** A frontier model running alone occasionally spirals on a routine problem it would normally solve. Because you cannot tell in advance which those will be, a few such runs dominate the bill. A coordinator that hands routine work to a lower-cost worker caps that tail, because any spiraling now happens at worker rates.
419517 
420518Anthropic measured this on a deliberately easy slice of BrowseComp[4](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#refs) (10 problems the solo model reliably solves; 50 delegated and 70 solo runs). A Claude Fable 5 coordinator with one Claude Sonnet 5 worker cost about half as much as Claude Fable 5 alone on average and about a third as much at the 90th percentile ($12 compared with $33), and the solo model's single most expensive run, at $84, was also wrong:
from line 837
73983718. **Compaction timing measurement:** The triage agent's long variant from [Trim input and context tokens](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence#trim-input-and-context-tokens), run August 24, 2026, on Claude Sonnet 5 with the 5-minute cache, cost from the usage fields at list prices, five sessions per arm: a no-change arm at the default effort throughout ($0.81 per session), and two arms that start at low effort and make the same two cache-breaking changes, a switch to the default effort and one added tool, either mid-session at requests 12 and 17 ($0.95) or together on the first request after the first compaction ($0.75). A fourth arm of six sessions, run August 25, 2026, made the same two changes on the request that triggered the first compaction ($0.92 per session): that request's summarization pass wrote the 81,000-token context to the cache instead of reading it, so that pass cost $0.21 against $0.04 for the same pass in the boundary arm. Sessions first compacted at request 21 to 25 (16 of the 21 sessions at request 22), once the prompt passed the 80,000-token compaction trigger, and two no-change sessions compacted a second time near the end. The boundary arm's lower total than the no-change arm reflects its low-effort requests before the change and those second compactions rather than caching: the two arms' re-write costs differ by under a cent. The mid-session arm paid $0.23 per session in cache re-writes; the difference between the mid-session and boundary arms was $0.20 with a 95% confidence interval of $0.11 to $0.29. One mid-session session ran cheap ($0.82) after its model mis-called the search tool following compaction and got empty results; it is included, and without it the arm averages $0.98. Accuracy averaged 14.2 of 20 labels in each August 24 arm and 14.7 in the August 25 arm; cache reads were 91% of prompt tokens with no changes, 85% mid-session, 91% at the boundary, and 86% with the changes on the triggering request.
74083819. **Cache duration measurement on Claude Fable 5.1:** The same 20-issue triage job and harness as reference 16, run August 23 and August 26, 2026, on the Claude Fable 5.1 launch snapshot at its launch prices ($10 input, $12.50 5-minute write, $20 1-hour write, $0.25 cache read, $50 output per million tokens), three settings per schedule: the 5-minute cache, the 1-hour cache, and the 5-minute cache kept warm by a `max_tokens: 0` request on the unchanged prefix every 4 minutes, timed from the previous request's start (the August 23 runs sent keep-alive requests with `max_tokens: 1`; in the August 26 cells reported here, every keep-alive request refreshed the cache and billed no output). Schedules: no pauses, 10% of turns, and every turn at 6 minutes on all 20 issues, and 45-minute pauses on the 5-issue subset; three runs per cell (six for the August 26 keep-alive cell with 45-minute pauses), cost computed from each response's `usage` fields at list prices, accuracy against the same gold labels (12 to 17 exact labels of 20). Per-session means on August 26 for the 5-minute, 1-hour, and keep-alive settings: no pauses $2.42, $3.09, $2.29; 10% paused $4.50, $2.96, $2.36; every turn $22.89, $3.01, $2.62; the August 23 cells agree within 6%. The 45-minute figures ($1.68, $0.59, and $0.71 per 5-issue session) are from a clean re-run on August 26 after a cache-billing incident spoiled that day's first cells; the August 23 runs gave $1.67, $0.58, and $0.70. The crossover between the 5-minute and 1-hour settings is 3.1% of turns, the same measure as reference 16.
74183920. **Terminal-Bench 3:** the public terminal-agent benchmark's 74 tasks, run on [Claude Managed Agents](https://platform.claude.com/docs/en/managed-agents/overview) with two custom tools, a shell and a file editor that the evaluation harness runs in each task's own container, in place of the platform's built-in tools, and otherwise at the platform's default settings for external accounts, two runs per model at `high` effort, August 27 to 28, 2026. These runs used Terminal-Bench version 3.0, and their scores are not comparable with the public Terminal-Bench leaderboard or with the Terminal-Bench 4.0 results in the Claude Opus 5.5 system card, which come from runs in Claude Code at `max` effort. Each task's time limits are 2.5 times the benchmark's own, which gives the agent between 75 minutes and 20 hours per task (5 hours for the median task), and each task gets three times the memory it specifies, from 6 GiB to 96 GiB, with extra memory for the 12 tasks that run helper services. The agent had no general internet access: its containers could reach an internal package mirror, a short list of download sites including GitHub and the Python Package Index, and a few sites specific to some tasks, and eight of the tasks had no network access at all. Scores are raw pass rates over the 148 attempts per model; single runs swing by 5 to 11 points. Costs are what a customer would be billed at list prices, re-priced request by request from the runs' usage records with the 5-minute cache lifetime. Claude Opus 4.7 ended 11 of its 148 attempts at its output cap.
84021. **DRACO:** Perplexity, "DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity," arXiv:2602.11685, 2026. Its 100 research tasks across 10 domains are graded against expert-written rubrics, and the score is the benchmark's normalized score. Every configuration ran on the Claude API with Claude Fable 5.1, the default adaptive thinking, the production safety classifiers on, and `max_tokens` at 128,000: a single agent at `high` and at `medium` effort, the single agent at `high` with the instruction and the clock, and a team at `high` with and without them. The team is a lead agent that starts helper agents of the same model through a tool, with no cap on their number. On DRACO, the lead started a median of 4 helpers per attempt. Each configuration made three attempts at each task, run September 8 to 10, 2026. An attempt that hit the four-hour limit was run again, and the new attempt counts. The only attempts left out are all 3 attempts at one task for the single agent at `medium` effort, so that configuration covers 99 tasks. That task timed out on every attempt, in the original run and in the re-run. Scoring those 3 attempts as 0, as the benchmark's own scoring would, affects only the two comparisons with `medium` effort. The score change at `medium` effort against `high` moves from 0.7 to 1.7 points lower, and the score change with both changes against `medium` effort moves from 1.2 to 0.2 points lower. The agents used a search tool and a fetch tool that the evaluation harness hosts over a pinned web index. Those tools set part of the time, and yours will run at a different speed, so the page gives time as a ratio between configurations, not in minutes. Time is the wall-clock time per task, from the first request to the last request on any agent, minus the estimated time spent waiting to retry requests after rate-limit or overload errors. Those errors came from the test account's shared limits. All configurations of a set started together. The slower ones finished hours later, so part of their time ran under different load. Each task's cost is its requests priced at public list prices, with prompt caching billed as it would be for a customer who sets a cache breakpoint at the end of each request and uses the 5-minute cache lifetime, for model tokens only. The harness's tools add no charges. Score changes are paired differences over tasks, with 95% bootstrap intervals. A change counts as inside the margin when its interval stays within 1.5 points on DRACO and 2.5 points on HLE. Anthropic set those margins before the runs. Claude Opus 5 grades the answers. Against each set's own grader, Opus 5 scored 1.9 to 2.4 points higher on DRACO, 2.2 to 2.9 points lower on HLE (Opus 5 graded 495 of the 500 questions, and the benchmark's grader graded all 500), and 1.3 to 2.0 points lower on the physics set, whose own grader also uses the expert reference solutions, in every configuration. The two graders agree on the direction of every change.
84122. **HLE:** Phan et al., "Humanity's Last Exam," arXiv:2501.14249, 2025. Expert-written questions with exact answers, graded against the reference answers. Measured on the first 500 questions, with the benchmark's own sources blocked from search, and the same setup as reference 21. Each configuration made three attempts at each question, run September 8 to 10, 2026. Claude Opus 5 compares each answer with the reference answer, with adaptive thinking on, as it is by default. The judge graded 495 of the 500 questions in every configuration, and the scores cover those 495. For the other 5, the grading request was over the judge's 1M-token limit. An attempt that hit the four-hour limit was run again, and the new attempt counts, so every configuration has all 1,500 attempts. Scoring the 5 ungraded questions as 0, as the benchmark's own scoring would, changes no finding.
84223. **Physics set:** An internal set of 70 research-level physics problems, adapted from the public CritPt benchmark: Zhu et al., "Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark," arXiv:2509.26574, 2025. Expert reviewers corrected the problem statements. Claude Opus 5 grades each answer against expert reference solutions that aren't public, so the scores can't be compared with published results. The score is the mean grade over a problem's attempts, averaged over problems. Measured on all 70 problems, four attempts per problem, run September 8 to 9, 2026. Every agent had a Python tool, a shell, and a file editor in a sandbox container with no network access, and no search or fetch tools. Otherwise the setup is that of reference 21. No score margin was set for the physics set before the runs, so the page gives its score changes with their 95% intervals and doesn't describe them as inside a margin.
742843 
743844## Next steps
744845 
Feedback