task-budgets
build-with-claude/task-budgets
History
build-with-claude/task-budgets First recorded · 633 lines, first recorded
## Compatibility ## When to use task budgets ## Setting a task budget ## How the budget countdown works ### Worked example: budget counting across turns ### Carrying a budget across compaction with `remaining` ## Changing the budget mid-conversation ## Task budgets are advisory, not enforced ## Choosing a budget ### Measure your current usage ## Interaction with other parameters ## Feature support ## Next steps
The first capture of this source. The page was already there, and this is what it said.
---
title: Task budgets
url: https://platform.claude.com/docs/en/build-with-claude/task-budgets
description: Give Claude an advisory token budget for the full agentic loop to help the model self-regulate on long agentic tasks.
---
## Compatibility
- Status: Beta
- [Beta header](https://platform.claude.com/docs/en/api/beta-headers): `task-budgets-2026-03-13`
- Supported models: `claude-fable-5`, `claude-mythos-5`, `claude-opus-5`, `claude-opus-4-8`, `claude-opus-4-7`
Task budgets let you tell Claude how many tokens it has for a full agentic loop, including thinking, tool calls, tool results, and output. The model sees a running countdown and uses it to prioritize work and finish gracefully as the budget is consumed.
## When to use task budgets
Task budgets work best for agentic workflows where Claude makes multiple tool calls and decisions before finalizing its output to await the next human response. Use them when:
* You want Claude to self-regulate token spend on long-horizon tasks.
* You have a predictable per-task cost or latency ceiling to enforce.
* You want the model to finish gracefully (summarize findings, report progress) as it approaches the budget rather than cutting off mid-action.
Task budgets complement the [effort parameter](https://platform.claude.com/docs/en/build-with-claude/effort): effort controls how thoroughly Claude reasons about each step, while task budgets cap the total work Claude can do across an agentic loop.
## Setting a task budget
Add `task_budget` to `output_config` and include the beta header:
<CodeGroup>
```bash cURL
curl https://api.anthropic.com/v1/messages \
-N \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "anthropic-beta: task-budgets-2026-03-13" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-5",
"max_tokens": 128000,
"stream": true,
"messages": [{
"role": "user",
"content": "Review the codebase and propose a refactor plan."
}],
"output_config": {
"effort": "high",
"task_budget": {"type": "tokens", "total": 64000}
}
}'
```
```bash CLI
ant beta:messages create --beta task-budgets-2026-03-13 \
--stream --format jsonl <<'YAML' | jq 'select(.type == "message_delta").usage'
model: claude-opus-5
max_tokens: 128000
messages:
- role: user
content: Review the codebase and propose a refactor plan.
output_config:
effort: high
task_budget:
type: tokens
total: 64000
YAML
```
```python Python
client = anthropic.Anthropic()
with client.beta.messages.stream(
model="claude-opus-5",
max_tokens=128000,
output_config={
"effort": "high",
"task_budget": {"type": "tokens", "total": 64000},
},
messages=[
{"role": "user", "content": "Review the codebase and propose a refactor plan."}
],
betas=["task-budgets-2026-03-13"],
) as stream:
response = stream.get_final_message()
print(response.usage)
```
```typescript TypeScript
const client = new Anthropic();
const stream = client.beta.messages.stream({
model: "claude-opus-5",
max_tokens: 128000,
output_config: {
effort: "high",
task_budget: { type: "tokens", total: 64000 }
},
messages: [{ role: "user", content: "Review the codebase and propose a refactor plan." }],
betas: ["task-budgets-2026-03-13"]
});
const response = await stream.finalMessage();
console.log(response.usage);
```
```csharp C#
var client = new AnthropicClient();
var responseUpdates = client.Beta.Messages.CreateStreaming(new MessageCreateParams
{
Model = Messages::Model.ClaudeOpus5,
MaxTokens = 128000,
Messages = [new() { Role = Role.User, Content = "Review the codebase and propose a refactor plan." }],
OutputConfig = new BetaOutputConfig
{
Effort = Effort.High,
TaskBudget = new BetaTokenTaskBudget { Total = 64000 },
},
Betas = ["task-budgets-2026-03-13"],
});
var response = await responseUpdates.Aggregate();
Console.WriteLine(response.Usage);
```
```go Go
client := anthropic.NewClient()
stream := client.Beta.Messages.NewStreaming(context.TODO(), anthropic.BetaMessageNewParams{
Model: anthropic.ModelClaudeOpus5,
MaxTokens: 128000,
Betas: []anthropic.AnthropicBeta{"task-budgets-2026-03-13"},
Messages: []anthropic.BetaMessageParam{{
Role: anthropic.BetaMessageParamRoleUser,
Content: []anthropic.BetaContentBlockParamUnion{{
OfText: &anthropic.BetaTextBlockParam{Text: "Review the codebase and propose a refactor plan."},
}},
}},
OutputConfig: anthropic.BetaOutputConfigParam{
Effort: anthropic.BetaOutputConfigEffortHigh,
TaskBudget: anthropic.BetaTokenTaskBudgetParam{
Total: 64000,
},
},
})
message := anthropic.BetaMessage{}
for stream.Next() {
event := stream.Current()
if err := message.Accumulate(event); err != nil {
panic(err)
}
}
if stream.Err() != nil {
panic(stream.Err())
}
fmt.Printf("Usage: input_tokens=%d, output_tokens=%d\n", message.Usage.InputTokens, message.Usage.OutputTokens)
```
```java Java
AnthropicClient client = AnthropicOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model(Model.CLAUDE_OPUS_5)
.maxTokens(128000L)
.addUserMessage("Review the codebase and propose a refactor plan.")
.outputConfig(BetaOutputConfig.builder()
.effort(BetaOutputConfig.Effort.HIGH)
.taskBudget(BetaTokenTaskBudget.builder().total(64000L).build())
.build())
.addBeta("task-budgets-2026-03-13")
.build();
BetaMessageAccumulator accumulator = BetaMessageAccumulator.create();
try (StreamResponse<BetaRawMessageStreamEvent> stream =
client.beta().messages().createStreaming(params)) {
stream.stream().forEach(accumulator::accumulate);
}
BetaMessage response = accumulator.message();
IO.println(response.usage());
```
```php PHP
use Anthropic\Beta\Messages\BetaRawMessageDeltaEvent;
$client = new Client();
$stream = $client->beta->messages->createStream(
model: 'claude-opus-5',
maxTokens: 128000,
messages: [
['role' => 'user', 'content' => 'Review the codebase and propose a refactor plan.'],
],
outputConfig: [
'effort' => 'high',
'taskBudget' => ['type' => 'tokens', 'total' => 64000],
],
betas: ['task-budgets-2026-03-13'],
);
// The final message_delta event carries the cumulative token usage for the request.
$usage = null;
foreach ($stream as $event) {
if ($event instanceof BetaRawMessageDeltaEvent) {
$usage = $event->usage;
}
}
echo $usage;
```
```ruby Ruby
client = Anthropic::Client.new
stream = client.beta.messages.stream(
model: "claude-opus-5",
max_tokens: 128_000,
messages: [
{ role: "user", content: "Review the codebase and propose a refactor plan." }
],
output_config: {
effort: :high,
task_budget: { type: :tokens, total: 64_000 }
},
betas: ["task-budgets-2026-03-13"]
)
response = stream.accumulated_message
puts response.usage
```
</CodeGroup>
The `task_budget` object has three fields:
* `type`: always `"tokens"`.
* `total`: the number of tokens Claude can spend across the agentic loop, including thinking, tool calls, tool results, and output.
* `remaining` (optional): the budget remainder carried over from a prior request. Defaults to `total` when omitted.
## How the budget countdown works
Claude sees a budget-countdown marker injected server-side throughout the conversation. The marker shows how many tokens remain in the current agentic loop and updates as the model generates thinking, tool calls, and output, and as it processes tool results. Claude uses this signal to pace itself and finish gracefully as the budget is consumed.
<Note>
**The countdown is visible only to the model.** API responses do not include a remaining-budget field: there is no `task_budget` information in the response `usage` object, and SDKs have no accessor for it. To track spend client-side, sum token usage across the requests in your loop as shown in [Measure your current usage](https://platform.claude.com/docs/en/build-with-claude/task-budgets#measure-your-current-usage), or pass your own figure forward with `remaining` when [carrying a budget across compaction](https://platform.claude.com/docs/en/build-with-claude/task-budgets#carrying-a-budget-across-compaction-with-remaining).
</Note>
<Warning>
**The countdown reflects tokens Claude has processed in the current agentic loop, not tokens you resend between turns.** If your client sends the full conversation history on every follow-up request, your client-side token count may differ from the budget Claude is tracking. If you also decrement `remaining` while resending full history, the model sees an under-reported budget and the countdown drops faster than it should, causing Claude to wrap up earlier than the budget actually allows. Set a generous budget and let the model self-regulate against the countdown rather than trying to mirror it client-side.
</Warning>
### Worked example: budget counting across turns
The task budget counts what Claude **sees** (thinking, tool calls and results, and text), not what's in your request payload. In an agentic loop your client resends the full conversation on every request, so the payload grows turn over turn, but the budget only decrements by the tokens Claude sees this turn.
Consider a loop with `task_budget: {type: "tokens", total: 100000}` and a single `bash` tool.
**Turn 1.** You send the initial request:
```json
{
"messages": [
{ "role": "user", "content": "Audit this repo for security issues and report findings." }
]
}
```
Claude thinks, then emits a tool call and stops with `stop_reason: "tool_use"`:
```json
{
"role": "assistant",
"content": [
{
"type": "thinking",
"thinking": "I'll start by listing dependencies to look for known-vulnerable packages..."
},
{
"type": "tool_use",
"id": "toolu_01",
"name": "bash",
"input": { "command": "cat package.json && npm audit --json" }
}
]
}
```
Suppose this assistant turn (thinking plus the tool call) totals 5,000 generated tokens. The countdown Claude saw during generation ended near `remaining` ≈ 95,000.
**Turn 2.** Your client runs the tool, then resends the full history with the tool result appended:
```json
{
"messages": [
{ "role": "user", "content": "Audit this repo for security issues and report findings." },
{
"role": "assistant",
"content": [
Cut at 300 lines.