reduce-latency changedtest-and-evaluate/strengthen-guardrails/reduce-latency
Nearest release: v2.1.293, published an hour before this site recorded the change. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.
Recorded here
Lines+70added
Lines−37removed
From line
1
where the diff opens
First seen
14 Aug 2026
this site's first read of the page
Recorded edits4to this page, all time
The whole hunk
from line 1, old and new numbered
/
from line 1
11---
22title: Reducing latency
33url: https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-latency
4description: Reduce Claude's response latency by choosing a faster model like Claude Haiku 4.5, trimming prompt and output tokens, and streaming responses.
4description: Reduce Claude's response latency by choosing a faster model like Claude Haiku 5.5, trimming prompt and output tokens, and streaming responses.
55---
66
77Latency refers to the time it takes for the model to process a prompt and generate an output. Latency can be influenced by various factors, such as the size of the model, the complexity of the prompt, and the underlying infrastructure supporting the model and point of interaction.
from line 29
2929
3030One of the most direct ways to reduce latency is to select the appropriate model for your use case. Anthropic offers a [range of models](https://platform.claude.com/docs/en/models/overview) with different capabilities and performance characteristics. Consider your specific requirements and choose the model that best fits your needs in terms of speed and output quality.
3131
32For speed-critical applications, **Claude Haiku 4.5** offers the fastest response times while maintaining high intelligence:
32For speed-critical applications, **Claude Haiku 5.5** offers the fastest response times while maintaining high intelligence. [Effort](https://platform.claude.com/docs/en/build-with-claude/effort) is its main control for speed and cost. See [Use effort to control thinking](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5#use-effort-to-control-thinking). The following example runs it at `low`, the cheapest and fastest level, and leaves room in `max_tokens` for thinking:
3333
3434<CodeGroup>
3535 ```bash cURL
36 # For time-sensitive applications, use Claude Haiku 4.5
36 # For time-sensitive applications, use Claude Haiku 5.5 at low effort
3737 curl https://api.anthropic.com/v1/messages \
3838 -H "x-api-key: $ANTHROPIC_API_KEY" \
3939 -H "anthropic-version: 2023-06-01" \
4040 -H "content-type: application/json" \
4141 -d '{
42 "model": "claude-haiku-4-5",
43 "max_tokens": 100,
42 "model": "claude-haiku-5-5",
43 "max_tokens": 1024,
44 "output_config": {"effort": "low"},
4445 "messages": [{"role": "user", "content": "Summarize this customer feedback in 2 sentences: [feedback text]"}]
4546 }'
4647 ```
4748
4849 ```bash CLI
49 # For time-sensitive applications, use Claude Haiku 4.5
50 # For time-sensitive applications, use Claude Haiku 5.5 at low effort
5051 ant messages create \
51 --model claude-haiku-4-5 \
52 --max-tokens 100 \
52 --model claude-haiku-5-5 \
53 --max-tokens 1024 \
54 --output-config '{effort: low}' \
5355 --message '{"role": "user", "content": "Summarize this customer feedback in 2 sentences: [feedback text]"}'
5456 ```
5557
from line 58
5658 ```python Python
5759 client = anthropic.Anthropic()
5860
59 # For time-sensitive applications, use Claude Haiku 4.5
61 # For time-sensitive applications, use Claude Haiku 5.5 at low effort
6062 message = client.messages.create(
61 model="claude-haiku-4-5",
62 max_tokens=100,
63 model="claude-haiku-5-5",
64 max_tokens=1024,
65 output_config={"effort": "low"},
6366 messages=[
6467 {
6568 "role": "user",
from line 70
6770 }
6871 ],
6972 )
70 print(message.content[0].text)
73 print(next(block.text for block in message.content if block.type == "text"))
7174 ```
7275
7376 ```typescript TypeScript
7477 const client = new Anthropic();
7578
76 // For time-sensitive applications, use Claude Haiku 4.5
79 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
7780 const message = await client.messages.create({
78 model: "claude-haiku-4-5",
79 max_tokens: 100,
81 model: "claude-haiku-5-5",
82 max_tokens: 1024,
83 output_config: { effort: "low" },
8084 messages: [
8185 {
8286 role: "user",
from line 95
9195 ```csharp C#
9296 AnthropicClient client = new();
9397
94 // For time-sensitive applications, use Claude Haiku 4.5
98 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
9599 var parameters = new MessageCreateParams
96100 {
97 Model = Model.ClaudeHaiku4_5,
98 MaxTokens = 100,
101 Model = Model.ClaudeHaiku5_5,
102 MaxTokens = 1024,
103 OutputConfig = new() { Effort = Effort.Low },
99104 Messages = [
100105 new()
101106 {
from line 110
105110 ]
106111 };
107112 var message = await client.Messages.Create(parameters);
108 message.Content[0].TryPickText(out var textBlock);
109 Console.WriteLine(textBlock?.Text);
113 foreach (var block in message.Content)
114 {
115 if (block.TryPickText(out var textBlock))
116 {
117 Console.WriteLine(textBlock.Text);
118 break;
119 }
120 }
110121 ```
111122
112123 ```go Go
113124 client := anthropic.NewClient()
114125
115 // For time-sensitive applications, use Claude Haiku 4.5
126 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
116127 message, err := client.Messages.New(context.TODO(), anthropic.MessageNewParams{
117 Model: anthropic.ModelClaudeHaiku4_5,
118 MaxTokens: 100,
128 Model: anthropic.ModelClaudeHaiku5_5,
129 MaxTokens: 1024,
130 OutputConfig: anthropic.OutputConfigParam{
131 Effort: anthropic.OutputConfigEffortLow,
132 },
119133 Messages: []anthropic.MessageParam{
120134 anthropic.NewUserMessage(anthropic.NewTextBlock("Summarize this customer feedback in 2 sentences: [feedback text]")),
121135 },
from line 137
123137 if err != nil {
124138 log.Fatal(err)
125139 }
126 fmt.Println(message.Content[0].Text)
140 for _, block := range message.Content {
141 if textBlock, ok := block.AsAny().(anthropic.TextBlock); ok {
142 fmt.Println(textBlock.Text)
143 break
144 }
145 }
127146 ```
128147
129148 ```java Java
130149 AnthropicClient client = AnthropicOkHttpClient.fromEnv();
131150
132 // For time-sensitive applications, use Claude Haiku 4.5
151 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
133152 MessageCreateParams params = MessageCreateParams.builder()
134 .model(Model.CLAUDE_HAIKU_4_5)
135 .maxTokens(100L)
153 .model(Model.CLAUDE_HAIKU_5_5)
154 .maxTokens(1024L)
155 .outputConfig(OutputConfig.builder()
156 .effort(OutputConfig.Effort.LOW)
157 .build())
136158 .addUserMessage("Summarize this customer feedback in 2 sentences: [feedback text]")
137159 .build();
138160 Message message = client.messages().create(params);
139 IO.println(message.content().get(0).text().map(TextBlock::text).orElse(""));
161 IO.println(message.content().stream()
162 .flatMap(block -> block.text().stream())
163 .map(TextBlock::text)
164 .findFirst()
165 .orElse(""));
140166 ```
141167
142168 ```php PHP
143169 $client = new Client();
144170
145 // For time-sensitive applications, use Claude Haiku 4.5
171 // For time-sensitive applications, use Claude Haiku 5.5 at low effort
146172 $message = $client->messages->create(
147 maxTokens: 100,
173 maxTokens: 1024,
148174 messages: [['role' => 'user', 'content' => 'Summarize this customer feedback in 2 sentences: [feedback text]']],
149 model: 'claude-haiku-4-5',
175 model: 'claude-haiku-5-5',
176 outputConfig: ['effort' => 'low'],
150177 );
151 echo $message->content[0]->text;
178 foreach ($message->content as $block) {
179 if ($block->type === 'text') {
180 echo $block->text;
181 break;
182 }
183 }
152184 ```
153185
154186 ```ruby Ruby
155187 client = Anthropic::Client.new
156188
157 # For time-sensitive applications, use Claude Haiku 4.5
189 # For time-sensitive applications, use Claude Haiku 5.5 at low effort
158190 message = client.messages.create(
159 model: "claude-haiku-4-5",
160 max_tokens: 100,
191 model: "claude-haiku-5-5",
192 max_tokens: 1024,
193 output_config: { effort: :low },
161194 messages: [{ role: "user", content: "Summarize this customer feedback in 2 sentences: [feedback text]" }]
162195 )
163 puts message.content.first.text
196 puts message.content.find { |block| block.type == :text }&.text
164197 ```
165198</CodeGroup>
166199
from line 222
189222
190223 tokens, the response will be cut off, perhaps mid-sentence or mid-word, so this is a blunt technique that might require post-processing and is usually most appropriate for multiple choice or short answer responses where the answer comes right at the beginning.
191224 </Note>
192* **Experiment with temperature:** The `temperature` [parameter](https://platform.claude.com/docs/en/api/messages/create) controls the randomness of the output. Lower values (for example, 0.2) can sometimes lead to more focused and shorter responses, while higher values (for example, 0.8) might result in more diverse but potentially longer outputs.
225* **Experiment with temperature:** The `temperature` [parameter](https://platform.claude.com/docs/en/api/messages/create) controls the randomness of the output. Lower values (for example, 0.2) can sometimes lead to more focused and shorter responses, while higher values (for example, 0.8) might result in more diverse but potentially longer outputs. Claude Haiku 5.5 accepts only the default `temperature` and returns a 400 error for any other value, so lower its [effort](https://platform.claude.com/docs/en/build-with-claude/effort) instead.
193226
194227Finding the right balance among prompt clarity, output quality, and token count might require some experimentation.
195228
No line in this hunk matches that.