Count tokens in a Message changedapi/messages/count_tokens
Nearest release: v2.1.296, published 4 hours before this site recorded the change. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.
Recorded here
Lines+9added
Lines−13removed
From line
41
where the diff opens
First seen
14 Aug 2026
this site's first read of the page
Recorded edits16to this page, all time
The whole hunk
from line 41, old and new numbered
/
from line 41
4141
4242 Each input message must be an object with a `role` and `content`. You can specify a single `user`-role message, or you can include multiple `user` and `assistant` messages.
4343
44 If the final message uses the `assistant` role, the response content will continue immediately from the content in that message. This can be used to constrain part of the model's response.
44 If the final message uses the `assistant` role, the response content will continue immediately from the content in that message. This can be used to constrain part of the model's response. This is called prefill. On models that don't support prefill, creating a message that ends with a partial `assistant` response returns a 400 error. See [Prefill not supported](https://platform.claude.com/docs/en/api/errors#prefill-not-supported).
4545
4646 Example with a single `user` message:
4747
from line 59
5959 ]
6060 ```
6161
62 Example with a partially-filled response from Claude:
62 Example with a partially-filled response from Claude, for models that support prefill:
6363
6464 ```json
6565 [
from line 1121
11211121
11221122 - `"claude-haiku-4-5"`
11231123
1124 Fastest model with near-frontier intelligence
1125
11261124 - `"claude-haiku-4-5-20251001"`
11271125
1128 Fastest model with near-frontier intelligence
1129
11301126 - `"claude-opus-4-5"`
11311127
11321128 Powerful intelligence for long-running agents and coding
from line 1209
12131209
12141210- `thinking: optional ThinkingConfigParam`
12151211
1216 Configuration for enabling Claude's extended thinking.
1212 Configuration for Claude's thinking.
12171213
1218 When enabled, responses include `thinking` content blocks showing Claude's thinking process before the final answer. Requires a minimum budget of 1,024 tokens and counts towards your `max_tokens` limit.
1214 With `{"type": "adaptive"}`, Claude decides when and how much to think. With `{"type": "enabled"}` (manual extended thinking), you set a `budget_tokens` of at least 1,024. Thinking tokens count toward your `max_tokens` limit.
12191215
1220 See [extended thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) for details.
1216 Which `type` values are accepted, and what happens when you omit `thinking`, depend on the model. See [thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#configuring-thinking) for each model's behavior.
12211217
12221218 - `ThinkingConfigEnabled object`
12231219
from line 1231
12351231
12361232 - `display: optional "summarized" or "omitted" or null`
12371233
1238 Controls how thinking content appears in the response. When set to `summarized`, thinking is returned normally. When set to `omitted`, thinking content is redacted but a signature is returned for multi-turn continuity. Defaults to `summarized`.
1234 Controls how thinking content appears in the response. When set to `summarized`, thinking is returned normally. When set to `omitted`, thinking content is redacted but a signature is returned for multi-turn continuity. The default depends on the model; see [Controlling thinking display](https://platform.claude.com/docs/en/build-with-claude/thinking#controlling-thinking-display).
12391235
12401236 - `"summarized"`
12411237
from line 1251
12551251
12561252 - `display: optional "summarized" or "omitted" or null`
12571253
1258 Controls how thinking content appears in the response. When set to `summarized`, thinking is returned normally. When set to `omitted`, thinking content is redacted but a signature is returned for multi-turn continuity. Defaults to `summarized`.
1254 Controls how thinking content appears in the response. When set to `summarized`, thinking is returned normally. When set to `omitted`, thinking content is redacted but a signature is returned for multi-turn continuity. The default depends on the model; see [Controlling thinking display](https://platform.claude.com/docs/en/build-with-claude/thinking#controlling-thinking-display).
12591255
12601256 - `"summarized"`
12611257
No line in this hunk matches that.