What probably matters to youSection of the release
Unclear What the per-upstream models list does is not stated.
What
The apps gateway sits between Claude Code and the model providers it forwards requests to, which are called upstreams. Its first-byte timeout setting, timeouts.upstream_ttfb_ms, controls how long the gateway waits for the first piece of a reply. It has been reworked, and upstreams gained new options.
timeouts.upstream_ttfb_ms is now optional in the gateway config. The schema no longer fills in a default. A default of 120000 milliseconds is still applied, but the gateway now records whether the operator wrote the value themselves.
The timeout, capped, is now passed on the upstream call for streaming requests and count_tokens requests. A request that times out returns an error message that names timeouts.upstream_ttfb_ms.
When an operator has set the value, the gateway logs an info message at startup. The message says the setting now also applies to streaming and count_tokens requests on non-anthropic upstreams. Before, the timeout applied only to provider: anthropic upstreams.
The anthropic, bedrock, anthropicAws, anthropicGoogleCloud, vertex and foundry upstreams now accept a models list.
count_tokens (a request that asks how many tokens a prompt uses) on Bedrock upstreams is now handled. Before, it always returned a 501 "not supported" error.
Why
If you run the gateway with non-Anthropic upstreams and have set timeouts.upstream_ttfb_ms, the timeout now also applies to those upstreams. Slow first replies there can therefore fail where they did not before. Bedrock users can now get token counts through the gateway.
Something disagreesSomething we can check disagrees with this entry, or the writer said they could not settle it.
The writer flagged doubtWhat the per-upstream `models` list does is not stated.
Anthropic's release notes agreeAdded support for timeouts.upstream_ttfb_ms on the Claude apps gateway's Bedrock, Vertex, Foundry and other cloud upstreams: a value you set…