One read of Claude Documentationclaude-docs-20261006T230705Z
1 pages moved out of 262 read.
Pages moved
1
significant first
Pages read
262
in this capture
Captured
23:07 UTC
Corpus hash
625e51792eb9
corpus-hash
What this read moved
1-1 of 1third-party/claude-desktop/gateway Changed · +31 / -0 lines
### Context window and compaction
from line 312
312312
313313The model picker then shows a second entry for the model, described as **1M context window**; the standard entry has no context-size label, and the default selection is unchanged. `supports1m` is an assertion about your gateway rather than something the app can verify: if the gateway does not accept 1M-token requests for that model, requests made from the 1M picker entry fail at inference time. Only set it on models you have confirmed against your deployment. The [Models section of the configuration reference](/docs/third-party/claude-desktop/configuration#models) documents the remaining entry fields, including display labels and tier mapping.
314314
315Users with no saved selection start on the standard entry. To start them on the **1M context window** entry of the default (first) model instead, add `prefer1m: true` next to `supports1m` on that model's entry, or set [`modelPrefer1mContext`](/docs/third-party/claude-desktop/configuration#modelprefer1mcontext) to `true` when discovery populates the picker.
316
317### Context window and compaction
318
319Cowork and Code sessions keep a long conversation within the model's context window by compacting it. To compact, the session sends one request that asks the model to summarize the conversation so far, then continues from the summary. Through a gateway, a session can't read the model's real context window from the provider, so it assumes one from the model ID and the picker entry:
320
321* **Standard entry**: 1M tokens for a model the session recognizes as having a native 1M context window, as listed under [the context window behind a gateway](https://code.claude.com/docs/en/model-config#context-window-behind-a-gateway) in the Claude Code documentation, and 200K tokens for every other model ID, including IDs the session doesn't recognize. On Claude Desktop versions earlier than 2.19675.0, a session assumes 200K tokens on this entry for every model.
322* **1M context window entry**: 1M tokens, regardless of the model ID.
323
324If the gateway or the provider accepts less than the window a session assumes, requests past that limit fail. The session then compacts and retries only when the provider's too-long error reaches it unchanged. To make sessions compact earlier than the window they assume, for example when your provider serves a natively 1M model with only a 200K window, see [Set the auto-compact window](https://code.claude.com/docs/en/model-config#set-the-auto-compact-window) in the Claude Code documentation. For a model ID the session doesn't recognize, see [Correct the window for a gateway or custom model ID](https://code.claude.com/docs/en/model-config#correct-the-window-for-a-gateway-or-custom-model-id).
325
326The following table lists what the model configuration and the gateway need to provide so that sessions use the full context window and compact when they reach it, and how to check each requirement.
327
328| Requirement | What fails without it | How to check |
329| - | - | - |
330| Set `supports1m`, or mark the model in discovery (see [Models](#models)), for each model your provider serves with a 1M-token context window but a session assumes 200K tokens for. Also accept the `context-1m-2025-08-07` value in the `anthropic-beta` request header. | Sessions on such a model's standard picker entry assume a 200K-token window, even when the provider accepts 1M. A gateway that rejects the header value fails every request made from the **1M context window** entry. | From a session on the **1M context window** entry, confirm in the gateway's logs that requests carry the header value and that requests above 200K input tokens succeed. |
331| [Stream responses](https://code.claude.com/docs/en/llm-gateway-protocol#streaming) through without buffering, including the provider's SSE `ping` events, and set the response timeout on every proxy and load balancer on the route longer than your slowest complete response. | The summary request carries the whole conversation, and while the provider works on it, pings can be the only bytes on the connection. If a proxy buffers them, the session sees no bytes and gives up after about five minutes, and if a proxy's response timeout is shorter than the response takes, it ends the request with its own error, such as `504`. | Run a streaming request with `curl -N` against the hostname clients use and confirm that events print as they are generated, not in one burst at the end. |
332| Return the provider's `usage` object unchanged on every response, including `cache_creation_input_tokens` and `cache_read_input_tokens`. | The session counts the conversation's tokens from the `usage` of the latest response. With the cache fields missing it undercounts, and requests reach the provider's limit before the session compacts. | Compare the `usage` object in the gateway's upstream response log with the one the client receives for the same request. They should be identical, and where the provider supports prompt caching, `cache_read_input_tokens` is above zero after the first turn. |
333| [Forward the provider's status code and error body unchanged](https://code.claude.com/docs/en/llm-gateway-protocol#automatic-retry-and-error-forwarding) when it rejects a request. | Where the provider's too-long error is the compaction trigger, the session compacts and retries only if that error reaches it unchanged. With a rewritten error, the session shows the error and doesn't compact. | Send an oversized request through the gateway and confirm that the client receives the provider's status code and error body. |
334
335On a response that delivers only keep-alive pings, a session accepts about five minutes of pings and then waits [`inferenceStreamIdleTimeoutSec`](/docs/third-party/claude-desktop/configuration#inferencestreamidletimeoutsec) seconds more for model output (300 by default, so about ten minutes in all). Raise that key if your provider needs longer than that before its first output on a large request.
336
315337### MCP tool search
316338
317339[MCP tool search](https://code.claude.com/docs/en/mcp#scale-with-mcp-tool-search) loads MCP tool schemas on demand instead of inlining every schema into the context window. It reduces context pressure when many MCP tools are configured (sessions that otherwise compact every turn or two).
from line 363
341363**Model picker is empty or missing models.** Check your gateway's `GET /v1/models` response and your `inferenceModels` list (see [Models](#models)). When `/v1/models` is unreachable or returns an error, the picker falls back to the `inferenceModels` list; if that list is empty, so is the picker.
342364
343365**The 1M context window entry does not appear in the picker.** `supports1m` takes effect only when the entry's `name` matches the model ID the picker uses. Setting it on a bare alias (for example `sonnet`) while discovery returns full model IDs produces no match. Set `supports1m` on an entry whose `name` is the exact ID your gateway's `/v1/models` endpoint returns.
366
367**`Your context window is full` in a Code session, or `This conversation is too long to continue. Start a new session, or remove some tools to free up space.` in a Cowork session.** The app shows these when Claude Code reports `Prompt is too long · automatic compaction failed: …`, meaning a session tried to compact the conversation and the summary request failed. Despite that wording, the conversation isn't lost. After you fix the cause, the user's next message in the same session retries compaction. The error the summary request returned is in the gateway's logs as the response to the session's last `POST /v1/messages` request before the failure. Match it to a cause as follows.
368
369* **A timeout, `504`, or `524`**: a proxy or load balancer on the route closed the summary response before it finished. See the streaming requirement in [Context window and compaction](#context-window-and-compaction).
370* **`401`, `402`, or `403`**: the gateway credential expired, or the gateway's quota for the user is spent. Have the user sign in again or issue a new key, or raise the quota.
371* **`API returned an empty or malformed response`**: a proxy, firewall, or sign-in page on the route answered in place of the provider. Look for the summary request in the gateway's logs, and if it isn't there, something in front of the gateway answered it.
372* **`400` with your gateway's own error text**: the gateway rejected the summary request itself, for example on a request-size or token limit of its own. Raise or remove that limit for these requests.
373
374For `Prompt is too long` without `automatic compaction failed`, or your gateway's own request-too-large error shown in the session, check the `usage` and error-forwarding requirements in [Context window and compaction](#context-window-and-compaction) first. [Prompt is too long](https://code.claude.com/docs/en/errors#prompt-is-too-long) in the Claude Code error reference covers the other causes.
344375