Claude Code now reads runtime.default_output_tokens and runtime.default_effort from the server's model catalog, which is the list of models the server offers. It records its own values and compares the two, noting these things:
- The default output tokens (the usual length limit for a reply), and whether that default was requested.
- The default effort, and whether it is pinned above what the server offers.
- Whether the server supplied each value, and whether it matches the local one.
The comparison report now includes default_output, plus default_effort_old and default_effort_new when the effort differs.
Models can also be marked rejects_disabled_thinking. For those models, turning thinking off is refused with "Thinking can't be turned off for this model". There is also new handling for the showThinkingSummaries setting, and a check for first-party entitlement. Context-window lookups now work from a single resolved model record.
If you try to turn thinking off on a model that requires it, you now get a clear refusal. The comparison reports show where Claude Code's defaults and the server's disagree.