{"version":"2.1.274","anchor":"self-hosted-runner-inference-token-refresh-now-retries-proac","canonical_anchor":"mcp-inference-token-refresh-rewritten-with-jittered-backoff","heading":"Inference token refresh rewritten with backoff, TTL tracking and refresh-on-error","tier":"internal","area":"Self-Hosted Runner","url":"https:\/\/changelogs.core-directive.com\/v\/2.1.274\/e\/self-hosted-runner-inference-token-refresh-now-retries-proac","release_url":"https:\/\/changelogs.core-directive.com\/v\/2.1.274","markdown":"### Inference token refresh rewritten with backoff, TTL tracking and refresh-on-error\n\nSelf-hosted runner's inference token refresh now uses jittered backoff, tracks token TTL, and refreshes immediately on a 401\/403 instead of waiting\n\n**What**\n\n- The self-hosted runner's inference-token auto-refresh logic was rewritten. It now tracks the token's time-to-live (`ttlMs`, 30 minutes by default), and retries failed refreshes with exponential backoff plus jitter (a random delay) instead of a fixed interval.\n\n- Tokens nearing expiry get a separate, shorter retry window (`expiredRetryMs`).\n\n- A new error classifier distinguishes errors worth retrying from ones that aren't: a new `RemoteConfigWithoutInferenceAuthError`, or an HTTP 401\/403, is treated as non-network and not retried early.\n\n- A new `onResultApiError` callback lets a 401\/403 returned by the model API during a child session's turn trigger an immediate, out-of-band token refresh (`refreshNow()`) instead of waiting for the next scheduled refresh.\n\n**Why** This makes token refresh more resilient to transient failures while reacting immediately to real auth failures, so a self-hosted runner recovers faster from an expired or rejected inference token instead of a child session stalling until the next scheduled refresh.\n\n- Area: Self-Hosted Runner\n- Tier: Under the hood\n- Useful: 3\/5\n- Signal: 2\/5"}