What
When Claude Code sends a request, it can store part of the conversation in a prompt cache so later requests are faster and cheaper. Cached entries last either 5 minutes or 1 hour. The API's usage report can now say how many cached tokens were written at each lifetime, and Claude Code reads that.
- The usage format now accepts a
cache_creationobject holdingephemeral_5m_input_tokensandephemeral_1h_input_tokens. - The cache fields and
server_tool_usemay now be missing or empty without causing a problem. - When usage from several responses is combined,
cache_creationis kept, andcache_creation_input_tokensis worked out from it. - A new check reports whether the cache lifetime was
1hor5mfrom these counts. - The fallback that picks the cache lifetime for the main conversation now calls a function. Before, it chose
1hfor the main thread and5motherwise. - Each usage record now carries
cache_creation_5m_input_tokensandcache_creation_1h_input_tokens, each 0 when not reported.
Why
The cache lifetime is now taken from what the API reported rather than assumed. That feeds Claude Code's decisions about cache misses and keeping the cache alive, and it lets the cost of cache writes be split by lifetime.
Something disagreesSomething we can check disagrees with this entry, or the writer said they could not settle it.
The writer flagged doubt
What the new fallback rule for the main conversation returns is not clear.