Source Intelligence

DisclaimerUnofficial, and not affiliated with Anthropic. Nearly all of this is read straight out of what ships: npm bundles, captured prompts, published docs. Anthropic's own notes go in verbatim, marked as theirs. The rest is my reading, and every entry carries the strings behind it. If one looks wrong, vote it down and say why.

All of v2.1.251 Home All releases olderv2.1.250
Claude Code v2.1.251

Eval failures say whether they can be reproduced

Under the hood
Useful1 Signal2
Plugin Eval

Eval failures now say whether they're reproducible and flag runs killed before reporting.

What

When the eval harness's integrity check fails, each failure is now labelled as provokable or unprovokable, and a note is added when the child agent was killed before it could report which calls it refused, so a killed run is no longer indistinguishable from a genuine mismatch.

Evidence

the child was killed before it reported which calls it refused itself

Strings lifted out of the shipped bundle, so the claim above can be checked against them.

Related

Other releases about the same thing. Found by shared names or similar wording; neither means one caused the other.

See this entry in the whole of v2.1.251 →