Source Intelligence

DisclaimerUnofficial, and not affiliated with Anthropic. Nearly all of this is read straight out of what ships: npm bundles, captured prompts, published docs. Anthropic's own notes go in verbatim, marked as theirs. The rest is my reading, and every entry carries the strings behind it. If one looks wrong, vote it down and say why.

All of v2.1.234 Home All releases olderv2.1.233 v2.1.235newer
Claude Code v2.1.234

Plugin eval harness can grade images

Use it now
Useful2 Signal3
Plugin Evals

Eval graders can now judge images, so rendered slides and charts get scored on appearance.

llmregex
What

A grader of type llm pointed at a PNG, JPEG, GIF or WebP file now shows that file to the judging model as an image, so rendered slides, charts and screenshots can be graded on what they look like. A regex grader over an image always fails and says to use an llm grader instead; over other binary files it still matches ASCII sequences.

Details
  • The HTML report now explains cases that ran only one arm and therefore have no baseline to compare against.
  • Non-image binaries are unchanged: regex matching against extractable ASCII.
Evidence

`an \llm\ grader shows it to the judge as an image`

Strings lifted out of the shipped bundle, so the claim above can be checked against them.

Related

Other releases about the same thing. Found by shared names or similar wording; neither means one caused the other.

See this entry in the whole of v2.1.234 →