Eval graders can now judge images, so rendered slides and charts get scored on appearance.
What's wrong with this entry?
A grader of type llm pointed at a PNG, JPEG, GIF or WebP file now shows that file to the judging model as an image, so rendered slides, charts and screenshots can be graded on what they look like. A regex grader over an image always fails and says to use an llm grader instead; over other binary files it still matches ASCII sequences.
- The HTML report now explains cases that ran only one arm and therefore have no baseline to compare against.
- Non-image binaries are unchanged: regex matching against extractable ASCII.
`an \llm\ grader shows it to the judge as an image`
Strings lifted out of the shipped bundle, so the claim above can be checked against them.
Related
Other releases about the same thing. Found by shared names or similar wording; neither means one caused the other.
-
v2.1.235
plugin evalnow tells you when a nearby plugin is not being loadedBoth mention eval
-
v2.1.235
--eval-diron the eval command, with warnings that name the fixBoth mention eval
-
v2.1.235
Eval setup refuses directories it cannot vouch for
Both mention eval