Follow Discord
Sweep 08 Oct 2026 · 18:53Z Build v2.1.295 516 read Stable v2.1.286 Latest v2.1.295 Next v2.1.295 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · claude-code

Test plugins with evals changedplugin-evals

Nearest release: v2.1.293, published under an hour after upstream edited the page. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Upstream edited this page at 7 Oct 2026 17:14 UTC, give or take a minute or two: the time comes from Anthropic’s own sitemap rather than from a commit. This site recorded the change at 7 Oct 2026 17:37 UTC.

Upstream edited
Recorded here
Lines+1added
Lines−1removed
From line 106 where the diff opens
First seen 11 Sep 2026 this site's first read of the page
Recorded edits16to this page, all time

The whole hunk

from line 106, old and new numbered
/
lines
from line 106
106106 
107107 The most common first finding is a `Δ` near zero with the case's `tool_used: Skill` grader failing, which means Claude isn't choosing your skill on natural phrasing. Adjust the skill's [`description`](/docs/en/skills#frontmatter-reference), run `claude plugin eval .` again, and compare.
108108 
109 To iterate on one case cheaply, run a single arm once. A single run is noisy, so confirm any change at the default three runs before you trust it. With one arm the table shows `SCORE` and `PASS%` columns instead of `WITH`, `W/OUT`, and `Δ`:
109 To iterate on one case with fewer runs, run a single arm once. A single run is noisy, so confirm any change at the default three runs before you trust it. With one arm the table shows `SCORE` and `PASS%` columns instead of `WITH`, `W/OUT`, and `Δ`:
110110 
111111 ```bash theme={null}
112112 claude plugin eval . --case <case-name> --runs 1 --ablation none
Feedback