Test plugins with evals changedplugin-evals
Nearest release: v2.1.293, published under an hour after upstream edited the page. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.
Upstream edited this page at 7 Oct 2026 17:14 UTC, give or take a minute or two: the time comes from Anthropic’s own sitemap rather than from a commit. This site recorded the change at 7 Oct 2026 17:37 UTC.
Upstream edited
Recorded here
Lines+1added
Lines−1removed
From line
106
where the diff opens
First seen
11 Sep 2026
this site's first read of the page
Recorded edits16to this page, all time
The whole hunk
from line 106, old and new numbered
/
from line 106
106106
107107 The most common first finding is a `Δ` near zero with the case's `tool_used: Skill` grader failing, which means Claude isn't choosing your skill on natural phrasing. Adjust the skill's [`description`](/docs/en/skills#frontmatter-reference), run `claude plugin eval .` again, and compare.
108108
109 To iterate on one case cheaply, run a single arm once. A single run is noisy, so confirm any change at the default three runs before you trust it. With one arm the table shows `SCORE` and `PASS%` columns instead of `WITH`, `W/OUT`, and `Δ`:
109 To iterate on one case with fewer runs, run a single arm once. A single run is noisy, so confirm any change at the default three runs before you trust it. With one arm the table shows `SCORE` and `PASS%` columns instead of `WITH`, `W/OUT`, and `Δ`:
110110
111111 ```bash theme={null}
112112 claude plugin eval . --case <case-name> --runs 1 --ablation none
No line in this hunk matches that.