Follow Discord
Sweep 25 Sep 2026 · 19:33Z Build v2.1.283 504 read Stable v2.1.274 Latest v2.1.283 Next v2.1.283 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · claude-code

Test plugins with evals changedplugin-evals

Nearest release: v2.1.283, published 2 hours before upstream edited the page. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Upstream edited this page at 25 Sep 2026 21:04 UTC, give or take a minute or two: the time comes from Anthropic’s own sitemap rather than from a commit. This site recorded the change at 25 Sep 2026 21:07 UTC.

Upstream edited
Recorded here
Lines+2added
Lines−2removed
From line 12 where the diff opens
First seen 11 Sep 2026 this site's first read of the page
Recorded edits7to this page, all time

The whole hunk

from line 12, old and new numbered
/
lines
from line 12
1212* Catch regressions when you change the plugin or a new model is released
1313* See what the plugin contributes compared with no plugin
1414 
15This page is for plugin and skill authors who have a working plugin and want to test its behavior, and for teams that gate plugin changes in CI. Its case format is separate from the `evals/evals.json` file the [skill-creator plugin](/docs/en/skills#run-evals-with-skill-creator) uses. To create a plugin, see [Create a plugin](/docs/en/plugins/create); to check a plugin's files for syntax and schema errors rather than its behavior, use [`claude plugin validate`](/docs/en/plugins/cli-reference#plugin-validate).
15This page is for plugin and skill authors who have a working plugin and want to test its behavior, and for teams that gate plugin changes in CI. For iterating on one skill inside a Claude Code conversation, the [skill-creator plugin](/docs/en/skills#run-evals-with-skill-creator) runs a similar comparison with its own `evals/evals.json` format, and neither tool reads the other's case files. To create a plugin, see [Create a plugin](/docs/en/plugins/create); to check a plugin's files for syntax and schema errors rather than its behavior, use [`claude plugin validate`](/docs/en/plugins/cli-reference#plugin-validate).
1616 
1717<Note>
1818 Every eval run and every judge grader is a real model call on your account, counted against your plan's usage or your API bill, so check the [requirements](#requirements) first. Then [create your first eval suite](#create-your-first-eval-suite), or go to [Run evals in CI](#run-evals-in-ci) if you already have one.
from line 408
408408| 130 | Interrupted. Partial results are written |
409409| 143 | Terminated, such as by a CI timeout |
410410 
411Problems writing or publishing the HTML report never change the exit code.
411The with-minus-without delta is reported but never changes the exit code, and neither do problems writing or publishing the HTML report.
412412 
413413To see why a case scored low, run it locally without `--json` so the per-run progress and grader lines print.
414414 
Feedback