Group of 5 You'll notice
claude plugin eval is now generally available everywhere but requires trusting a plugin directory before it runs, with clearer help text and init delegation
What
claude plugin evalnow reports itself as generally available for every client type (Bedrock/Vertex/Foundry, gateways, telemetry-disabled clients, CI), unless a server-side flag turns it off — in which case the CLI now says the command exists but is switched off, rather than claiming it doesn't exist.- Before loading a target,
claude plugin evalnow runs a trust check that must resolve totrusted: true— via an already-installed plugin, an existing trust marker, the--trust-pluginflag, or an interactive prompt — or the command errors out and exits without loading or running anything. - The command's help text was expanded to explain that it loads and runs a plugin's eval suite on your machine, that sandboxing limits but doesn't guarantee safety, and that the first run against an untrusted plugin directory will prompt for confirmation (use
--trust-pluginto pre-answer that, e.g. for CI). claude plugin eval init, when run without a terminal but detected as launched from inside another Claude Code session, can now delegate its authoring interview to that parent session instead of failing with "no TTY available". In that delegated interview flow, the interviewer must now explicitly ask whether you trust the plugin directory before piloting eval cases, and only adds--trust-pluginon an explicit yes; case files are still written on a no/no-answer, just not piloted.
Why
These changes make claude plugin eval a supported, documented feature everywhere while making sure a plugin's eval code — which runs on your machine — isn't executed against an untrusted directory without your explicit confirmation.
Names in the bundleclaude plugin eval
claude plugin eval
Claude Code changelog modified, high confidence
* Added `claude plugin eval`: run a plugin's eval suite against Claude Code and get scored, reproducible results (JSON + HTML report); see `claude plugin eval --help`see the edit
claude plugin eval
Error reference modified, high confidence
You ran [`claude plugin eval`](/docs/en/plugin-evals) or `claude plugin eval init` and it exited 1 with one of these messages before doing anything:see the edit
claude plugin eval
Plugins reference modified, high confidence
| `experimental.evals` | string\|array | Directory below the plugin root that holds the plugin's [eval cases](/docs/en/plugin-evals#use-a-different-eval-directory), when it isn't the default `evals/`. `claude plugin eval --eval-dir` overri…see the edit
claude plugin eval
Test plugins with evals modified, high confidence
You don't have to write the suite manually. `claude plugin eval init` asks you about your plugin, proposes the cases and graders, tries them, and writes the files. You can also ask Claude to do the same from a session you already have open.see the edit
claude plugin eval
Create and distribute a plugin marketplace modified, medium confidence
Test your marketplace before sharing. Validation checks file structure; to test whether a plugin changes what Claude does on realistic prompts, run its eval suite with [`claude plugin eval`](/docs/en/plugin-evals) before you publish a new …see the edit
claude plugin eval
Create plugins modified, medium confidence
Trying the plugin with `--plugin-dir` tells you it can work. To find out how often Claude actually reaches for it and gets the right result, run it against a set of test prompts with [`claude plugin eval`](/docs/en/plugin-evals). Each prom…see the edit
claude plugin eval
Extend Claude with skills modified, medium confidence
Two tools automate that comparison. For a skill that ships in a [plugin](/docs/en/plugins), [`claude plugin eval`](/docs/en/plugin-evals) runs each prompt in an isolated session with and without the plugin, scores it with graders you defin…see the edit
The entry above is what we published on the day. These lines were added later, as Anthropic's own pages caught up, and they sit beside the original rather than replacing it.
Confirmed since
Anthropic's documentation has since written up claude plugin eval, on What's new.
**`claude plugin eval`**: run your plugin against a suite of test cases, score the results, and compare against a no-plugin baseline. `claude plugin eval init` drafts the cases and graders for you.whats-new/index see the edit
Two sources agreeTwo things we can check say the same as this entry.
Anthropic's documentation agrees
claude plugin eval on Claude Code changelog
Anthropic's release notes agree
Added claude plugin eval: run a plugin's eval suite against Claude Code and get scored, reproducible results (JSON + HTML report); see…