Test plugins with evals changedplugin-evals
Nearest release: v2.1.282, published 7 hours before upstream edited the page. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.
Upstream edited this page at 24 Sep 2026 23:46 UTC, give or take a minute or two: the time comes from Anthropic’s own sitemap rather than from a commit. This site recorded the change at 25 Sep 2026 00:07 UTC.
Upstream edited
Recorded here
Lines+134added
Lines−75removed
From line
2
where the diff opens
First seen
11 Sep 2026
this site's first read of the page
Recorded edits7to this page, all time
The whole hunk
from line 2, old and new numbered
/
from line 2
22
33> Write eval cases for your Claude Code plugin, run them with claude plugin eval, grade the results, compare against a no-plugin baseline, and gate CI on the score.
44
5`claude plugin eval` runs your [plugin](/docs/en/plugins) against a suite of test cases and scores the results. Each case is a realistic prompt plus one or more graders. A grader is a pass/fail check on what Claude produced, such as a regex over the reply, whether a particular tool was called, or a rubric that a second model judges the reply against.
5The `claude plugin eval` shell command runs your [plugin](/docs/en/plugins/overview) against a suite of test cases and scores the results. Each case is a realistic prompt plus one or more graders. A grader is a pass/fail check on what Claude produced, such as a regex over the reply, whether a particular tool was called, or a rubric that a second model judges the reply against.
66
77You don't have to write the suite manually. `claude plugin eval init` asks you about your plugin, proposes the cases and graders, tries them, and writes the files. You can also ask Claude to do the same from a session you already have open.
88
9Use evals to measure how reliably your plugin steers Claude to the right outcome, to catch regressions when you change the plugin or a new model ships, and to see what the plugin contributes compared with no plugin at all.
9Use evals to:
1010
11This page is for plugin and skill authors who have a working plugin and want to test its behavior, and for teams that gate plugin changes in CI. Its case format is separate from the `evals/evals.json` file the [skill-creator plugin](/docs/en/skills#run-evals-with-skill-creator) uses. To create a plugin, see [Create plugins](/docs/en/plugins); to check a plugin's files for syntax and schema errors rather than its behavior, use [`claude plugin validate`](/docs/en/plugins-reference#plugin-validate).
11* Measure how reliably your plugin leads Claude to produce the right outcome
12* Catch regressions when you change the plugin or a new model is released
13* See what the plugin contributes compared with no plugin
1214
15This page is for plugin and skill authors who have a working plugin and want to test its behavior, and for teams that gate plugin changes in CI. Its case format is separate from the `evals/evals.json` file the [skill-creator plugin](/docs/en/skills#run-evals-with-skill-creator) uses. To create a plugin, see [Create a plugin](/docs/en/plugins/create); to check a plugin's files for syntax and schema errors rather than its behavior, use [`claude plugin validate`](/docs/en/plugins/cli-reference#plugin-validate).
16
1317<Note>
1418 Every eval run and every judge grader is a real model call on your account, counted against your plan's usage or your API bill, so check the [requirements](#requirements) first. Then [create your first eval suite](#create-your-first-eval-suite), or go to [Run evals in CI](#run-evals-in-ci) if you already have one.
1519</Note>
from line 23
1923To run plugin evals you need:
2024
2125* Claude Code v2.1.269 or later. Run `claude --version` to check and `claude update` to upgrade.
22* A plugin directory with a `plugin.json` or `.claude-plugin/plugin.json` manifest, or a [skills-directory plugin](/docs/en/plugins-reference#skills-directory-plugins).
26* A plugin directory with a `plugin.json` or `.claude-plugin/plugin.json` manifest, or a [skills-directory plugin](/docs/en/plugins/loading#plugins-shared-through-a-repository).
2327* The same authentication and model provider your normal Claude Code sessions use. Eval runs, judge-scored graders, and `claude plugin eval init` call the model with your credentials, so they count against your plan's usage limits or your API bill. When the command reports a cost, the figure is a [list-price estimate](/docs/en/costs) of those calls.
2428
2529## How an eval run works
from line 36
3236
3337### How a case is scored
3438
35One run of a non-deterministic agent tells you little, so each case runs three times by default. A run's score is the fraction of its graders that passed, weighted if you set weights, and the case's score is the mean across its runs. A case passes when its score meets the [`--threshold`](#command-options), `1.0` by default. In model calls, a suite makes roughly cases × runs agent runs with the plugin and as many again for the [no-plugin baseline](#the-no-plugin-baseline), plus three short judge calls per `llm` or `baseline` grader per run.
39One run of a non-deterministic agent tells you little, so each case runs three times by default. A run's score is the fraction of its graders that passed, weighted if you set weights, and the case's score is the mean across its runs. A case passes when its score meets the [`--threshold`](#command-options), `1.0` by default.
3640
41In model calls, a suite makes roughly cases × runs agent runs with the plugin and the same number again for the [no-plugin baseline](#the-no-plugin-baseline), plus three short judge calls per `llm` or `baseline` grader per run.
42
3743### The no-plugin baseline
3844
39A high score on its own doesn't tell you the plugin helped, because Claude might do as well without it. To separate the two, each case's runs are repeated with no plugin loaded by default, and you get two scores, `WITH` and `W/OUT`. Their difference, `Δ`, is what the plugin contributed. If a case scores 1.0 both with and without the plugin, the plugin isn't what made it pass. The two sets of runs are called the with-arm and the without-arm; [Compare against a no-plugin baseline](#compare-against-a-no-plugin-baseline) covers how graders are scored across them and how to turn the baseline off.
45A high score on its own doesn't tell you the plugin helped, because Claude might do as well without it. To separate the two, each case's runs are repeated with no plugin loaded by default, and you get two scores, `WITH` and `W/OUT`. Their difference, `Δ`, is what the plugin contributed. If a case scores 1.0 both with and without the plugin, the plugin isn't what made it pass.
4046
47The two sets of runs are called the with-arm and the without-arm; [Score against the no-plugin baseline](#compare-against-a-no-plugin-baseline) covers how graders are scored across them and how to turn the baseline off.
48
4149## Create your first eval suite
4250
4351This walkthrough writes one case for your own plugin, runs it, and reads the result. Before you start, make sure you have:
from line 62
5462 claude plugin eval init
5563 ```
5664
57 If Claude Code doesn't already trust this directory it first asks `Trust this plugin directory?`; answer `y`. An interactive Claude Code session then opens. Claude reads your plugin and asks you what a good result looks like, proposes prompts that should and shouldn't trigger the plugin, designs graders for each, pilots them once to check they behave, and writes one case directory per prompt under `evals/`, each named after its prompt. When Claude tells you the suite is ready, exit that session with `/exit` or Ctrl+D to return to your shell.
65 If Claude Code doesn't already trust this directory it first asks `Trust this plugin directory?`; answer `y`.
5866
67 An interactive Claude Code session then opens. Claude reads your plugin and asks you what a good result looks like, proposes prompts that should and shouldn't trigger the plugin, designs graders for each, runs them once as a trial to check they behave, and writes one case directory per prompt under `evals/`, each named after its prompt.
68
69 When Claude tells you the suite is ready, exit that session with `/exit` or Ctrl+D to return to your shell.
70
5971 If you already have a Claude Code session open at the plugin root, you can instead ask Claude there to run `claude plugin eval init`. Claude runs the command and then asks you the same questions in that conversation.
6072
6173 If you'd rather write a case yourself to see exactly what the files contain, follow [Write a case manually](#write-a-case-manually) and come back here to run it.
from line 165
153165Write me a commit message for this change: I renamed getUser to fetchUser and updated the three call sites.
154166```
155167
156Each run starts in an empty working directory, so put whatever the task needs in the prompt itself, or [set up the workspace](#add-setup-or-history-with-case-yaml) first. The [full list of frontmatter fields](#prompt-md-fields) covers the model, timeout, tags, and environment variables.
168Each run starts in an empty working directory, so put whatever the task needs in the prompt itself, or [set up the workspace](#add-setup-or-history-with-case-yaml) first.
157169
170The [full list of frontmatter fields](#prompt-md-fields) covers the model, timeout, tags, and environment variables.
171
158172Each file under `graders/` is one check applied after the run. Open `evals/first-case/graders/criteria.md` and replace the placeholder with a rubric for the judge model, written as concrete PASS and FAIL conditions:
159173
160174```markdown theme={null}
from line 180
166180FAIL if <what a wrong or missing response looks like>.
167181```
168182
169Then add a second grader that checks whether your skill is what produced the answer. Create `evals/first-case/graders/skill-fired.md`, replacing `your-skill-name` with the `name` from your skill's `SKILL.md`:
183Then add a second grader that checks whether your skill is what produced the answer. Create `evals/first-case/graders/skill-fired.md`, replacing `your-skill-name` with the skill's directory name under `skills/`, which is the name Claude invokes it by:
170184
171185```markdown theme={null}
172186---
from line 190
176190---
177191```
178192
179This passes when Claude invoked that skill at least once during the run, including by its namespaced `plugin-name:skill-name` form. [Grader types](#grader-types) lists the other checks available, such as matching a regex or confirming a file was created.
193This passes when Claude invoked that skill at least once during the run, including by its namespaced `plugin-name:skill-name` form.
180194
195[Grader types](#grader-types) lists the other checks available, such as matching a regex or confirming a file was created.
196
181197With both files saved, run the case the way the [quickstart](#create-your-first-eval-suite) does, with `claude plugin eval .` from the plugin root.
182198
183199<h3 id="set-run-limits-and-tools-in-prompt-md">
from line 200
184200 Set run limits and tools in prompt.md
185201</h3>
186202
187Set a case's `max_turns`, `timeout_seconds`, `model`, `tags`, and the `allowed_tools` it may use in `prompt.md` frontmatter; the [prompt.md frontmatter](#prompt-md-fields) reference lists every field and its default. Claude receives the body exactly as you wrote it. `@path` mentions in it aren't expanded into file attachments, so if Claude needs to read a file, grant a tool for it in `allowed_tools`.
203Set a case's `max_turns`, `timeout_seconds`, `model`, `tags`, and the `allowed_tools` it may use in `prompt.md` frontmatter; the [prompt.md frontmatter](#prompt-md-fields) reference lists every field and its default.
188204
205Claude receives the body exactly as you wrote it. `@path` mentions in it aren't expanded into file attachments, so if Claude needs to read a file, grant a tool for it in `allowed_tools`.
206
189207<h3 id="grade-the-result">
190208 Choose and weight graders
191209</h3>
from line 210
192210
193211A grader's frontmatter sets its `type`, and optionally a `weight` that makes it count for more of the run's score and an [`arm`](#compare-against-a-no-plugin-baseline) that controls how it's scored against the baseline. Of the six types, `regex`, `tool_used`, `tool_order`, and `file_exists` are computed from the transcript and files and cost nothing, while `llm` and `baseline` call a judge model and add to the run's cost.
194212
195There are no custom-code graders. [Grader types](#grader-types) lists each type's options and pass condition, and [what a grader can look at](#what-a-grader-can-look-at) lists the values `target` and `focus` accept.
213There are no custom-code graders.
196214
215[Grader types](#grader-types) lists each type's options and pass condition, and [what a grader can look at](#what-a-grader-can-look-at) lists the values `target` and `focus` accept.
216
197217The judge for `llm` and `baseline` graders is a small fast model by default. Pass `--judge-model sonnet` or a full model ID to use a stronger one for nuanced rubrics.
198218
199219#### Choose graders that give a stable signal
from line 221
201221An `llm` grader asks a model for a verdict, so its answer can differ between runs, and it differs more the longer the text it has to read. These habits keep a suite's scores steady enough to trust:
202222
203223* For long output such as a generated file, grade it with a `regex` grader over the file's contents, which checks the whole file the same way every time. Keep `llm` graders for short outputs, with rubrics written as concrete PASS and FAIL conditions.
204* Give each case one grader on the result, such as the final message or a produced file, and one on how Claude got there, such as `tool_used` or `tool_order`. Together they tell you both whether the answer was right and whether your plugin produced it.
224* Give each case one grader on the result, such as the final message or a produced file, and one on the steps Claude took to produce it, such as `tool_used` or `tool_order`. Together they tell you both whether the answer was right and whether your plugin produced it.
205225* If a case's `tool_used: Skill` grader passes but `Δ` is negative, suspect the judge before the plugin. A small judge model can mark a correct answer wrong because it's formatted differently from what the rubric describes. Re-run with `--judge-model sonnet`, and tighten the rubric so formatting doesn't decide the verdict.
206226* To check that a build or test passed inside the run, have the prompt ask Claude to run it and write the outcome to a file, grade that file, and assert the command ran with a `tool_used` grader whose `input_match` names the command.
207227
from line 229
209229 Score against the no-plugin baseline
210230</h3>
211231
212When a plugin is under test, each case runs in two arms by default. The with-arm is its runs with the plugin loaded, and the without-arm is the same number of runs with no plugin at all. The summary and report show both scores and `Δ`, the with-arm score minus the without-arm score. Pass `--ablation none` to run only the with-arm, which halves the cost when you don't need the comparison, such as while iterating on graders.
232When a plugin is under test, each case runs in two arms by default. The with-arm is its runs with the plugin loaded, and the without-arm is the same number of runs with no plugin at all. The summary and report show both scores and `Δ`, the with-arm score minus the without-arm score.
213233
234Pass `--ablation none` to run only the with-arm, which halves the cost when you don't need the comparison, such as while iterating on graders.
235
214236In a two-arm run, some graders are reported with `scored: false`. A check like "the skill was invoked" can never pass without the plugin, so counting it would push the without-arm toward zero and inflate `Δ`. To keep the two arms comparable, Claude Code excludes such graders from the score in both arms and reports them in the with-arm as pass/fail indicators only. That includes:
215237
216238* Every `tool_used` grader whose `tool` is `Skill`
239* Every `regex` grader with `target: mock_calls` and every `llm` grader with `focus: mock_calls`, when each [mocked server](#mock-mcp-servers) in the case is one your plugin declares
217240* Any grader you mark `arm: with-only`
218241
219If every grader in a case is one of these, they're scored normally instead, since there would be nothing left to score. Set `arm: both` on a grader to score it in both arms regardless, which is what you want for a "must not invoke the skill" check with `min: 0` and `max: 0`. Under `--ablation none` nothing is excluded, so the same suite can produce a different absolute score in the two modes.
242Three settings change that exclusion:
220243
244* **Every grader excluded**: if every grader in a case is in the excluded set, they're scored normally instead, since there would be nothing left to score.
245* **`arm: both`**: set `arm: both` on a grader to score it in both arms regardless, which is what you want for a "must not invoke the skill" check with `min: 0` and `max: 0`.
246* **`--ablation none`**: under `--ablation none` nothing is excluded, so the same suite can produce a different absolute score in the two modes.
247
221248### Use a different eval directory
222249
223250If `evals/` is already taken by another tool, keep the suite in a different directory. You can record that directory in the plugin's `plugin.json` so every run and every collaborator uses it, or pass it on the command line for a single run:
from line 256
229256
230257## Set up fixtures and mocks
231258
232A case can need more than a prompt: files or a git repository in the workspace, an earlier conversation to continue, or answers from the MCP servers your plugin talks to. Each of those is set up beside the case so runs stay repeatable.
259A case can need more than a prompt: files or a git repository in the workspace, an earlier conversation to continue, or answers from the MCP servers your plugin connects to. Each of those is set up beside the case so runs stay repeatable.
233260
234261<h3 id="add-setup-or-history-with-case-yaml">
235262 Seed the workspace or conversation
236263</h3>
237264
238Each run starts in an empty workspace. When a case needs more than the prompt, add a `case.yaml` beside `prompt.md` with a `context` block.
265Each run starts in an empty workspace. When a case needs more than the prompt, add a `case.yaml` beside `prompt.md` with a `context` block:
239266
240To create fixture files or a git repository first, write a Bash script in the case directory and name it in `context.scaffold_script`. The script runs as you, outside the agent's sandbox, and only when you pass `--scaffold`, so pass that flag only for suites you or your organization wrote. To continue an earlier conversation, save the transcript as a `.jsonl` file and name it in `context.history_file`, and the case's prompt becomes the next user turn. To let Claude read fixture directories in the case during the run, list them in `context.add_dirs`.
267* **Fixture files or a git repository**: write a Bash script in the case directory and name it in `context.scaffold_script`. The script runs as you, outside the agent's sandbox, and only when you pass `--scaffold`, so pass that flag only for suites you or your organization wrote.
268* **An earlier conversation to continue**: save the transcript as a `.jsonl` file and name it in `context.history_file`, and the case's prompt becomes the next user turn.
269* **Fixture directories Claude can read during the run**: list them in `context.add_dirs`.
241270
242271A `case.yaml` also needs `schema_version: "1.1"` and `name`; the [case.yaml fields](#case-yaml-fields) reference has the full list.
243272
from line 285
256285 Mock MCP servers
257286</h3>
258287
259You can evaluate a plugin whose skills call MCP tools without the real service behind them. Put one Markdown file per tool under `evals/mocks/<server>/<tool>.md` for the whole suite, or under a case's own `mocks/` directory for one case, where `<server>` is the server's name in your plugin's [MCP configuration](/docs/en/plugins-reference#mcp-servers).
288You can evaluate a plugin whose skills call MCP tools without the real service behind them. Put one Markdown file per tool under `evals/mocks/<server>/<tool>.md` for the whole suite, or under a case's own `mocks/` directory for one case, where `<server>` is the server's name in your plugin's [MCP configuration](/docs/en/plugins/components#mcp-servers).
260289
261A run never starts your plugin's real MCP servers unless you ask. Claude Code registers a stand-in under each server's own name. Tools with a mock file answer from it and are allowed without an `--allow-tools` grant, and a tool with no mock file isn't available to Claude. A server with no mocks at all appears in the case's `mocked:` progress line as `plugin_<plugin>_<server>[not started: no mock]`.
290A run never starts your plugin's real MCP servers unless you ask. Claude Code registers a substitute server under each server's own name. Tools with a mock file answer from it and are allowed without an `--allow-tools` grant, and a tool with no mock file isn't available to Claude. A server with no mocks at all appears in the case's `mocked:` progress line as `plugin_<plugin>_<server>[not started: no mock]`.
262291
263The file's body is what the tool returns to Claude. This mock stands in for a `create_issue` tool on a server named `tracker`, checks the input Claude sends, and echoes the title back. Save it as `evals/mocks/tracker/create_issue.md`:
292The file's body is what the tool returns to Claude. This mock substitutes for a `create_issue` tool on a server named `tracker`, checks the input Claude sends, and echoes the title back. Save it as `evals/mocks/tracker/create_issue.md`:
264293
265294```markdown theme={null}
266295---
from line 301
272301Created issue #4821: {{input.title}}
273302```
274303
275Insert fields from the call's input with `{{input.<field>}}`, and the contents of a fixture file beside the mock with `{{file:fixtures/{input.<field>}.json}}`. The `expect:` block guards the input. If a call violates it, the run aborts with score 0 and records why, so a case can assert what your plugin asked the server to do. Set `error: true` to return the body as a tool error instead, or `type: agent` to have a small model answer as the server from instructions in the body. The [mock file reference](#mock-files) lists every key and the `_server.md` and `_tools.json` files.
304A mock file's body and frontmatter accept these options:
276305
306* **Substitutions**: insert fields from the call's input with `{{input.<field>}}`, and the contents of a fixture file beside the mock with `{{file:fixtures/{input.<field>}.json}}`.
307* **`expect:`**: the `expect:` block guards the input. If a call violates it, the run aborts with score 0 and records why, so a case can assert what your plugin asked the server to do.
308* **`error: true`**: set `error: true` to return the body as a tool error instead.
309* **`type: agent`**: set `type: agent` to have a small model answer as the server from instructions in the body.
310
311The [mock file reference](#mock-files) lists every key and the `_server.md` and `_tools.json` files.
312
277313To grade the calls themselves, point a grader at `target: mock_calls`.
278314
279315To run against the plugin's real MCP servers instead, pass one of these flags. Either way those processes run as you, outside the run's sandbox, and their tools need an [`--allow-tools` grant](#grant-tools):
from line 336
300336| A plugin's root directory, such as `.` | Every case under its eval directory, with that plugin loaded |
301337| A single `prompt.md` or `case.yaml` file | That case, with its enclosing plugin loaded |
302338| An installed plugin by name, `name` or `name@marketplace` | The cases in the installed copy's eval directory, with the installed copy loaded. Results are written under `./evals/results/` in your current directory, or `./<dir>/results/` with `--eval-dir` |
303| `name@skills-dir` | The same, for a [skills-directory plugin](/docs/en/plugins-reference#skills-directory-plugins) |
339| `name@skills-dir` | The same, for a [skills-directory plugin](/docs/en/plugins/loading#plugins-shared-through-a-repository) |
304340| Omitted | The current directory as a path |
305341
306Add `--case <glob>` to filter by case name and `--tag <tag>` to keep cases with any of the given tags. Put the target before `--tag`, `--allow-tools`, and `--json`. The first two take a list and `--json` takes an optional path, so each of them reads a target that follows as its own value.
342Add `--case <glob>` to filter by case name and `--tag <tag>` to keep cases with any of the given tags.
307343
344Put the target before `--tag`, `--allow-tools`, and `--json`. The first two take a list and `--json` takes an optional path, so each of them reads a target that follows as its own value.
345
308346### Grant tools
309347
310348Runs never stop to ask for permission. Built-in tools that need a grant you didn't give, such as `Bash`, `Write`, `Edit`, `WebFetch`, and `WebSearch`, are removed from the session, so Claude can't call them at all.
311349
312The allowlist is the read-only tools the case lists in `allowed_tools`, from `Read`, `Glob`, `Grep`, `NotebookRead`, `Skill`, `Agent`, `TodoWrite`, and the task tools `TaskCreate`, `TaskGet`, `TaskList`, `TaskUpdate`, and `TaskStop`, plus whatever you grant with `--allow-tools`. That grant applies to every case in the run. To let cases use `Bash`, `Write`, `Edit`, `WebFetch`, or `WebSearch`, grant them yourself:
350A run allows only the read-only tools the case lists in `allowed_tools`, from `Read`, `Glob`, `Grep`, `NotebookRead`, `Skill`, `AskUserQuestion`, `Agent`, `TodoWrite`, and the task tools `TaskCreate`, `TaskGet`, `TaskList`, `TaskUpdate`, and `TaskStop`, plus whatever you grant with `--allow-tools`. That grant applies to every case in the run. To let cases use `Bash`, `Write`, `Edit`, `WebFetch`, or `WebSearch`, grant them yourself:
313351
314352```bash theme={null}
315353claude plugin eval . --allow-tools Write Edit "Bash(npm test *)"
316354```
317355
318When a case asked for a tool you didn't grant, the run lists it on stderr as `not granted`. Tools on a [mocked](#mock-mcp-servers) MCP server need no grant. Tools on a real plugin MCP server need both the server started, with `--allow-real-servers` or `--mocks off`, and a grant by name, such as `--allow-tools "mcp__plugin_my-plugin_github__*"`; a plugin's MCP tools are named `mcp__plugin_<plugin>_<server>__<tool>`.
356When a case asked for a tool you didn't grant, the progress output lists it as `not granted`. Tools on a [mocked](#mock-mcp-servers) MCP server need no grant. Tools on a real plugin MCP server need both the server started, with `--allow-real-servers` or `--mocks off`, and a grant by name, such as `--allow-tools "mcp__plugin_my-plugin_github__*"`; a plugin's MCP tools are named `mcp__plugin_<plugin>_<server>__<tool>`.
319357
320358When you grant `Bash` in any form, every command runs under Claude Code's [OS-level sandbox](/docs/en/sandboxing). Writes are confined to the run's workspace, your home directory and Claude Code configuration are unreadable, and network access is limited to domains you grant with `--allow-tools "WebFetch(domain:example.com)"`. If you grant Bash or PowerShell on a machine with no sandbox backend, Claude Code refuses each run rather than running it unconfined, and the case shows a run error and usually scores 0. Native Windows has no backend, so run shell-granting suites under WSL2; on Linux, install `bubblewrap` and `socat` first. See the [sandboxing prerequisites](/docs/en/sandboxing).
321359
from line 361
323361
324362This table covers the options for run count, models, scoring, cost, tool grants, mocks, and output. Run `claude plugin eval --help` for the complete list, which also includes `--case`, `--tag`, `--eval-dir`, `--no-scaffold`, `--report`, and `--verbose`.
325363
326| Option | Default | Effect |
327| :------------------------- | :----------------------------------------------------------------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
328| `--runs <n>` | Each case's `runs`, else 3 | Runs per case per arm |
329| `-j`, `--concurrency <n>` | `1` | Run up to this many agent runs at once, from 1 to 8. They share your account's rate limit, so this shortens wall-clock time rather than raising throughput past that limit. Results keep case order |
330| `--model <model>` | Each case's `model`, else `ANTHROPIC_MODEL` if set, else Claude Code's default | Model for the agent under test. Pin it in CI so a model rollout isn't mistaken for a plugin regression |
331| `--judge-model <model>` | A small fast model | Model for `llm` and `baseline` graders |
332| `--ablation <mode>` | `with-without` when a plugin resolves, else `none` | Whether to also run each case without the plugin to measure what it adds. `none` runs one arm; `with-without` adds the no-plugin baseline |
333| `--threshold <0..1>` | `1.0` | A case passes when its with-arm score is at least this. Any case below it makes the command exit 1 |
334| `--max-cost-usd <usd>` | No ceiling | A ceiling on the run's list-price cost estimate, not on plan usage. Checked before each run starts. Once spent, nothing further starts; runs already in flight finish, so spend can pass the ceiling by those runs. If any run is left unstarted, the command exits 2 with partial results |
335| `--allow-tools <tools...>` | None | Grant tools beyond the read-only set. See [Grant tools](#grant-tools) |
336| `--scaffold` | Off | Run each case's [`scaffold_script`](#add-setup-or-history-with-case-yaml) |
337| `--trust-plugin` | Off | Skip the first-run trust prompt for a plugin whose code and suite you'd run yourself. Pass it in CI so the job is never refused by or left waiting at the prompt. See [What a run can access](#security) |
338| `--mocks <mode>` | `record` | `record` answers MCP tool calls from [mocks](#mock-mcp-servers), doesn't start the plugin's real servers, and saves agent-mock answers for replay. `off` ignores mocks and starts the plugin's real MCP servers |
339| `--allow-real-servers` | Off | With `--mocks record`, also start the plugin's real MCP servers for servers that have no mock |
340| `--json [path]` | Off | Print the [result document](#json-result) to stdout, or write it to a path ending in `.json`. The run is quiet: no progress lines or summary table |
341| `--output-dir <dir>` | `<eval dir>/results/<timestamp>/` | Where `aggregate-result.json` and `report.html` go |
342| `--no-publish` | | Keep the HTML report local. See [HTML report](#html-report) |
343| `--publish-report` | | Publish the report even where it would stay local by default, such as a run a Claude Code session started |
344| `--keep-temp` | Off | Keep every run's sandbox directory and print its path, for debugging what Claude produced |
364| Option | Default | Effect |
365| :------------------------- | :----------------------------------------------------------------------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
366| `--runs <n>` | Each case's `runs`, else 3 | Runs per case per arm |
367| `-j`, `--concurrency <n>` | `1` | Run up to this many agent runs at once, from 1 to 8. They share your account's rate limit, so this shortens wall-clock time rather than raising throughput past that limit. Results keep case order |
368| `--model <model>` | Each case's `model`, else `ANTHROPIC_MODEL` if set, else Claude Code's default | Model for the agent under test. Pin it in CI so a model rollout isn't mistaken for a plugin regression |
369| `--judge-model <model>` | A small fast model | Model for `llm` and `baseline` graders |
370| `--ablation <mode>` | `with-without` when a plugin resolves, else `none` | Whether to also run each case without the plugin to measure what it adds. `none` runs one arm; `with-without` adds the no-plugin baseline |
371| `--threshold <0..1>` | `1.0` | A case passes when its with-arm score is at least this. Any case below it makes the command exit 1 |
372| `--max-cost-usd <usd>` | No ceiling | A ceiling on the run's list-price cost estimate, not on plan usage. Checked before each run starts. Once spent, nothing further starts; runs that already started finish, so spend can pass the ceiling by those runs. If any run is left unstarted, the command exits 2 with partial results |
373| `--allow-tools <tools...>` | None | Grant tools beyond the read-only set. See [Grant tools](#grant-tools) |
374| `--scaffold` | Off | Run each case's [`scaffold_script`](#add-setup-or-history-with-case-yaml) |
375| `--trust-plugin` | Off | Skip the first-run trust prompt for a plugin whose code and suite you'd run yourself. Pass it in CI so the job is never refused by or left waiting at the prompt. See [What a run can access](#security) |
376| `--mocks <mode>` | `record` | `record` answers MCP tool calls from [mocks](#mock-mcp-servers), doesn't start the plugin's real servers, and saves agent-mock answers for replay. `off` ignores mocks and starts the plugin's real MCP servers |
377| `--allow-real-servers` | Off | With `--mocks record`, also start the plugin's real MCP servers for servers that have no mock |
378| `--json [path]` | Off | Print the [result document](#json-result) to stdout, or write it to a path ending in `.json`. The run is quiet: no progress lines or summary table |
379| `--output-dir <dir>` | `<eval dir>/results/<timestamp>/` | Where `aggregate-result.json` and `report.html` go |
380| `--no-publish` | | Keep the HTML report local. See [HTML report](#html-report) |
381| `--publish-report` | | Publish the report even where it would stay local by default, such as a run a Claude Code session started |
382| `--keep-temp` | Off | Keep every run's sandbox directory and print its path, for debugging what Claude produced |
345383
346384<h3 id="run-evals-in-ci">
347385 Run evals in CI
from line 408
370408| 130 | Interrupted. Partial results are written |
371409| 143 | Terminated, such as by a CI timeout |
372410
373Problems writing or publishing the HTML report never change the exit code. To see why a case scored low, run it locally without `--json` so the per-run progress and grader lines print.
411Problems writing or publishing the HTML report never change the exit code.
374412
375A CI runner needs a Claude Code install and [credentials in the environment](/docs/en/authentication) such as `ANTHROPIC_API_KEY`. Without `--trust-plugin`, a job whose checkout directory Claude Code doesn't already trust is refused with exit 1 when it has no terminal, or waits at the prompt when the runner allocates one. `claude plugin eval init` needs a terminal to ask you its questions; in CI, run `claude plugin eval init --bare <name>` to get the blank template.
413To see why a case scored low, run it locally without `--json` so the per-run progress and grader lines print.
376414
415A CI runner also needs these in place:
416
417* **Install and credentials**: a CI runner needs a Claude Code install and [credentials in the environment](/docs/en/authentication) such as `ANTHROPIC_API_KEY`.
418* **Trust**: without `--trust-plugin`, a job whose checkout directory Claude Code doesn't already trust needs the [first-run trust prompt](#trust-the-plugin-directory), and a run that can't ask is refused with exit 1.
419* **`init` in CI**: `claude plugin eval init` needs a terminal to ask you its questions; in CI, run `claude plugin eval init --bare <name>` to get the blank template.
420
377421To keep costs predictable, give quick every-change suites only graders that don't call a judge, use `--ablation none` where you don't need `Δ`, and leave `partial: true` documents and runs with `skippedPaidGraders` out of any trend you chart.
378422
379423## Read the results
from line 432
388432
389433Read it from the top down:
390434
391* **The verdict line and tiles** answer whether the plugin helped across the whole suite. Suite score is the mean of the per-case with-plugin scores, Ablation Δ is how far that sits above or below the baseline score, and Cases counts how many met the threshold. Perfect runs is the share of with-plugin runs where every grader passed.
435* **The verdict line and tiles** answer whether the plugin helped across the whole suite. Suite score is the mean of the per-case with-plugin scores, Ablation Δ is how far that is above or below the baseline score, and Cases counts how many met the threshold. Perfect runs is the share of with-plugin runs where every grader passed.
392436* **Each case card** shows the case's own `Δ` and with-plugin score, with a tick on the bar at the threshold. A case whose `Δ` is negative gets a red left edge, so regressions stand out when you scroll.
393437* **Inside a case**, the with-plugin runs come first and the baseline runs after. Each run lists its graders with a pass or fail chip. A failed grader is already expanded with its explanation, and an `llm` grader also shows the judge's votes and the evidence it was shown, which is where you find out why a run scored low. Graders that don't count toward the score, such as `tool_used: Skill`, carry a `plugin-fired indicator` badge.
394438* **Prompt and Graders**, below the runs, show the case's prompt and each grader's rubric or pattern, so someone reading the report without the suite can see what was asked and what counted as good.
from line 465
421465 What a run can access
422466</h2>
423467
424`claude plugin eval` loads the target plugin's skills, hooks, and agents and runs its eval suite on your machine, as you. Pointing it at a plugin is the same trust decision as `claude --plugin-dir`, so only evaluate plugins you trust. The isolation described in this section limits what the agent under test can reach; it isn't a boundary against the plugin's own code, and a suite that passes says nothing about whether the plugin is safe.
468`claude plugin eval` loads the target plugin's skills, hooks, and agents and runs its eval suite on your machine, as you. Pointing it at a plugin is the same trust decision as `claude --plugin-dir`, so only evaluate plugins you trust.
425469
470The isolation described in this section limits what the agent under test can reach; it isn't a boundary against the plugin's own code, and a suite that passes says nothing about whether the plugin is safe.
471
426472### Trust the plugin directory
427473
428The first time you run `claude plugin eval` against a directory, Claude Code asks `Trust this plugin directory?` before it loads anything from it, unless you already accepted the trust prompt there in an interactive `claude` session. Inside a git repository, answering yes trusts the whole repository, for interactive sessions too. When stdin or stdout isn't a terminal, or under `--json`, the run can't ask and is refused with exit 1; pass `--trust-plugin` to assert the trust yourself, only for a plugin you'd run on your own machine. A target you name rather than give as a path, meaning an installed plugin or a skills-directory plugin, skips the prompt.
474The first time you run `claude plugin eval` against a directory, Claude Code asks `Trust this plugin directory?` before it loads anything from it, unless you already accepted the trust prompt there in an interactive `claude` session. Inside a git repository, answering yes trusts the whole repository, for interactive sessions too. When stdin or stdout isn't a terminal, under `--json`, or when the `CI` environment variable is set to a true value such as `true`, the run can't ask and is refused with exit 1; pass `--trust-plugin` to assert the trust yourself, only for a plugin you'd run on your own machine. A target you name rather than give as a path, meaning an installed plugin or a skills-directory plugin, skips the prompt.
429475
430Some parts of the plugin and suite run only when you pass their flag for that run: a case's [`scaffold_script`](#add-setup-or-history-with-case-yaml) with `--scaffold`, [tools beyond the read-only set](#grant-tools) with `--allow-tools`, and the plugin's [real MCP servers](#mock-mcp-servers) with `--allow-real-servers` or `--mocks off`. A case's `allowed_tools` and a skill's own `allowed-tools` frontmatter can't widen any of them. When the plugin ships hooks you didn't write, or you start its real MCP servers, treat its scores as advisory unless you ran it in an isolated environment such as a container or CI runner, since hooks and servers run outside the agent's sandbox and could touch the files the graders read.
476Some parts of the plugin and suite run only when you pass their flag for that run:
431477
478* A case's [`scaffold_script`](#add-setup-or-history-with-case-yaml) with `--scaffold`
479* [Tools beyond the read-only set](#grant-tools) with `--allow-tools`
480* The plugin's [real MCP servers](#mock-mcp-servers) with `--allow-real-servers` or `--mocks off`
481
482A case's `allowed_tools` and a skill's own `allowed-tools` frontmatter can't widen any of them.
483
484When the plugin includes hooks you didn't write, or you start its real MCP servers, treat its scores as advisory unless you ran it in an isolated environment such as a container or CI runner, since hooks and servers run outside the agent's sandbox and could modify the files the graders read.
485
432486<h3 id="how-runs-are-isolated">
433487 How runs are isolated
434488</h3>
435489
436Each run gets a throwaway home directory, working directory, and Claude Code configuration, and the agent under test runs there as a `claude -p` child process with only your plugin loaded. Keep these consequences in mind when you write cases:
490Each run gets a temporary home directory, working directory, and Claude Code configuration, and the agent under test runs there as a `claude -p` child process with only your plugin loaded. Keep these consequences in mind when you write cases:
437491
438492* **Nothing personal or project-level loads.** Your user settings, hooks, `CLAUDE.md` files, MCP servers, other installed plugins, memory, and skills are absent, and no project-scoped `.claude/` or `.mcp.json` above the sandbox is read. Most of your shell environment is withheld too; only an [allowlist](#prompt-md-fields) and `EVAL_*` variables reach the run. If the plugin needs setup, ship it in the plugin, create it in a `scaffold_script`, or pass `EVAL_*` variables.
439493* **Managed policy can still restrict a run.** Restrictions in [managed settings](/docs/en/managed-settings) an administrator deployed to the machine apply inside a run, so results on a managed machine can differ from an unmanaged one by that policy.
from line 541
487541| `timeout_seconds` | `300` | Wall-clock cap per run, up to 3600 |
488542| `allowed_tools` | `[]` | Tools the case wants, such as `[Read, Glob, Grep, Skill]`. Read-only tools are granted when listed here; for anything else, see [Grant tools](#grant-tools) |
489543| `append_system_prompt` | | Text appended to the child session's system prompt |
490| `env` | `{}` | Extra environment variables for the child session. Keys must match `EVAL_[A-Z0-9_]*`; any other key fails the run. The run inherits only an allowlist from your shell: basics such as `PATH` and locale, proxy and certificate settings, the variables that select and authenticate your model provider, most `ANTHROPIC_*` and `CLAUDE_CODE_*` configuration, and `EVAL_*`. To hand the plugin anything else, such as a toolchain setting, export it as an `EVAL_*` variable |
544| `env` | `{}` | Extra environment variables for the child session. Keys must match `EVAL_[A-Z0-9_]*`; any other key fails the run. The run inherits only an allowlist from your shell: basics such as `PATH` and locale, proxy and certificate settings, the variables that select and authenticate your model provider, most `ANTHROPIC_*` and `CLAUDE_CODE_*` configuration, and `EVAL_*`. To pass the plugin anything else, such as a toolchain setting, export it as an `EVAL_*` variable |
491545
492546<h3 id="case-yaml-fields">
493547 case.yaml fields
494548</h3>
495549
496`case.yaml` describes the same case in YAML and adds the fields that point at other files. It requires `schema_version: "1.1"` and `name`. The `prompt.md` fields `description`, `tags`, `plugins`, `runs`, and `expected_outcome` go at the top level; `model`, `max_turns`, `timeout_seconds`, `allowed_tools`, `append_system_prompt`, and `env` go under `execution:`. When both files exist, `prompt.md` frontmatter overrides the matching `case.yaml` fields, the `prompt.md` body is the prompt, and `graders/*.md` are added after any graders listed in `case.yaml`.
550`case.yaml` is an alternative or companion to `prompt.md`: it describes a case in YAML and adds the fields that point at other files. It requires `schema_version: "1.1"` and `name`. The `prompt.md` fields `description`, `tags`, `plugins`, `runs`, and `expected_outcome` go at the top level; `model`, `max_turns`, `timeout_seconds`, `allowed_tools`, `append_system_prompt`, and `env` go under `execution:`. When both files exist, `prompt.md` frontmatter overrides the matching `case.yaml` fields, the `prompt.md` body is the prompt, and `graders/*.md` are added after any graders listed in `case.yaml`.
497551
498552These fields exist only in `case.yaml`:
499553
from line 563
509563
510564Every grader file under `graders/` takes these keys in frontmatter, plus the options for its type. The grader's name is the filename without `.md`:
511565
512| Key | Default | Purpose |
513| :------- | :------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
514| `type` | required | One of the [grader types](#grader-types) |
515| `weight` | `1` | Relative weight in the run's score. Any positive number |
516| `arm` | unset | `with-only` excludes the grader from scoring in a [two-arm run](#compare-against-a-no-plugin-baseline); `both` forces a `tool_used: Skill` grader to be scored in both arms |
566| Key | Default | Purpose |
567| :------- | :------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
568| `type` | required | One of the [grader types](#grader-types) |
569| `weight` | `1` | Relative weight in the run's score. Any positive number |
570| `arm` | unset | `with-only` excludes the grader from scoring in a [two-arm run](#compare-against-a-no-plugin-baseline); `both` forces a grader Claude Code would otherwise exclude to be scored in both arms |
517571
518572#### What a grader can look at
519573
from line 602
548602
549603| Key | Default | Purpose |
550604| :----------- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
551| `type` | `fixed` | `fixed` returns the body as written. `agent` treats the body as instructions for a small model that plays the server for the run and sees earlier calls as history |
605| `type` | `fixed` | `fixed` returns the body as written. `agent` treats the body as instructions for a small model that acts as the server for the run and sees earlier calls as history |
552606| `expect` | unset | A map from dotted input paths to a type name such as `string`, `number`, `boolean`, `array`, or `object`, a `/regex/`, a literal, or a list of allowed literals. A call that violates it aborts the run with score 0 and is reported as `aborted` with the server, tool, and reason |
553607| `error` | `false` | `fixed` only. Return the body as a tool error |
554608| `abort_when` | unset | `agent` only. Prose listing the only conditions under which the agent may abort the run |
from line 616
562616
563617## Troubleshooting
564618
565These are the problems authors hit most often, keyed on what you see.
619These are the problems authors encounter most often, keyed on what you see.
566620
567621### "plugin eval is currently in early access"
568622
from line 628
574628
575629### "is not a trusted plugin directory, and this run cannot stop to ask you about it"
576630
577This is the first run against a directory Claude Code doesn't trust yet, and it can't ask you because stdin or stdout isn't a terminal or you passed `--json`. Run `claude plugin eval <dir>` once in a terminal and answer the prompt, or pass `--trust-plugin` if you trust the plugin's code and suite. See [What a run can access](#security).
631This is the first run against a directory Claude Code doesn't trust yet, and it can't ask you because stdin or stdout isn't a terminal, you passed `--json`, or the `CI` environment variable is set to a true value such as `true`. Run `claude plugin eval <dir>` once in a terminal and answer the prompt, or pass `--trust-plugin` if you trust the plugin's code and suite. See [What a run can access](#security).
578632
579633### "No eval cases found"
580634
from line 650
596650
597651### Everything scores zero although the right files were produced
598652
599Your graders target `files`, the list of created paths, when you meant the file's contents. Use `{ source: file, path: <path> }` as the `target` or `focus`. Separately, `file_exists` counts only files created during the run, so a file the scaffold created or that Claude only edited is invisible to it; grade its contents, or use `tool_used` on `Edit`.
653Your graders target `files`, the list of created paths, when you meant the file's contents. Use `{ source: file, path: <path> }` as the `target` or `focus`.
600654
655Separately, `file_exists` counts only files created during the run, so a file the scaffold created or that Claude only edited is invisible to it; grade its contents, or use `tool_used` on `Edit`.
656
601657### A regex over the trace doesn't match text I can see
602658
603The default `target` is `last_message`, not the trace. When you do target `trace`, it's JSON per line, so quotes appear as `\"`. Regexes use JavaScript syntax, so put `i` in `flags` rather than writing `(?i)`.
659* **Wrong target**: the default `target` is `last_message`, not the trace.
660* **JSON escaping**: when you do target `trace`, it's JSON per line, so quotes appear as `\"`.
661* **Regex syntax**: regexes use JavaScript syntax, so put `i` in `flags` rather than writing `(?i)`.
604662
605663### Tools are denied, MCP tools are missing, or Bash won't run
606664
from line 666
608666
609667### The run exits 1 but the results look fine
610668
611The default `--threshold` is 1.0, so the command exits 1 when any case scores below perfect. Set a threshold that matches your bar. Exit 1 also covers a case file that failed to load, which is reported on stderr above the table.
669The default `--threshold` is 1.0, so the command exits 1 when any case scores below perfect. Set a threshold that matches the score you require. Exit 1 also covers a case file that failed to load, which is reported on stderr above the table.
612670
613671### "--json output path must end in .json"
614672
from line 674
616674
617675### A grader shows passed: false under a run that scored 1.0
618676
619That grader is excluded from the score by design in a two-arm run, and its `scored` field is `false`. See [Compare against a no-plugin baseline](#compare-against-a-no-plugin-baseline).
677That grader is excluded from the score by design in a two-arm run, and its `scored` field is `false`. See [Score against the no-plugin baseline](#compare-against-a-no-plugin-baseline).
620678
621679### Runs fail with a usage-limit or rate-limit error partway through
622680
from line 686
628686
629687## See also
630688
631* [Create plugins](/docs/en/plugins): build the plugin you're testing, and load it with `--plugin-dir` during development
632* [Plugins reference](/docs/en/plugins-reference#plugin-eval): the `plugin eval` and `plugin eval init` command entries and the manifest's `experimental.evals` key
689* [Create a plugin](/docs/en/plugins/create): build the plugin you're testing, and load it with `--plugin-dir` during development
690* [Plugin commands reference](/docs/en/plugins/cli-reference#plugin-eval): the `plugin eval` and `plugin eval init` command entries. The manifest's [`experimental.evals`](/docs/en/plugins/manifest-reference#fields) key is on the manifest reference
633691* [Skills](/docs/en/skills): how a skill's description decides when Claude invokes it, which is what a case that checks whether the skill triggers is measuring
634692* [Sandboxing](/docs/en/sandboxing): the OS-level sandbox that applies when you grant Bash to a run
635* [Create and distribute a plugin marketplace](/docs/en/plugin-marketplaces): publish the plugin once its suite passes
693* [Publish a plugin](/docs/en/plugins/publish): publish the plugin once its suite passes
694* [Measure plugin cost and usage](/docs/en/plugins/measure): what the plugin adds to each session's context and whether people still use it
636695
No line in this hunk matches that.