Follow Discord
Sweep 08 Oct 2026 · 18:53Z Build v2.1.295 516 read Stable v2.1.286 Latest v2.1.295 Next v2.1.295 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · claude-code

Test plugins with evals changedplugin-evals

Nearest release: v2.1.288, published 4 hours after upstream edited the page. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Upstream edited this page at 2 Oct 2026 13:59 UTC, give or take a minute or two: the time comes from Anthropic’s own sitemap rather than from a commit. This site recorded the change at 2 Oct 2026 14:07 UTC.

Upstream edited
Recorded here
Lines+4added
Lines−4removed
From line 215 where the diff opens
First seen 11 Sep 2026 this site's first read of the page
Recorded edits16to this page, all time

The whole hunk

from line 215, old and new numbered
/
lines
from line 215
215215 
216216[Grader types](#grader-types) lists each type's options and pass condition, and [what a grader can look at](#what-a-grader-can-look-at) lists the values `target` and `focus` accept.
217217 
218The judge for `llm` and `baseline` graders is a small fast model by default. Pass `--judge-model sonnet` or a full model ID to use a stronger one for nuanced rubrics.
218By default, the judge for `llm` and `baseline` graders is the model Claude Code uses for background tasks. Pass `--judge-model sonnet` or a full model ID to choose the judge yourself.
219219 
220220#### Choose graders that give a stable signal
221221 
from line 313
313313* **Substitutions**: insert fields from the call's input with `{{input.<field>}}`, and the contents of a fixture file beside the mock with `{{file:fixtures/{input.<field>}.json}}`.
314314* **`expect:`**: the `expect:` block guards the input. If a call violates it, the run aborts with score 0 and records why, so a case can assert what your plugin asked the server to do.
315315* **`error: true`**: set `error: true` to return the body as a tool error instead.
316* **`type: agent`**: set `type: agent` to have a small model answer as the server from instructions in the body.
316* **`type: agent`**: set `type: agent` to have the judge model answer as the server from instructions in the body.
317317 
318318The [mock file reference](#mock-files) lists every key and the `_server.md` and `_tools.json` files.
319319 
from line 373
373373| `--runs <n>` | Each case's `runs`, else 3 | Runs per case per arm |
374374| `-j`, `--concurrency <n>` | `1` | Run up to this many agent runs at once, from 1 to 8. They share your account's rate limit, so this shortens wall-clock time rather than raising throughput past that limit. Results keep case order |
375375| `--model <model>` | Each case's `model`, else `ANTHROPIC_MODEL` if set, else Claude Code's default | Model for the agent under test. Pin it in CI so a model rollout isn't mistaken for a plugin regression |
376| `--judge-model <model>` | A small fast model | Model for `llm` and `baseline` graders |
376| `--judge-model <model>` | The model for [background tasks](#grade-the-result) | Model for `llm` and `baseline` graders |
377377| `--ablation <mode>` | Decided per case; see [Score against the no-plugin baseline](#compare-against-a-no-plugin-baseline) | Whether to also run each case without the plugin to measure what it adds. `none` runs one arm; `with-without` adds the no-plugin baseline |
378378| `--threshold <0..1>` | `1.0` | A case passes when its with-arm score is at least this. Any case below it makes the command exit 1 |
379379| `--max-cost-usd <usd>` | No ceiling | A ceiling on the run's list-price cost estimate, not on plan usage. Checked before each run starts. Once spent, nothing further starts; runs that already started finish, so spend can pass the ceiling by those runs. If any run is left unstarted, the command exits 2 with partial results |
from line 609
609609 
610610| Key | Default | Purpose |
611611| :- | :- | :- |
612| `type` | `fixed` | `fixed` returns the body as written. `agent` treats the body as instructions for a small model that acts as the server for the run and sees earlier calls as history |
612| `type` | `fixed` | `fixed` returns the body as written. `agent` treats the body as instructions for the [judge model](#command-options), which acts as the server for the run and sees earlier calls as history |
613613| `expect` | unset | A map from dotted input paths to a type name such as `string`, `number`, `boolean`, `array`, or `object`, a `/regex/`, a literal, or a list of allowed literals. A call that violates it aborts the run with score 0 and is reported as `aborted` with the server, tool, and reason |
614614| `error` | `false` | `fixed` only. Return the body as a tool error |
615615| `abort_when` | unset | `agent` only. Prose listing the only conditions under which the agent may abort the run |
Feedback