Follow Discord
Sweep 08 Oct 2026 · 18:53Z Build v2.1.295 516 read Stable v2.1.286 Latest v2.1.295 Next v2.1.295 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · claude-code

Test plugins with evals changedplugin-evals

Nearest release: v2.1.294, published an hour after upstream edited the page. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Upstream edited this page at 8 Oct 2026 02:20 UTC, give or take a minute or two: the time comes from Anthropic’s own sitemap rather than from a commit. This site recorded the change at 8 Oct 2026 03:07 UTC.

Upstream edited
Recorded here
Lines+25added
Lines−4removed
From line 313 where the diff opens
First seen 11 Sep 2026 this site's first read of the page
Recorded edits16to this page, all time

The whole hunk

from line 313, old and new numbered
/
lines
from line 313
313313* **Substitutions**: insert fields from the call's input with `{{input.<field>}}`, and the contents of a fixture file beside the mock with `{{file:fixtures/{input.<field>}.json}}`.
314314* **`expect:`**: the `expect:` block guards the input. If a call violates it, the run aborts with score 0 and records why, so a case can assert what your plugin asked the server to do.
315315* **`error: true`**: set `error: true` to return the body as a tool error instead.
316* **`type: agent`**: set `type: agent` to have the judge model answer as the server from instructions in the body.
316* **`type: agent`**: set `type: agent` to have the judge model answer as the server from instructions in the body. Calls to agent mocks share one [budget per run](#mock-call-budget-exceeded) of four times the case's `max_turns`, and a call past it aborts the run with score 0.
317317 
318318The [mock file reference](#mock-files) lists every key and the `_server.md` and `_tools.json` files.
319319 
from line 464
464464| `cases[].aggregates.score` | Mean with-arm run score for the case |
465465| `cases[].aggregates.delta` | With-arm score minus without-arm score. Omitted when the case ran one arm or the arms aren't comparable |
466466| `cases[].arms.with[].error` | `null`, or why a run ended abnormally, such as `timed out after 300s`. A run that started but ended badly is still graded on what it produced, so a non-null error doesn't imply score 0 |
467| `cases[].arms.with[].aborted` | Present when a [mock](#mock-mcp-servers)'s `expect:` or `abort_when` stopped the run, with `server`, `tool`, and `reason`. The run scores 0 and `error` stays `null` |
467| `cases[].arms.with[].aborted` | Present when a [mock](#mock-mcp-servers) stopped the run through `expect:`, `abort_when`, or the [agent-mock call budget](#mock-call-budget-exceeded), with `server`, `tool`, and `reason`. The run scores 0 and `error` stays `null` |
468468| `cases[].arms.with[].skippedPaidGraders` | `true` when the cost ceiling skipped this run's judge graders, so its score isn't comparable |
469469| `costUsd`, `durationSeconds`, `claudeVersion` | Estimated cost at list price including judge calls, wall-clock seconds, and the Claude Code version that ran the suite |
470470 
from line 610
610610| Key | Default | Purpose |
611611| :- | :- | :- |
612612| `type` | `fixed` | `fixed` returns the body as written. `agent` treats the body as instructions for the [judge model](#command-options), which acts as the server for the run and sees earlier calls as history |
613| `expect` | unset | A map from dotted input paths to a type name such as `string`, `number`, `boolean`, `array`, or `object`, a `/regex/`, a literal, or a list of allowed literals. A call that violates it aborts the run with score 0 and is reported as `aborted` with the server, tool, and reason |
613| `expect` | unset | A map from dotted input paths to a type name such as `string`, `number`, `boolean`, `array`, or `object`, a [`/regex/`](#expect-patterns), a literal, or a list of allowed literals. A call that violates it aborts the run with score 0 and is reported as `aborted` with the server, tool, and reason |
614614| `error` | `false` | `fixed` only. Return the body as a tool error |
615615| `abort_when` | unset | `agent` only. Prose listing the only conditions under which the agent may abort the run |
616616 
617617Two optional files sit beside the tool files in a server's directory:
618618 
619* **`_server.md`**: a single `type: agent` mock that answers several tools, listed in its `tools:` frontmatter key. A `<tool>.md` for the same tool takes precedence. Put an `expect:` guard on the individual `<tool>.md`, not here
619* **`_server.md`**: a single `type: agent` mock that answers several tools, listed in its `tools:` frontmatter key. A `<tool>.md` for the same tool takes precedence. An `expect:` guard here is a load error unless `tools:` lists a single tool, so put the guard on the individual `<tool>.md` instead
620620* **`_tools.json`**: a saved `tools/list` response from the real server, so mocked tools carry their real descriptions and input schemas instead of a permissive placeholder
621621 
622622A case's own `mocks/` directory uses the same layout and overrides the suite's mocks file by file.
623623 
624<h4 id="expect-patterns">
625 Regex patterns in expect
626</h4>
627 
628A `/regex/` value in `expect:` uses a small dialect that Claude Code checks when it loads the suite:
629 
630* Literal characters, `.`, escapes such as `\d`, and character classes such as `[a-z]`
631* The quantifiers `*`, `+`, `?`, and the `{m,n}` forms, each on a single character, escape, or class
632* An optional `^` at the start and `$` at the end
633* The flags `i` and `s` only
634 
635A pattern outside the dialect, such as one with a group, alternation, a backreference, lookaround, or another flag, stops the case from loading: the case scores 0 and its error names the pattern. To allow several exact values, write a list of literals instead of an alternation.
636 
637Each pattern checks values only up to a maximum length, and a longer value counts as a violation. Quantifiers can lower that length, and a leading `^` raises it, so anchor patterns with `^` and keep quantifiers few.
638 
624639## Troubleshooting
625640 
626641These are the problems authors encounter most often, keyed on what you see.
from line 717
702717### Runs fail with a usage-limit or rate-limit error partway through
703718 
704719If your account reaches its plan's usage limit or an API rate limit while a suite is running, each later run ends with that error, is graded on what it produced, and usually scores 0. The suite still finishes and isn't marked `partial`, so the result can look like a regression. Check the `NOTES` column or `cases[].arms.with[].error` in the JSON for the limit message before trusting the scores, then re-run after the limit resets, with `--runs 1` or a `--case` filter if you need to stay under it.
720 
721<h3 id="mock-call-budget-exceeded">
722 "mock call budget exceeded"
723</h3>
724 
725Every `type: agent` [mock](#mock-mcp-servers) in a run draws on one call budget of four times the case's `max_turns`, which is 40 calls at the default of 10. Calls answered from `.replay/` recordings count too, and the case's `mock budget` progress line prints the budget. A call past it aborts the run with score 0 and this reason, so raise `max_turns` in the case for a skill that makes many calls to agent mocks.
705726 
706727### Runs time out or hit the turn cap
707728 
Feedback