Follow Discord
Sweep 29 Sep 2026 · 18:10Z Build v2.1.285 506 read Stable v2.1.280 Latest v2.1.285 Next v2.1.285 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · claude-code

Test plugins with evals changedplugin-evals

Nearest release: v2.1.283, published 4 hours before upstream edited the page. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Upstream edited this page at 25 Sep 2026 23:00 UTC, give or take a minute or two: the time comes from Anthropic’s own sitemap rather than from a commit. This site recorded the change at 28 Sep 2026 23:37 UTC.

Upstream edited
Recorded here
Lines+91added
Lines−91removed
From line 332 where the diff opens
First seen 11 Sep 2026 this site's first read of the page
Recorded edits9to this page, all time

The whole hunk

from line 332, old and new numbered
/
lines
from line 332
332332 
333333Most of the time you run `claude plugin eval .` from the plugin root, which runs every case in the suite with the plugin you're standing in loaded. To run a single case file, or to evaluate a plugin you installed rather than one you're developing, pass a different target:
334334 
335| Target | What runs |
336| :-------------------------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
337| A plugin's root directory, such as `.` | Every case under its eval directory, with that plugin loaded |
338| A single `prompt.md` or `case.yaml` file | That case, with its enclosing plugin loaded |
335| Target | What runs |
336| :- | :- |
337| A plugin's root directory, such as `.` | Every case under its eval directory, with that plugin loaded |
338| A single `prompt.md` or `case.yaml` file | That case, with its enclosing plugin loaded |
339339| An installed plugin by name, `name` or `name@marketplace` | The cases in the installed copy's eval directory, with the installed copy loaded. Results are written under `./evals/results/` in your current directory, or `./<dir>/results/` with `--eval-dir` |
340| `name@skills-dir` | The same, for a [skills-directory plugin](/docs/en/plugins/loading#plugins-shared-through-a-repository) |
341| Omitted | The current directory as a path |
340| `name@skills-dir` | The same, for a [skills-directory plugin](/docs/en/plugins/loading#plugins-shared-through-a-repository) |
341| Omitted | The current directory as a path |
342342 
343343Add `--case <glob>` to filter by case name and `--tag <tag>` to keep cases with any of the given tags.
344344 
from line 362
362362 
363363This table covers the options for run count, models, scoring, cost, tool grants, mocks, and output. Run `claude plugin eval --help` for the complete list, which also includes `--case`, `--tag`, `--eval-dir`, `--no-scaffold`, `--report`, and `--verbose`.
364364 
365| Option | Default | Effect |
366| :------------------------- | :----------------------------------------------------------------------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
367| `--runs <n>` | Each case's `runs`, else 3 | Runs per case per arm |
368| `-j`, `--concurrency <n>` | `1` | Run up to this many agent runs at once, from 1 to 8. They share your account's rate limit, so this shortens wall-clock time rather than raising throughput past that limit. Results keep case order |
369| `--model <model>` | Each case's `model`, else `ANTHROPIC_MODEL` if set, else Claude Code's default | Model for the agent under test. Pin it in CI so a model rollout isn't mistaken for a plugin regression |
370| `--judge-model <model>` | A small fast model | Model for `llm` and `baseline` graders |
371| `--ablation <mode>` | `with-without` when a plugin resolves, else `none` | Whether to also run each case without the plugin to measure what it adds. `none` runs one arm; `with-without` adds the no-plugin baseline |
372| `--threshold <0..1>` | `1.0` | A case passes when its with-arm score is at least this. Any case below it makes the command exit 1 |
373| `--max-cost-usd <usd>` | No ceiling | A ceiling on the run's list-price cost estimate, not on plan usage. Checked before each run starts. Once spent, nothing further starts; runs that already started finish, so spend can pass the ceiling by those runs. If any run is left unstarted, the command exits 2 with partial results |
374| `--allow-tools <tools...>` | None | Grant tools beyond the read-only set. See [Grant tools](#grant-tools) |
375| `--scaffold` | Off | Run each case's [`scaffold_script`](#add-setup-or-history-with-case-yaml) |
376| `--trust-plugin` | Off | Skip the first-run trust prompt for a plugin whose code and suite you'd run yourself. Pass it in CI so the job is never refused by or left waiting at the prompt. See [What a run can access](#security) |
377| `--mocks <mode>` | `record` | `record` answers MCP tool calls from [mocks](#mock-mcp-servers), doesn't start the plugin's real servers, and saves agent-mock answers for replay. `off` ignores mocks and starts the plugin's real MCP servers |
378| `--allow-real-servers` | Off | With `--mocks record`, also start the plugin's real MCP servers for servers that have no mock |
379| `--json [path]` | Off | Print the [result document](#json-result) to stdout, or write it to a path ending in `.json`. The run is quiet: no progress lines or summary table |
380| `--output-dir <dir>` | `<eval dir>/results/<timestamp>/` | Where `aggregate-result.json` and `report.html` go |
381| `--no-publish` | | Keep the HTML report local. See [HTML report](#html-report) |
382| `--publish-report` | | Publish the report even where it would stay local by default, such as a run a Claude Code session started |
383| `--keep-temp` | Off | Keep every run's sandbox directory and print its path, for debugging what Claude produced |
365| Option | Default | Effect |
366| :- | :- | :- |
367| `--runs <n>` | Each case's `runs`, else 3 | Runs per case per arm |
368| `-j`, `--concurrency <n>` | `1` | Run up to this many agent runs at once, from 1 to 8. They share your account's rate limit, so this shortens wall-clock time rather than raising throughput past that limit. Results keep case order |
369| `--model <model>` | Each case's `model`, else `ANTHROPIC_MODEL` if set, else Claude Code's default | Model for the agent under test. Pin it in CI so a model rollout isn't mistaken for a plugin regression |
370| `--judge-model <model>` | A small fast model | Model for `llm` and `baseline` graders |
371| `--ablation <mode>` | `with-without` when a plugin resolves, else `none` | Whether to also run each case without the plugin to measure what it adds. `none` runs one arm; `with-without` adds the no-plugin baseline |
372| `--threshold <0..1>` | `1.0` | A case passes when its with-arm score is at least this. Any case below it makes the command exit 1 |
373| `--max-cost-usd <usd>` | No ceiling | A ceiling on the run's list-price cost estimate, not on plan usage. Checked before each run starts. Once spent, nothing further starts; runs that already started finish, so spend can pass the ceiling by those runs. If any run is left unstarted, the command exits 2 with partial results |
374| `--allow-tools <tools...>` | None | Grant tools beyond the read-only set. See [Grant tools](#grant-tools) |
375| `--scaffold` | Off | Run each case's [`scaffold_script`](#add-setup-or-history-with-case-yaml) |
376| `--trust-plugin` | Off | Skip the first-run trust prompt for a plugin whose code and suite you'd run yourself. Pass it in CI so the job is never refused by or left waiting at the prompt. See [What a run can access](#security) |
377| `--mocks <mode>` | `record` | `record` answers MCP tool calls from [mocks](#mock-mcp-servers), doesn't start the plugin's real servers, and saves agent-mock answers for replay. `off` ignores mocks and starts the plugin's real MCP servers |
378| `--allow-real-servers` | Off | With `--mocks record`, also start the plugin's real MCP servers for servers that have no mock |
379| `--json [path]` | Off | Print the [result document](#json-result) to stdout, or write it to a path ending in `.json`. The run is quiet: no progress lines or summary table |
380| `--output-dir <dir>` | `<eval dir>/results/<timestamp>/` | Where `aggregate-result.json` and `report.html` go |
381| `--no-publish` | | Keep the HTML report local. See [HTML report](#html-report) |
382| `--publish-report` | | Publish the report even where it would stay local by default, such as a run a Claude Code session started |
383| `--keep-temp` | Off | Keep every run's sandbox directory and print its path, for debugging what Claude produced |
384384 
385385<h3 id="run-evals-in-ci">
386386 Run evals in CI
from line 401
401401 
402402The job's exit code tells you what happened:
403403 
404| Exit code | Meaning |
405| :-------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
406| 0 | Every case scored at or above `--threshold` and every case file loaded |
407| 1 | A case scored below the threshold, a case file failed to load, no cases were found, a run couldn't be started, the plugin directory isn't trusted and `--trust-plugin` wasn't passed, or an option was invalid |
408| 2 | Partial run: the `--max-cost-usd` ceiling was hit, or your credential was rejected before or at the first run. `results.json` is still written with `partial: true` and the reason |
409| 130 | Interrupted. Partial results are written |
410| 143 | Terminated, such as by a CI timeout |
404| Exit code | Meaning |
405| :- | :- |
406| 0 | Every case scored at or above `--threshold` and every case file loaded |
407| 1 | A case scored below the threshold, a case file failed to load, no cases were found, a run couldn't be started, the plugin directory isn't trusted and `--trust-plugin` wasn't passed, or an option was invalid |
408| 2 | Partial run: the `--max-cost-usd` ceiling was hit, or your credential was rejected before or at the first run. `results.json` is still written with `partial: true` and the reason |
409| 130 | Interrupted. Partial results are written |
410| 143 | Terminated, such as by a CI timeout |
411411 
412412The with-minus-without delta is reported but never changes the exit code, and neither do problems writing or publishing the HTML report.
413413 
from line 448
448448 
449449These are the fields a gating script usually reads. The document also carries the suite configuration, every grader definition, and per-run grader results with explanations and evidence:
450450 
451| Field | Meaning |
452| :------------------------------------------------ | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
453| `partial`, `partialReason` | `true` with `cost_ceiling`, `interrupted`, or `auth_failed` when the suite didn't finish. Leave partial results out of trend charts |
454| `aggregates.overallScore` | Mean case score across the suite |
455| `aggregates.casesPassed`, `aggregates.casesTotal` | Cases at or above `--threshold`, and the total |
456| `aggregates.meanDelta` | Mean `Δ` across cases, under the two-arm mode |
457| `cases[].name` | Case name |
458| `cases[].aggregates.score` | Mean with-arm run score for the case |
459| `cases[].aggregates.delta` | With-arm score minus without-arm score. Omitted when the arms aren't comparable |
460| `cases[].arms.with[].error` | `null`, or why a run ended abnormally, such as `timed out after 300s`. A run that started but ended badly is still graded on what it produced, so a non-null error doesn't imply score 0 |
461| `cases[].arms.with[].aborted` | Present when a [mock](#mock-mcp-servers)'s `expect:` or `abort_when` stopped the run, with `server`, `tool`, and `reason`. The run scores 0 and `error` stays `null` |
462| `cases[].arms.with[].skippedPaidGraders` | `true` when the cost ceiling skipped this run's judge graders, so its score isn't comparable |
463| `costUsd`, `durationSeconds`, `claudeVersion` | Estimated cost at list price including judge calls, wall-clock seconds, and the Claude Code version that ran the suite |
451| Field | Meaning |
452| :- | :- |
453| `partial`, `partialReason` | `true` with `cost_ceiling`, `interrupted`, or `auth_failed` when the suite didn't finish. Leave partial results out of trend charts |
454| `aggregates.overallScore` | Mean case score across the suite |
455| `aggregates.casesPassed`, `aggregates.casesTotal` | Cases at or above `--threshold`, and the total |
456| `aggregates.meanDelta` | Mean `Δ` across cases, under the two-arm mode |
457| `cases[].name` | Case name |
458| `cases[].aggregates.score` | Mean with-arm run score for the case |
459| `cases[].aggregates.delta` | With-arm score minus without-arm score. Omitted when the arms aren't comparable |
460| `cases[].arms.with[].error` | `null`, or why a run ended abnormally, such as `timed out after 300s`. A run that started but ended badly is still graded on what it produced, so a non-null error doesn't imply score 0 |
461| `cases[].arms.with[].aborted` | Present when a [mock](#mock-mcp-servers)'s `expect:` or `abort_when` stopped the run, with `server`, `tool`, and `reason`. The run scores 0 and `error` stays `null` |
462| `cases[].arms.with[].skippedPaidGraders` | `true` when the cost ceiling skipped this run's judge graders, so its score isn't comparable |
463| `costUsd`, `durationSeconds`, `claudeVersion` | Estimated cost at list price including judge calls, wall-clock seconds, and the Claude Code version that ran the suite |
464464 
465465<h2 id="security">
466466 What a run can access
from line 528
528528 
529529`prompt.md` frontmatter accepts these fields. An unknown key is an error:
530530 
531| Field | Default | Purpose |
532| :--------------------- | :--------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
533| `schema_version` | `"1.1"`, set for you | Case format version. Cases written as `prompt.md` get it automatically, so you rarely set it |
534| `name` | The directory name | Case name. `--case` globs match it and the report keys on it |
535| `description` | | For humans. Not used at run time |
536| `tags` | `[]` | Labels for `--tag` filtering. A case runs if any of its tags matches |
537| `plugins` | The nearest enclosing plugin | Plugin directories under test, relative to the case directory. Set `plugins: ["../.."]` when auto-detection doesn't find your plugin; see [the plugin didn't load](#the-baseline-arm-shows-no-plugin-or-delta-is-zero) |
538| `runs` | `3` | Runs per arm, 1 to 50. `--runs` overrides it |
539| `expected_outcome` | | For humans. Not used at run time |
540| `model` | The child session's default | Model for the agent under test. `--model` overrides it |
541| `max_turns` | `10` | Turn cap, up to 200. Hitting it is recorded as a run error and usually lowers the score, so set it generously |
542| `timeout_seconds` | `300` | Wall-clock cap per run, up to 3600 |
543| `allowed_tools` | `[]` | Tools the case wants, such as `[Read, Glob, Grep, Skill]`. Read-only tools are granted when listed here; for anything else, see [Grant tools](#grant-tools) |
544| `append_system_prompt` | | Text appended to the child session's system prompt |
545| `env` | `{}` | Extra environment variables for the child session. Keys must match `EVAL_[A-Z0-9_]*`; any other key fails the run. The run inherits only an allowlist from your shell: basics such as `PATH` and locale, proxy and certificate settings, the variables that select and authenticate your model provider, most `ANTHROPIC_*` and `CLAUDE_CODE_*` configuration, and `EVAL_*`. To pass the plugin anything else, such as a toolchain setting, export it as an `EVAL_*` variable |
531| Field | Default | Purpose |
532| :- | :- | :- |
533| `schema_version` | `"1.1"`, set for you | Case format version. Cases written as `prompt.md` get it automatically, so you rarely set it |
534| `name` | The directory name | Case name. `--case` globs match it and the report keys on it |
535| `description` | | For humans. Not used at run time |
536| `tags` | `[]` | Labels for `--tag` filtering. A case runs if any of its tags matches |
537| `plugins` | The nearest enclosing plugin | Plugin directories under test, relative to the case directory. Set `plugins: ["../.."]` when auto-detection doesn't find your plugin; see [the plugin didn't load](#the-baseline-arm-shows-no-plugin-or-delta-is-zero) |
538| `runs` | `3` | Runs per arm, 1 to 50. `--runs` overrides it |
539| `expected_outcome` | | For humans. Not used at run time |
540| `model` | The child session's default | Model for the agent under test. `--model` overrides it |
541| `max_turns` | `10` | Turn cap, up to 200. Hitting it is recorded as a run error and usually lowers the score, so set it generously |
542| `timeout_seconds` | `300` | Wall-clock cap per run, up to 3600 |
543| `allowed_tools` | `[]` | Tools the case wants, such as `[Read, Glob, Grep, Skill]`. Read-only tools are granted when listed here; for anything else, see [Grant tools](#grant-tools) |
544| `append_system_prompt` | | Text appended to the child session's system prompt |
545| `env` | `{}` | Extra environment variables for the child session. Keys must match `EVAL_[A-Z0-9_]*`; any other key fails the run. The run inherits only an allowlist from your shell: basics such as `PATH` and locale, proxy and certificate settings, the variables that select and authenticate your model provider, most `ANTHROPIC_*` and `CLAUDE_CODE_*` configuration, and `EVAL_*`. To pass the plugin anything else, such as a toolchain setting, export it as an `EVAL_*` variable |
546546 
547547<h3 id="case-yaml-fields">
548548 case.yaml fields
from line 552
552552 
553553These fields exist only in `case.yaml`:
554554 
555| Field | Purpose |
556| :------------------------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
555| Field | Purpose |
556| :- | :- |
557557| `context.scaffold_script` | A Bash script in the case directory that runs in the empty workspace before Claude starts, to create fixture files or a git repository. It runs only when you pass [`--scaffold`](#add-setup-or-history-with-case-yaml) |
558| `context.history_file` | A `.jsonl` transcript in the case directory to resume. The case's prompt becomes the next user turn |
559| `context.add_dirs` | Directories inside the case directory that Claude may read during the run, granted read-only |
560| `execution.prompt` | The prompt, when you keep the whole case in `case.yaml` and omit `prompt.md` |
561| `graders` | A list of graders, each with a `name` plus the same keys a `graders/*.md` file takes in frontmatter. For `llm` graders, put the rubric in `criteria` |
558| `context.history_file` | A `.jsonl` transcript in the case directory to resume. The case's prompt becomes the next user turn |
559| `context.add_dirs` | Directories inside the case directory that Claude may read during the run, granted read-only |
560| `execution.prompt` | The prompt, when you keep the whole case in `case.yaml` and omit `prompt.md` |
561| `graders` | A list of graders, each with a `name` plus the same keys a `graders/*.md` file takes in frontmatter. For `llm` graders, put the rubric in `criteria` |
562562 
563563### Grader frontmatter
564564 
565565Every grader file under `graders/` takes these keys in frontmatter, plus the options for its type. The grader's name is the filename without `.md`:
566566 
567| Key | Default | Purpose |
568| :------- | :------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
569| `type` | required | One of the [grader types](#grader-types) |
570| `weight` | `1` | Relative weight in the run's score. Any positive number |
571| `arm` | unset | `with-only` excludes the grader from scoring in a [two-arm run](#compare-against-a-no-plugin-baseline); `both` forces a grader Claude Code would otherwise exclude to be scored in both arms |
567| Key | Default | Purpose |
568| :- | :- | :- |
569| `type` | required | One of the [grader types](#grader-types) |
570| `weight` | `1` | Relative weight in the run's score. Any positive number |
571| `arm` | unset | `with-only` excludes the grader from scoring in a [two-arm run](#compare-against-a-no-plugin-baseline); `both` forces a grader Claude Code would otherwise exclude to be scored in both arms |
572572 
573573#### What a grader can look at
574574 
575575`regex` graders take a `target` and `llm` graders take a `focus`. Both accept the same values:
576576 
577| Value | What the grader sees |
578| :------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
579| `last_message` | Claude's final response text. This is the default |
580| `trace` | The session as JSON, one message per line. A `regex` grader sees every message; an `llm` judge sees the first 12 and the last 12. Quotes and newlines inside it are JSON-escaped, so a regex matches `\"` rather than `"` |
581| `files` | The list of paths Claude created during the run, one per line. Not their contents, and not files that a scaffold created or that Claude only modified |
577| Value | What the grader sees |
578| :- | :- |
579| `last_message` | Claude's final response text. This is the default |
580| `trace` | The session as JSON, one message per line. A `regex` grader sees every message; an `llm` judge sees the first 12 and the last 12. Quotes and newlines inside it are JSON-escaped, so a regex matches `\"` rather than `"` |
581| `files` | The list of paths Claude created during the run, one per line. Not their contents, and not files that a scaffold created or that Claude only modified |
582582| `{ source: file, path: <path> }` | The contents of one file in the workspace after the run. Use this to grade what the plugin produced. A PNG, JPEG, GIF, or WebP file is shown to an `llm` judge as an image. An `llm` judge refuses other binary files such as `.pptx` or PDF; render them to an image or write them out as text and grade that |
583| `mock_calls` | Each call Claude made to a [mocked MCP tool](#mock-mcp-servers), with its input and the mock's answer |
583| `mock_calls` | Each call Claude made to a [mocked MCP tool](#mock-mcp-servers), with its input and the mock's answer |
584584 
585585#### Grader types
586586 
587587Each grader type below lists its options and when it passes:
588588 
589| Type | Options | Passes when |
590| :------------ | :------------------------------------ | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
591| `regex` | `pattern`, `flags`, `match`, `target` | The JavaScript regex `pattern` is found in the target. Set `match: not_contains` to require absence or `match: "count:N"` to require exactly N matches. Put case-insensitivity in `flags: i`; inline `(?i)` isn't supported |
592| `tool_used` | `tool`, `input_match`, `min`, `max` | The number of calls to `tool` whose JSON-encoded input matches the optional `input_match` regex is between `min`, default 1, and `max`, default unlimited. To assert a tool was never called, set both `min: 0` and `max: 0` |
593| `tool_order` | `before`, `after` | Both tools were called and the first matching `before` call precedes the first matching `after` call. Each is a tool name or `{ tool, input_match }` |
594| `file_exists` | `path`, `exists` | A file Claude created matches the `path` glob, or none does with `exists: false`. Only files created during the run count |
595| `llm` | `criteria`, `focus` | A judge model votes PASS on the rubric in at least two of three votes. In the `.md` layout the file body is the criteria |
596| `baseline` | `baseline_file`, `criteria` | A judge finds the run satisfies the criteria at least as well as the reference transcript at `baseline_file`, a `.jsonl` in the case directory |
589| Type | Options | Passes when |
590| :- | :- | :- |
591| `regex` | `pattern`, `flags`, `match`, `target` | The JavaScript regex `pattern` is found in the target. Set `match: not_contains` to require absence or `match: "count:N"` to require exactly N matches. Put case-insensitivity in `flags: i`; inline `(?i)` isn't supported |
592| `tool_used` | `tool`, `input_match`, `min`, `max` | The number of calls to `tool` whose JSON-encoded input matches the optional `input_match` regex is between `min`, default 1, and `max`, default unlimited. To assert a tool was never called, set both `min: 0` and `max: 0` |
593| `tool_order` | `before`, `after` | Both tools were called and the first matching `before` call precedes the first matching `after` call. Each is a tool name or `{ tool, input_match }` |
594| `file_exists` | `path`, `exists` | A file Claude created matches the `path` glob, or none does with `exists: false`. Only files created during the run count |
595| `llm` | `criteria`, `focus` | A judge model votes PASS on the rubric in at least two of three votes. In the `.md` layout the file body is the criteria |
596| `baseline` | `baseline_file`, `criteria` | A judge finds the run satisfies the criteria at least as well as the reference transcript at `baseline_file`, a `.jsonl` in the case directory |
597597 
598598<h3 id="mock-files">
599599 Mock files
from line 601
601601 
602602A `<tool>.md` file under `mocks/<server>/` answers one tool. Its body is the tool result, with `{{input.<field>}}` and `{{file:fixtures/<name>}}` substitutions. Its frontmatter accepts these keys:
603603 
604| Key | Default | Purpose |
605| :----------- | :------ | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
606| `type` | `fixed` | `fixed` returns the body as written. `agent` treats the body as instructions for a small model that acts as the server for the run and sees earlier calls as history |
607| `expect` | unset | A map from dotted input paths to a type name such as `string`, `number`, `boolean`, `array`, or `object`, a `/regex/`, a literal, or a list of allowed literals. A call that violates it aborts the run with score 0 and is reported as `aborted` with the server, tool, and reason |
608| `error` | `false` | `fixed` only. Return the body as a tool error |
609| `abort_when` | unset | `agent` only. Prose listing the only conditions under which the agent may abort the run |
604| Key | Default | Purpose |
605| :- | :- | :- |
606| `type` | `fixed` | `fixed` returns the body as written. `agent` treats the body as instructions for a small model that acts as the server for the run and sees earlier calls as history |
607| `expect` | unset | A map from dotted input paths to a type name such as `string`, `number`, `boolean`, `array`, or `object`, a `/regex/`, a literal, or a list of allowed literals. A call that violates it aborts the run with score 0 and is reported as `aborted` with the server, tool, and reason |
608| `error` | `false` | `fixed` only. Return the body as a tool error |
609| `abort_when` | unset | `agent` only. Prose listing the only conditions under which the agent may abort the run |
610610 
611611Two optional files sit beside the tool files in a server's directory:
612612 
Feedback