Test plugins with evals
Write eval cases for a Claude Code plugin, run them with claude plugin eval, compare against a no-plugin baseline, mock MCP servers and gate CI on the score.
A plugin that loads cleanly can still do nothing useful. Maybe Claude never picks your skill on natural phrasing, or it would have produced the same answer without it. claude plugin eval answers that: it runs realistic prompts against your plugin, checks the results with graders, and compares the score against the same prompts with no plugin loaded.
Use it to:
- measure how reliably the plugin produces the outcome you want;
- catch regressions when you change the plugin or a new model ships;
- see what the plugin actually contributes over plain Claude.
Note: Every eval run and every judge call is a real model request on your account, counted against your plan or API bill. Check the requirements, then start small.
Checking files for schema errors is a different job: that is claude plugin validate (see /docs/plugins/create). And the skill-creator plugin has its own evals/evals.json format for iterating on a single skill in conversation; the two tools do not read each other's files.
Requirements
- Claude Code v2.1.269 or later (
claude --version,claude update). - Git 2.31 or later if git is installed. With older git, the command stops before running anything. With no git at all, it runs normally.
- A plugin directory with
plugin.jsonor.claude-plugin/plugin.json, or a skills-directory plugin. - The same authentication and provider as your normal sessions. On Bedrock, Vertex or Foundry, run from a shell that exports the same provider variables, since runs inherit them. Reported costs are list-price estimates.
How it works
Cases, runs and scores
The suite lives in evals/ inside your plugin. Each case is a subfolder with a prompt and one or more graders. For each run of a case, Claude Code starts a fresh, isolated non-interactive session with only your plugin loaded, sends the prompt, lets Claude work to completion or to the case's turn or time limit, and then each grader passes or fails on the final reply, the transcript or a file Claude created.
Agents are non-deterministic, so each case runs three times by default. A run's score is the fraction of graders that passed (weighted if you set weights). The case score is the mean across runs. A case passes if it meets --threshold, which defaults to 1.0.
The baseline
A high score alone proves nothing: Claude might ace the task without your plugin. So by default each case also runs the same number of times with no plugin. You get WITH, W/OUT and their difference Δ. A case that scores 1.0 in both arms is not being helped by your plugin.
Rough cost: cases × runs with the plugin, the same again without, plus three short judge calls per llm or baseline grader per run.
Your first suite
Say the plugin has a release-notes skill that turns a list of merged PRs into customer-facing notes. Open a terminal at the plugin root.
1. Let Claude write the cases
claude plugin eval init
If the directory is not trusted yet you get Trust this plugin directory?; answer y. An interactive session opens in which Claude reads the plugin, asks what a good result looks like, proposes prompts that should and should not trigger it, designs graders, trial-runs them once, and writes one case folder per prompt under evals/. Exit with /exit or Ctrl+D when it says the suite is ready.
You can also ask Claude to run claude plugin eval init from a session already open at the plugin root; it will ask the same questions there.
2. Run it
claude plugin eval .
Each case runs three times with and three without, so six runs per case. A progress line prints per run with its score and each grader's verdict.
3. Read the summary
CASE WITH W/OUT Δ RUNS COST NOTES
notes-from-pr-list 1.00 0.33 +0.67 6 $0.38
ignores-unrelated-request 1.00 1.00 +0.00 6 $0.12
2 case(s) · mean Δ +0.33 · 81s · $0.50
Report: /Users/cam/plugins/release-notes/evals/results/2026-10-08T09-14-03-210Z/report.html
COST is a list-price estimate. NOTES shows the highest-weight failing grader's explanation (or the run error) from the with-plugin arm. If your account can publish reports you also get a Published: line with a URL.
4. Iterate
Open the report. The most common first finding is Δ near zero with the tool_used: Skill grader failing: Claude is not choosing your skill. Rewrite the skill's description (see /docs/skills) and run again.
For quick iteration on one case, run one arm once, then confirm at the default three runs before you believe it:
claude plugin eval . --case notes-from-pr-list --runs 1 --ablation none
With a single arm the table shows SCORE and PASS% instead of WITH, W/OUT and Δ.
Writing cases by hand
Everything init writes is plain files. A case is a folder containing prompt.md, case.yaml or both, with at least one grader (a case without graders fails to load). To group cases, nest them under a folder that is not itself a case.
release-notes/
├── .claude-plugin/plugin.json
├── skills/release-notes/SKILL.md
└── evals/
├── notes-from-pr-list/
│ ├── prompt.md
│ ├── graders/
│ │ ├── customer-tone.md
│ │ └── skill-fired.md
│ └── case.yaml # only if you need context.* fields
├── ignores-unrelated-request/
└── results/ # written by runs; add to .gitignore
Get a blank template without running anything:
claude plugin eval init --bare notes-from-pr-list
The prompt
Write it the way a user would, without naming the skill:
---
max_turns: 8
allowed_tools: [Read, Skill]
tags: [smoke]
---
Turn these into release notes our customers can read:
- #812 fix: retry webhook delivery on 502
- #815 feat: CSV export for invoices
- #817 chore: bump eslint
Each run starts in an empty working directory, so include what the task needs in the prompt or seed the workspace. @path mentions are not expanded; if Claude must read a file, allow a tool for it.
The graders
A rubric for the judge model, written as concrete PASS and FAIL conditions, in graders/customer-tone.md:
---
type: llm
---
PASS if the notes mention CSV export for invoices and more reliable webhooks in plain language, and omit the eslint bump.
FAIL if internal PR numbers or the word "chore" appear, or either user-facing change is missing.
And proof that your skill did the work, in graders/skill-fired.md:
---
type: tool_used
tool: Skill
input_match: '"skill"\s*:\s*"(?:[\w-]+:)?release-notes"'
---
This passes if Claude invoked the skill at least once, including via its namespaced plugin:skill form.
Choosing graders that give a stable signal
There are six grader types. regex, tool_used, tool_order and file_exists are computed from the transcript and files and cost nothing. llm and baseline call a judge model. There are no custom-code graders. The judge defaults to the model Claude Code uses for background tasks; override with --judge-model sonnet or a full model ID.
Habits that keep scores trustworthy:
- Long output, use regex. A judge reading a long file gets flaky. Grade generated files with a
regexon their contents; keepllmfor short outputs with crisp PASS/FAIL rubrics. - One outcome grader, one process grader. Pair a check on the result with a
tool_usedortool_ordercheck on how it was produced. Together they tell you whether it was right and whether your plugin did it. - Negative Δ with the skill firing? Suspect the judge. A small judge may fail a correct answer for formatting. Re-run with
--judge-model sonnetand tighten the rubric so formatting does not decide the verdict. - Verifying a build or test inside the run: have the prompt ask Claude to run it and write the outcome to a file, grade that file, and assert the command ran with
tool_usedplus aninput_matchnaming it.
Scoring against the baseline
In a two-arm run, some graders would always fail without the plugin (a "skill was invoked" check, for instance), which would drag W/OUT towards zero and inflate Δ. So Claude Code excludes these from the score in both arms and reports them in the with-arm as indicators only, with scored: false:
- every
tool_usedgrader withtool: Skill; - every
regexgrader withtarget: mock_callsand everyllmgrader withfocus: mock_calls, when each mocked server in the case is one your plugin declares; - any grader marked
arm: with-only.
Overrides: if every grader in a case would be excluded, they are scored normally. arm: both forces scoring in both arms, which is what you want for a "must not invoke the skill" check (min: 0, max: 0). Under --ablation none nothing is excluded, so absolute scores differ between modes.
A case runs the with-arm only (no W/OUT, no Δ) when you pass --ablation none; when it resumes a transcript via context.history_file and the target is a path (a single-arm (no Δ) notice prints on stderr; pass --ablation with-without to force both); or when no plugin could be found for it.
A different eval folder
If evals/ is taken, add "experimental": { "evals": "quality/evals" } to plugin.json, or pass --eval-dir quality/evals to both claude plugin eval and init. The flag wins over the manifest. Paths must be relative, plain folder names: absolute paths or .. are an error as a flag, and as a manifest value print a Warning: and fall back to evals/.
Fixtures and mocks
Seeding the workspace or conversation
Add a case.yaml next to prompt.md with a context block:
schema_version: "1.1"
name: notes-from-git-log
tags: [nightly]
context:
scaffold_script: make-repo.sh
add_dirs: [samples]
scaffold_script: a Bash script in the case folder that builds fixture files or a git repo in the empty workspace. It runs as you, outside the sandbox, only when you pass--scaffold. It gets a minimal environment: yourPATH,HOMEset to the run's temporary home,TMPDIR, and a few constants likeTERM=dumb; not your other shell variables and not the case'sEVAL_*variables. Non-zero exit or more than 120 seconds scores the run 0 withscaffold failed. Project config it writes is not loaded, so use it for files and git state only.history_file: a.jsonltranscript to resume; the prompt becomes the next user turn.add_dirs: folders inside the case that Claude may read, read-only.
#!/bin/bash
# make-repo.sh
git init -q && git config user.email t@example.com && git config user.name test
echo "v1" > app.txt && git add . && git commit -qm "feat: initial invoices page"
echo "v2" > app.txt && git commit -qam "fix: rounding on VAT totals"
Mocking MCP servers
You can evaluate skills that call MCP tools without the real service. Put one Markdown file per tool at evals/mocks/<server>/<tool>.md (suite-wide) or in a case's own mocks/ folder, where <server> is the server name from your plugin's MCP config.
Runs never start your real servers unless you ask. Claude Code registers a stand-in under each server's name: mocked tools answer from their files and need no grant; tools with no mock are unavailable. A server with no mocks shows in the mocked: progress line as plugin_<plugin>_<server>[not started: no mock].
The file body is what the tool returns. For a post_message tool on a chat server:
---
expect:
channel: /^#[a-z-]+$/
text: string
---
Posted to {{input.channel}} (ts 1728380000.000100)
Options:
{{input.<field>}}inserts call input;{{file:fixtures/{input.<field>}.json}}inserts a fixture file beside the mock.expect:guards the input. A violating call aborts the run with score 0 and records why, so you can assert what your plugin asked for.error: truereturns the body as a tool error.type: agentmakes the judge model act as the server, following instructions in the body. All agent-mock calls in a run share a budget of four timesmax_turns; exceeding it aborts with score 0.
Grade the calls themselves with target: mock_calls.
To hit real servers instead, use --allow-real-servers (start real servers only for unmocked ones) or --mocks off (ignore mocks, start everything). Either way those processes run as you outside the sandbox and their tools need an --allow-tools grant.
Replaying agent mocks. Agent-mock answers vary run to run. After a clean run, Claude Code saves each answer under the results folder in mock-recordings/. ADOPT.txt there lists each recording and the .replay/<server>/ folder to copy it into. Once copied, identical calls are answered from the recording with no model call. Commit mocks/.replay/ so CI is repeatable.
Running evals
Targets
| Target | Runs |
|---|---|
A plugin root, such as . | Every case in its eval folder, with that plugin |
One prompt.md or case.yaml | That case, with its enclosing plugin |
An installed plugin, name or name@marketplace | Cases in the installed copy, with that copy loaded; results go to ./evals/results/ in your current folder (or ./<dir>/results/ with --eval-dir) |
name@skills-dir | Same, for a skills-directory plugin |
| Nothing | The current folder as a path |
Filter with --case <glob> and --tag <tag>. Put the target before --tag, --allow-tools and --json, because those flags would swallow it as a value.
Granting tools
Runs never pause for permission. Tools that need a grant you did not give (Bash, Write, Edit, WebFetch, WebSearch) are removed entirely.
A case may use the read-only tools it lists in allowed_tools, from: Read, Glob, Grep, NotebookRead, Skill, AskUserQuestion, Agent, TodoWrite, TaskCreate, TaskGet, TaskList, TaskUpdate, TaskStop. Anything else you grant for the whole run:
claude plugin eval . --allow-tools Write Edit "Bash(pnpm test *)"
Ungranted requests show as not granted in progress output. Real plugin MCP tools need the server started and a grant such as --allow-tools "mcp__plugin_release-notes_chat__*".
Granting Bash in any form runs every command under the OS-level sandbox: writes confined to the workspace, home and Claude Code config unreadable, network limited to domains you grant with --allow-tools "WebFetch(domain:example.com)". With no sandbox backend, shell-granting runs are refused (usually scoring 0). Native Windows has none, so use WSL2; on Linux install bubblewrap and socat. See /docs/sandboxing.
Options
claude plugin eval --help has the full list, including --case, --tag, --eval-dir, --no-scaffold, --report and --verbose.
| Option | Default | What it does |
|---|---|---|
--runs <n> | Case's runs, else 3 | Runs per case per arm |
-j, --concurrency <n> | 1 | 1 to 8 parallel runs; shares your rate limit; results keep case order |
--model <model> | Case model, else ANTHROPIC_MODEL, else default | Agent under test. Pin it in CI |
--judge-model <model> | Background-task model | Judge for llm and baseline |
--ablation <mode> | Per case | none for one arm, with-without for the baseline too |
--threshold <0..1> | 1.0 | Minimum with-arm score; any case below makes exit 1 |
--max-cost-usd <usd> | None | Ceiling on estimated list-price cost, checked before each run starts. Started runs finish, so spend can overshoot. Unstarted runs mean exit 2 |
--allow-tools <tools...> | None | Extra tool grants |
--scaffold | Off | Run scaffold_scripts |
--trust-plugin | Off | Skip the first-run trust prompt; use in CI |
--mocks <mode> | record | record uses mocks and saves agent-mock answers; off starts real servers |
--allow-real-servers | Off | With record, also start real servers that have no mock |
--json [path] | Off | Result document to stdout or a .json path; suppresses progress and table |
--output-dir <dir> | <eval dir>/results/<timestamp>/ | Where results go |
--no-publish | Keep the HTML report local | |
--publish-report | Publish even where it would stay local by default | |
--keep-temp | Off | Keep each run's sandbox folder and print its path |
In CI
claude plugin eval . \
--trust-plugin \
--json eval-results.json \
--threshold 0.75 \
--model claude-sonnet-5 \
--judge-model claude-haiku-4-5 \
--no-publish \
--max-cost-usd 15
Pin both models so a model rollout is not mistaken for a plugin regression.
| Exit | Meaning |
|---|---|
| 0 | All cases at or above threshold and all case files loaded |
| 1 | A case below threshold, a case failed to load, no cases found, a run could not start, untrusted directory without --trust-plugin, or a bad option |
| 2 | Partial: cost ceiling hit, or credentials rejected at or before the first run. JSON still written with partial: true |
| 130 | Interrupted; partial results written |
| 143 | Terminated (for example a CI timeout) |
Δ never affects the exit code, nor do report write or publish failures. The runner needs Claude Code installed and credentials such as ANTHROPIC_API_KEY (see /docs/authentication). eval init needs a terminal, so in CI use init --bare <name>.
To keep spend predictable on every-push suites, I use only free graders, --ablation none, and keep the judge-based, two-arm suite for a nightly job. Exclude partial: true results and runs with skippedPaidGraders from any trend chart.
Reading results
Each run with at least one case writes results/<timestamp>/ with aggregate-result.json and report.html (under the plugin for path targets, under your current folder for named targets).
The HTML report
A single self-contained file with no external requests, fine to attach to a CI job. Read it top down:
- Verdict line and tiles: suite score (mean with-plugin case score), ablation Δ against the baseline, cases meeting the threshold, and perfect runs (share of with-plugin runs where every grader passed).
- Case cards: each case's
Δand score with a threshold tick. NegativeΔcases get a red left edge. - Inside a case: with-plugin runs first, then baseline. Failed graders are pre-expanded with their explanation;
llmgraders show the judge's votes and the evidence it saw. Unscored graders carry aplugin-fired indicatorbadge. - Prompt and Graders: the case's prompt and every rubric or pattern, so readers do not need the suite.
Signed in with a claude.ai subscription and with artifacts available, the report is also published as a private artifact (Published: <url>). --no-publish keeps it local. Runs started by a Claude Code session stay local (kept local) unless you add --publish-report.
The JSON document
schemaVersion: 1, camelCase fields, new fields added without renames, so ignore unknown keys. Fields gating scripts usually read:
| Field | Meaning |
|---|---|
partial, partialReason | true with cost_ceiling, interrupted or auth_failed if unfinished |
aggregates.overallScore | Mean case score |
aggregates.casesPassed, aggregates.casesTotal | Cases at threshold, and total |
aggregates.meanDelta | Mean Δ in two-arm mode |
cases[].name | Case name |
cases[].aggregates.score | Mean with-arm score |
cases[].aggregates.delta | With minus without; omitted if single-arm or not comparable |
cases[].arms.with[].error | null or why the run ended abnormally (still graded on what it produced) |
cases[].arms.with[].aborted | Set when a mock stopped the run via expect:, abort_when or the budget; includes server, tool, reason; scores 0 with error null |
cases[].arms.with[].skippedPaidGraders | Judge graders skipped by the cost ceiling |
costUsd, durationSeconds, claudeVersion | Estimated cost, wall time, version |
A tiny gate on top of the exit code:
jq -e '.partial == false and .aggregates.meanDelta > 0.2' eval-results.json
What a run can access
claude plugin eval loads the plugin's skills, hooks and agents and runs on your machine as you. Pointing it at a plugin is the same trust decision as claude --plugin-dir. The isolation below limits the agent under test, not the plugin's own code, and a passing suite says nothing about whether a plugin is safe.
Trust
First run against a directory asks Trust this plugin directory? unless you already accepted trust there interactively. Inside a git repo, yes trusts the whole repo, interactive sessions included. When stdin or stdout is not a terminal, or with --json, it cannot ask and exits 1; --trust-plugin asserts trust. Named targets (installed or skills-directory plugins) skip the prompt.
Only the matching flag enables: scaffold_script (--scaffold), non-read-only tools (--allow-tools), real MCP servers (--allow-real-servers or --mocks off). Neither a case's allowed_tools nor a skill's allowed-tools can widen these. If the plugin has hooks you did not write or you start real servers, treat scores as advisory unless run in a container or CI runner, since those run outside the sandbox and could tamper with what graders read.
Isolation
Each run gets a temporary home, working directory and Claude Code config, and runs as a claude -p child with only your plugin.
- Nothing personal or project-level loads: no user settings, hooks,
CLAUDE.md, MCP servers, other plugins, memory or skills; no.claude/,CLAUDE.mdor.mcp.jsonfrom the workspace or above it, even one a scaffold wrote. Most of your environment is withheld. Ship everything a case depends on inside the plugin. - Managed policy still applies, so managed machines can score differently.
- The Artifact tool is off. Grade what a skill produces before it would publish.
- Case definitions are hidden: Claude cannot read the eval folder.
- No network sandbox outside shell commands:
WebFetch(domain:...)grants reach the domain directly, and hooks and real servers can reach anything.
Suite reference
evals/
├── <case>/
│ ├── prompt.md # frontmatter: case/run fields; body: prompt
│ ├── case.yaml # optional: context.*, or the whole case
│ ├── graders/<name>.md # one grader per file
│ ├── mocks/ # optional per-case mocks
│ └── <fixtures, scripts, transcripts>
├── mocks/<server>/
│ ├── <tool>.md # one mocked tool
│ ├── _server.md # optional agent answering several tools
│ ├── _tools.json # optional saved tools/list response
│ └── fixtures/
├── mocks/.replay/<server>/ # adopted agent-mock recordings
└── results/<timestamp>/
├── aggregate-result.json
├── report.html
└── mock-recordings/ # with ADOPT.txt
prompt.md frontmatter
Unknown keys are an error.
| Field | Default | Purpose |
|---|---|---|
schema_version | "1.1" (automatic) | Case format version |
name | Folder name | Matched by --case; report key |
description | Human note | |
tags | [] | For --tag; any match runs the case |
plugins | Nearest enclosing plugin | Plugin folders under test, relative to the case. Use ["../.."] if detection fails |
runs | 3 | 1 to 50 per arm |
expected_outcome | Human note | |
model | Child session default | Agent model |
max_turns | 10 | Up to 200; hitting it is a run error |
timeout_seconds | 300 | Up to 3600 |
allowed_tools | [] | Tools the case wants |
append_system_prompt | Appended to the child's system prompt | |
env | {} | Extra variables; keys must match EVAL_[A-Z0-9_]*. Runs inherit only an allowlist (PATH, locale, proxy and certificate settings, provider and auth variables, most ANTHROPIC_* and CLAUDE_CODE_*, and EVAL_*) |
case.yaml
Requires schema_version: "1.1" and name. description, tags, plugins, runs and expected_outcome sit at top level; model, max_turns, timeout_seconds, allowed_tools, append_system_prompt and env go under execution:. When both files exist, prompt.md frontmatter overrides matching fields, its body is the prompt, and graders/*.md are appended after graders listed in case.yaml.
| Field (case.yaml only) | Purpose |
|---|---|
context.scaffold_script | Workspace setup script (needs --scaffold, 120 s limit) |
context.history_file | .jsonl transcript to resume |
context.add_dirs | Read-only folders inside the case |
execution.prompt | The prompt, if there is no prompt.md |
graders | List of graders, each with name plus the usual keys; llm rubric goes in criteria |
Grader keys and types
Every grader takes type (required), weight (default 1, any positive number) and arm (with-only or both). The grader's name is its file name without .md.
regex uses target and llm uses focus, with the same values:
| Value | Sees |
|---|---|
last_message | Final reply (default) |
trace | Session as JSON lines. Regex sees all; a judge sees the first 12 and last 12. Quotes appear escaped as \" |
files | Paths Claude created, not contents, not edited or scaffolded files |
{ source: file, path: <path> } | One file's contents after the run. PNG, JPEG, GIF and WebP go to the judge as images; other binaries are refused |
mock_calls | Each mocked tool call with input and answer |
| Type | Options | Passes when |
|---|---|---|
regex | pattern, flags, match, target | JavaScript regex found. match: not_contains for absence, match: "count:N" for exactly N. Use flags: i, not (?i) |
tool_used | tool, input_match, min, max | Matching calls between min (default 1) and max (unlimited). min: 0 and max: 0 asserts never |
tool_order | before, after | Both called and first before precedes first after. Each a name or { tool, input_match } |
file_exists | path, exists | A created file matches the glob, or none with exists: false |
llm | criteria, focus | Judge votes PASS in at least two of three |
baseline | baseline_file, criteria | Judge finds the run at least as good as the reference .jsonl |
Mock files
| Key | Default | Purpose |
|---|---|---|
type | fixed | fixed returns the body; agent uses it as instructions for the judge acting as the server, with earlier calls as history |
expect | Map of dotted input paths to a type (string, number, boolean, array, object), a /regex/, a literal, or a list of literals. Violations abort with score 0 | |
error | false | fixed only: return as a tool error |
abort_when | agent only: prose listing the only conditions to abort |
_server.md is one type: agent mock answering the tools in its tools: key; a <tool>.md for the same tool wins. An expect: there is a load error unless tools: lists one tool. _tools.json is a saved tools/list response so mocks carry real descriptions and schemas. Per-case mocks/ overrides suite mocks file by file.
/regex/ in expect: is a restricted dialect: literals, ., escapes like \d, classes like [a-z]; *, +, ? and {m,n} on a single item; optional leading ^ and trailing $; flags i and s only. Groups, alternation, backreferences and lookaround stop the case loading (score 0). Use a list of literals instead of alternation. Values are checked only up to a maximum length; anchor with ^ and keep quantifiers few.
Troubleshooting
| Symptom | Fix |
|---|---|
plugin eval is currently in early access | Your build is older than general availability. claude update, fresh session |
plugin eval is currently unavailable | Switched off server-side. Update and try later |
is not a trusted plugin directory, and this run cannot stop to ask you about it | Run once in a terminal and accept, or pass --trust-plugin |
is too old for claude plugin eval | Git below 2.31 cannot honour the config that disables repo hooks and helpers. Upgrade git. (Before v2.1.283 the version was not checked) |
No eval cases found | Run from the plugin root, check filters, or run init |
No W/OUT column, or ablation requested but no plugin resolved | Add plugins: ["../.."] to the case. Expected for history_file cases |
Δ near zero, skill grader failing | Real finding: improve the skill's description |
Agent type '<plugin>:<agent>' not found in baseline runs | Expected: the baseline has no plugin. Use --ablation none to skip it |
| Everything scores 0 though files exist | You targeted files (paths) instead of { source: file, path: ... }. file_exists ignores edited or scaffolded files |
| Regex over trace misses visible text | Default target is last_message; trace is JSON-escaped; use flags: i |
| Tools denied or MCP tools missing | Grant with --allow-tools; start real servers explicitly or mock them |
| Exit 1 but results look fine | Default threshold is 1.0; set your own. Also check stderr for load errors |
--json output path must end in .json | Target went after --json; put it first |
Grader passed: false in a 1.0 run | It is unscored by design (scored: false) |
| Usage or rate limit errors mid-suite | Later runs score ~0 but the suite is not partial. Check NOTES or error, re-run after reset |
mock call budget exceeded | Agent mocks share 4 × max_turns calls (replays count). Raise max_turns |
| Timeouts or turn cap | Raise max_turns and timeout_seconds; control spend with --max-cost-usd |