Skip to content

Test plugins with evals

Write eval cases for a Claude Code plugin, run them with claude plugin eval, compare against a no-plugin baseline, mock MCP servers and gate CI on the score.

A plugin that loads cleanly can still do nothing useful. Maybe Claude never picks your skill on natural phrasing, or it would have produced the same answer without it. claude plugin eval answers that: it runs realistic prompts against your plugin, checks the results with graders, and compares the score against the same prompts with no plugin loaded.

Use it to:

  • measure how reliably the plugin produces the outcome you want;
  • catch regressions when you change the plugin or a new model ships;
  • see what the plugin actually contributes over plain Claude.

Note: Every eval run and every judge call is a real model request on your account, counted against your plan or API bill. Check the requirements, then start small.

Checking files for schema errors is a different job: that is claude plugin validate (see /docs/plugins/create). And the skill-creator plugin has its own evals/evals.json format for iterating on a single skill in conversation; the two tools do not read each other's files.

Requirements

  • Claude Code v2.1.269 or later (claude --version, claude update).
  • Git 2.31 or later if git is installed. With older git, the command stops before running anything. With no git at all, it runs normally.
  • A plugin directory with plugin.json or .claude-plugin/plugin.json, or a skills-directory plugin.
  • The same authentication and provider as your normal sessions. On Bedrock, Vertex or Foundry, run from a shell that exports the same provider variables, since runs inherit them. Reported costs are list-price estimates.

How it works

Cases, runs and scores

The suite lives in evals/ inside your plugin. Each case is a subfolder with a prompt and one or more graders. For each run of a case, Claude Code starts a fresh, isolated non-interactive session with only your plugin loaded, sends the prompt, lets Claude work to completion or to the case's turn or time limit, and then each grader passes or fails on the final reply, the transcript or a file Claude created.

Agents are non-deterministic, so each case runs three times by default. A run's score is the fraction of graders that passed (weighted if you set weights). The case score is the mean across runs. A case passes if it meets --threshold, which defaults to 1.0.

The baseline

A high score alone proves nothing: Claude might ace the task without your plugin. So by default each case also runs the same number of times with no plugin. You get WITH, W/OUT and their difference Δ. A case that scores 1.0 in both arms is not being helped by your plugin.

Rough cost: cases × runs with the plugin, the same again without, plus three short judge calls per llm or baseline grader per run.

Your first suite

Say the plugin has a release-notes skill that turns a list of merged PRs into customer-facing notes. Open a terminal at the plugin root.

1. Let Claude write the cases

claude plugin eval init

If the directory is not trusted yet you get Trust this plugin directory?; answer y. An interactive session opens in which Claude reads the plugin, asks what a good result looks like, proposes prompts that should and should not trigger it, designs graders, trial-runs them once, and writes one case folder per prompt under evals/. Exit with /exit or Ctrl+D when it says the suite is ready.

You can also ask Claude to run claude plugin eval init from a session already open at the plugin root; it will ask the same questions there.

2. Run it

claude plugin eval .

Each case runs three times with and three without, so six runs per case. A progress line prints per run with its score and each grader's verdict.

3. Read the summary

CASE                       WITH  W/OUT Δ      RUNS COST    NOTES
notes-from-pr-list         1.00  0.33  +0.67  6    $0.38
ignores-unrelated-request  1.00  1.00  +0.00  6    $0.12

2 case(s) · mean Δ +0.33 · 81s · $0.50
Report: /Users/cam/plugins/release-notes/evals/results/2026-10-08T09-14-03-210Z/report.html

COST is a list-price estimate. NOTES shows the highest-weight failing grader's explanation (or the run error) from the with-plugin arm. If your account can publish reports you also get a Published: line with a URL.

4. Iterate

Open the report. The most common first finding is Δ near zero with the tool_used: Skill grader failing: Claude is not choosing your skill. Rewrite the skill's description (see /docs/skills) and run again.

For quick iteration on one case, run one arm once, then confirm at the default three runs before you believe it:

claude plugin eval . --case notes-from-pr-list --runs 1 --ablation none

With a single arm the table shows SCORE and PASS% instead of WITH, W/OUT and Δ.

Writing cases by hand

Everything init writes is plain files. A case is a folder containing prompt.md, case.yaml or both, with at least one grader (a case without graders fails to load). To group cases, nest them under a folder that is not itself a case.

release-notes/
├── .claude-plugin/plugin.json
├── skills/release-notes/SKILL.md
└── evals/
    ├── notes-from-pr-list/
    │   ├── prompt.md
    │   ├── graders/
    │   │   ├── customer-tone.md
    │   │   └── skill-fired.md
    │   └── case.yaml          # only if you need context.* fields
    ├── ignores-unrelated-request/
    └── results/               # written by runs; add to .gitignore

Get a blank template without running anything:

claude plugin eval init --bare notes-from-pr-list

The prompt

Write it the way a user would, without naming the skill:

---
max_turns: 8
allowed_tools: [Read, Skill]
tags: [smoke]
---

Turn these into release notes our customers can read:
- #812 fix: retry webhook delivery on 502
- #815 feat: CSV export for invoices
- #817 chore: bump eslint

Each run starts in an empty working directory, so include what the task needs in the prompt or seed the workspace. @path mentions are not expanded; if Claude must read a file, allow a tool for it.

The graders

A rubric for the judge model, written as concrete PASS and FAIL conditions, in graders/customer-tone.md:

---
type: llm
---

PASS if the notes mention CSV export for invoices and more reliable webhooks in plain language, and omit the eslint bump.
FAIL if internal PR numbers or the word "chore" appear, or either user-facing change is missing.

And proof that your skill did the work, in graders/skill-fired.md:

---
type: tool_used
tool: Skill
input_match: '"skill"\s*:\s*"(?:[\w-]+:)?release-notes"'
---

This passes if Claude invoked the skill at least once, including via its namespaced plugin:skill form.

Choosing graders that give a stable signal

There are six grader types. regex, tool_used, tool_order and file_exists are computed from the transcript and files and cost nothing. llm and baseline call a judge model. There are no custom-code graders. The judge defaults to the model Claude Code uses for background tasks; override with --judge-model sonnet or a full model ID.

Habits that keep scores trustworthy:

  • Long output, use regex. A judge reading a long file gets flaky. Grade generated files with a regex on their contents; keep llm for short outputs with crisp PASS/FAIL rubrics.
  • One outcome grader, one process grader. Pair a check on the result with a tool_used or tool_order check on how it was produced. Together they tell you whether it was right and whether your plugin did it.
  • Negative Δ with the skill firing? Suspect the judge. A small judge may fail a correct answer for formatting. Re-run with --judge-model sonnet and tighten the rubric so formatting does not decide the verdict.
  • Verifying a build or test inside the run: have the prompt ask Claude to run it and write the outcome to a file, grade that file, and assert the command ran with tool_used plus an input_match naming it.

Scoring against the baseline

In a two-arm run, some graders would always fail without the plugin (a "skill was invoked" check, for instance), which would drag W/OUT towards zero and inflate Δ. So Claude Code excludes these from the score in both arms and reports them in the with-arm as indicators only, with scored: false:

  • every tool_used grader with tool: Skill;
  • every regex grader with target: mock_calls and every llm grader with focus: mock_calls, when each mocked server in the case is one your plugin declares;
  • any grader marked arm: with-only.

Overrides: if every grader in a case would be excluded, they are scored normally. arm: both forces scoring in both arms, which is what you want for a "must not invoke the skill" check (min: 0, max: 0). Under --ablation none nothing is excluded, so absolute scores differ between modes.

A case runs the with-arm only (no W/OUT, no Δ) when you pass --ablation none; when it resumes a transcript via context.history_file and the target is a path (a single-arm (no Δ) notice prints on stderr; pass --ablation with-without to force both); or when no plugin could be found for it.

A different eval folder

If evals/ is taken, add "experimental": { "evals": "quality/evals" } to plugin.json, or pass --eval-dir quality/evals to both claude plugin eval and init. The flag wins over the manifest. Paths must be relative, plain folder names: absolute paths or .. are an error as a flag, and as a manifest value print a Warning: and fall back to evals/.

Fixtures and mocks

Seeding the workspace or conversation

Add a case.yaml next to prompt.md with a context block:

schema_version: "1.1"
name: notes-from-git-log
tags: [nightly]
context:
  scaffold_script: make-repo.sh
  add_dirs: [samples]
  • scaffold_script: a Bash script in the case folder that builds fixture files or a git repo in the empty workspace. It runs as you, outside the sandbox, only when you pass --scaffold. It gets a minimal environment: your PATH, HOME set to the run's temporary home, TMPDIR, and a few constants like TERM=dumb; not your other shell variables and not the case's EVAL_* variables. Non-zero exit or more than 120 seconds scores the run 0 with scaffold failed. Project config it writes is not loaded, so use it for files and git state only.
  • history_file: a .jsonl transcript to resume; the prompt becomes the next user turn.
  • add_dirs: folders inside the case that Claude may read, read-only.
#!/bin/bash
# make-repo.sh
git init -q && git config user.email t@example.com && git config user.name test
echo "v1" > app.txt && git add . && git commit -qm "feat: initial invoices page"
echo "v2" > app.txt && git commit -qam "fix: rounding on VAT totals"

Mocking MCP servers

You can evaluate skills that call MCP tools without the real service. Put one Markdown file per tool at evals/mocks/<server>/<tool>.md (suite-wide) or in a case's own mocks/ folder, where <server> is the server name from your plugin's MCP config.

Runs never start your real servers unless you ask. Claude Code registers a stand-in under each server's name: mocked tools answer from their files and need no grant; tools with no mock are unavailable. A server with no mocks shows in the mocked: progress line as plugin_<plugin>_<server>[not started: no mock].

The file body is what the tool returns. For a post_message tool on a chat server:

---
expect:
  channel: /^#[a-z-]+$/
  text: string
---

Posted to {{input.channel}} (ts 1728380000.000100)

Options:

  • {{input.<field>}} inserts call input; {{file:fixtures/{input.<field>}.json}} inserts a fixture file beside the mock.
  • expect: guards the input. A violating call aborts the run with score 0 and records why, so you can assert what your plugin asked for.
  • error: true returns the body as a tool error.
  • type: agent makes the judge model act as the server, following instructions in the body. All agent-mock calls in a run share a budget of four times max_turns; exceeding it aborts with score 0.

Grade the calls themselves with target: mock_calls.

To hit real servers instead, use --allow-real-servers (start real servers only for unmocked ones) or --mocks off (ignore mocks, start everything). Either way those processes run as you outside the sandbox and their tools need an --allow-tools grant.

Replaying agent mocks. Agent-mock answers vary run to run. After a clean run, Claude Code saves each answer under the results folder in mock-recordings/. ADOPT.txt there lists each recording and the .replay/<server>/ folder to copy it into. Once copied, identical calls are answered from the recording with no model call. Commit mocks/.replay/ so CI is repeatable.

Running evals

Targets

TargetRuns
A plugin root, such as .Every case in its eval folder, with that plugin
One prompt.md or case.yamlThat case, with its enclosing plugin
An installed plugin, name or name@marketplaceCases in the installed copy, with that copy loaded; results go to ./evals/results/ in your current folder (or ./<dir>/results/ with --eval-dir)
name@skills-dirSame, for a skills-directory plugin
NothingThe current folder as a path

Filter with --case <glob> and --tag <tag>. Put the target before --tag, --allow-tools and --json, because those flags would swallow it as a value.

Granting tools

Runs never pause for permission. Tools that need a grant you did not give (Bash, Write, Edit, WebFetch, WebSearch) are removed entirely.

A case may use the read-only tools it lists in allowed_tools, from: Read, Glob, Grep, NotebookRead, Skill, AskUserQuestion, Agent, TodoWrite, TaskCreate, TaskGet, TaskList, TaskUpdate, TaskStop. Anything else you grant for the whole run:

claude plugin eval . --allow-tools Write Edit "Bash(pnpm test *)"

Ungranted requests show as not granted in progress output. Real plugin MCP tools need the server started and a grant such as --allow-tools "mcp__plugin_release-notes_chat__*".

Granting Bash in any form runs every command under the OS-level sandbox: writes confined to the workspace, home and Claude Code config unreadable, network limited to domains you grant with --allow-tools "WebFetch(domain:example.com)". With no sandbox backend, shell-granting runs are refused (usually scoring 0). Native Windows has none, so use WSL2; on Linux install bubblewrap and socat. See /docs/sandboxing.

Options

claude plugin eval --help has the full list, including --case, --tag, --eval-dir, --no-scaffold, --report and --verbose.

OptionDefaultWhat it does
--runs <n>Case's runs, else 3Runs per case per arm
-j, --concurrency <n>11 to 8 parallel runs; shares your rate limit; results keep case order
--model <model>Case model, else ANTHROPIC_MODEL, else defaultAgent under test. Pin it in CI
--judge-model <model>Background-task modelJudge for llm and baseline
--ablation <mode>Per casenone for one arm, with-without for the baseline too
--threshold <0..1>1.0Minimum with-arm score; any case below makes exit 1
--max-cost-usd <usd>NoneCeiling on estimated list-price cost, checked before each run starts. Started runs finish, so spend can overshoot. Unstarted runs mean exit 2
--allow-tools <tools...>NoneExtra tool grants
--scaffoldOffRun scaffold_scripts
--trust-pluginOffSkip the first-run trust prompt; use in CI
--mocks <mode>recordrecord uses mocks and saves agent-mock answers; off starts real servers
--allow-real-serversOffWith record, also start real servers that have no mock
--json [path]OffResult document to stdout or a .json path; suppresses progress and table
--output-dir <dir><eval dir>/results/<timestamp>/Where results go
--no-publishKeep the HTML report local
--publish-reportPublish even where it would stay local by default
--keep-tempOffKeep each run's sandbox folder and print its path

In CI

claude plugin eval . \
  --trust-plugin \
  --json eval-results.json \
  --threshold 0.75 \
  --model claude-sonnet-5 \
  --judge-model claude-haiku-4-5 \
  --no-publish \
  --max-cost-usd 15

Pin both models so a model rollout is not mistaken for a plugin regression.

ExitMeaning
0All cases at or above threshold and all case files loaded
1A case below threshold, a case failed to load, no cases found, a run could not start, untrusted directory without --trust-plugin, or a bad option
2Partial: cost ceiling hit, or credentials rejected at or before the first run. JSON still written with partial: true
130Interrupted; partial results written
143Terminated (for example a CI timeout)

Δ never affects the exit code, nor do report write or publish failures. The runner needs Claude Code installed and credentials such as ANTHROPIC_API_KEY (see /docs/authentication). eval init needs a terminal, so in CI use init --bare <name>.

To keep spend predictable on every-push suites, I use only free graders, --ablation none, and keep the judge-based, two-arm suite for a nightly job. Exclude partial: true results and runs with skippedPaidGraders from any trend chart.

Reading results

Each run with at least one case writes results/<timestamp>/ with aggregate-result.json and report.html (under the plugin for path targets, under your current folder for named targets).

The HTML report

A single self-contained file with no external requests, fine to attach to a CI job. Read it top down:

  • Verdict line and tiles: suite score (mean with-plugin case score), ablation Δ against the baseline, cases meeting the threshold, and perfect runs (share of with-plugin runs where every grader passed).
  • Case cards: each case's Δ and score with a threshold tick. Negative Δ cases get a red left edge.
  • Inside a case: with-plugin runs first, then baseline. Failed graders are pre-expanded with their explanation; llm graders show the judge's votes and the evidence it saw. Unscored graders carry a plugin-fired indicator badge.
  • Prompt and Graders: the case's prompt and every rubric or pattern, so readers do not need the suite.

Signed in with a claude.ai subscription and with artifacts available, the report is also published as a private artifact (Published: <url>). --no-publish keeps it local. Runs started by a Claude Code session stay local (kept local) unless you add --publish-report.

The JSON document

schemaVersion: 1, camelCase fields, new fields added without renames, so ignore unknown keys. Fields gating scripts usually read:

FieldMeaning
partial, partialReasontrue with cost_ceiling, interrupted or auth_failed if unfinished
aggregates.overallScoreMean case score
aggregates.casesPassed, aggregates.casesTotalCases at threshold, and total
aggregates.meanDeltaMean Δ in two-arm mode
cases[].nameCase name
cases[].aggregates.scoreMean with-arm score
cases[].aggregates.deltaWith minus without; omitted if single-arm or not comparable
cases[].arms.with[].errornull or why the run ended abnormally (still graded on what it produced)
cases[].arms.with[].abortedSet when a mock stopped the run via expect:, abort_when or the budget; includes server, tool, reason; scores 0 with error null
cases[].arms.with[].skippedPaidGradersJudge graders skipped by the cost ceiling
costUsd, durationSeconds, claudeVersionEstimated cost, wall time, version

A tiny gate on top of the exit code:

jq -e '.partial == false and .aggregates.meanDelta > 0.2' eval-results.json

What a run can access

claude plugin eval loads the plugin's skills, hooks and agents and runs on your machine as you. Pointing it at a plugin is the same trust decision as claude --plugin-dir. The isolation below limits the agent under test, not the plugin's own code, and a passing suite says nothing about whether a plugin is safe.

Trust

First run against a directory asks Trust this plugin directory? unless you already accepted trust there interactively. Inside a git repo, yes trusts the whole repo, interactive sessions included. When stdin or stdout is not a terminal, or with --json, it cannot ask and exits 1; --trust-plugin asserts trust. Named targets (installed or skills-directory plugins) skip the prompt.

Only the matching flag enables: scaffold_script (--scaffold), non-read-only tools (--allow-tools), real MCP servers (--allow-real-servers or --mocks off). Neither a case's allowed_tools nor a skill's allowed-tools can widen these. If the plugin has hooks you did not write or you start real servers, treat scores as advisory unless run in a container or CI runner, since those run outside the sandbox and could tamper with what graders read.

Isolation

Each run gets a temporary home, working directory and Claude Code config, and runs as a claude -p child with only your plugin.

  • Nothing personal or project-level loads: no user settings, hooks, CLAUDE.md, MCP servers, other plugins, memory or skills; no .claude/, CLAUDE.md or .mcp.json from the workspace or above it, even one a scaffold wrote. Most of your environment is withheld. Ship everything a case depends on inside the plugin.
  • Managed policy still applies, so managed machines can score differently.
  • The Artifact tool is off. Grade what a skill produces before it would publish.
  • Case definitions are hidden: Claude cannot read the eval folder.
  • No network sandbox outside shell commands: WebFetch(domain:...) grants reach the domain directly, and hooks and real servers can reach anything.

Suite reference

evals/
├── <case>/
│   ├── prompt.md              # frontmatter: case/run fields; body: prompt
│   ├── case.yaml              # optional: context.*, or the whole case
│   ├── graders/<name>.md      # one grader per file
│   ├── mocks/                 # optional per-case mocks
│   └── <fixtures, scripts, transcripts>
├── mocks/<server>/
│   ├── <tool>.md              # one mocked tool
│   ├── _server.md             # optional agent answering several tools
│   ├── _tools.json            # optional saved tools/list response
│   └── fixtures/
├── mocks/.replay/<server>/    # adopted agent-mock recordings
└── results/<timestamp>/
    ├── aggregate-result.json
    ├── report.html
    └── mock-recordings/       # with ADOPT.txt

prompt.md frontmatter

Unknown keys are an error.

FieldDefaultPurpose
schema_version"1.1" (automatic)Case format version
nameFolder nameMatched by --case; report key
descriptionHuman note
tags[]For --tag; any match runs the case
pluginsNearest enclosing pluginPlugin folders under test, relative to the case. Use ["../.."] if detection fails
runs31 to 50 per arm
expected_outcomeHuman note
modelChild session defaultAgent model
max_turns10Up to 200; hitting it is a run error
timeout_seconds300Up to 3600
allowed_tools[]Tools the case wants
append_system_promptAppended to the child's system prompt
env{}Extra variables; keys must match EVAL_[A-Z0-9_]*. Runs inherit only an allowlist (PATH, locale, proxy and certificate settings, provider and auth variables, most ANTHROPIC_* and CLAUDE_CODE_*, and EVAL_*)

case.yaml

Requires schema_version: "1.1" and name. description, tags, plugins, runs and expected_outcome sit at top level; model, max_turns, timeout_seconds, allowed_tools, append_system_prompt and env go under execution:. When both files exist, prompt.md frontmatter overrides matching fields, its body is the prompt, and graders/*.md are appended after graders listed in case.yaml.

Field (case.yaml only)Purpose
context.scaffold_scriptWorkspace setup script (needs --scaffold, 120 s limit)
context.history_file.jsonl transcript to resume
context.add_dirsRead-only folders inside the case
execution.promptThe prompt, if there is no prompt.md
gradersList of graders, each with name plus the usual keys; llm rubric goes in criteria

Grader keys and types

Every grader takes type (required), weight (default 1, any positive number) and arm (with-only or both). The grader's name is its file name without .md.

regex uses target and llm uses focus, with the same values:

ValueSees
last_messageFinal reply (default)
traceSession as JSON lines. Regex sees all; a judge sees the first 12 and last 12. Quotes appear escaped as \"
filesPaths Claude created, not contents, not edited or scaffolded files
{ source: file, path: <path> }One file's contents after the run. PNG, JPEG, GIF and WebP go to the judge as images; other binaries are refused
mock_callsEach mocked tool call with input and answer
TypeOptionsPasses when
regexpattern, flags, match, targetJavaScript regex found. match: not_contains for absence, match: "count:N" for exactly N. Use flags: i, not (?i)
tool_usedtool, input_match, min, maxMatching calls between min (default 1) and max (unlimited). min: 0 and max: 0 asserts never
tool_orderbefore, afterBoth called and first before precedes first after. Each a name or { tool, input_match }
file_existspath, existsA created file matches the glob, or none with exists: false
llmcriteria, focusJudge votes PASS in at least two of three
baselinebaseline_file, criteriaJudge finds the run at least as good as the reference .jsonl

Mock files

KeyDefaultPurpose
typefixedfixed returns the body; agent uses it as instructions for the judge acting as the server, with earlier calls as history
expectMap of dotted input paths to a type (string, number, boolean, array, object), a /regex/, a literal, or a list of literals. Violations abort with score 0
errorfalsefixed only: return as a tool error
abort_whenagent only: prose listing the only conditions to abort

_server.md is one type: agent mock answering the tools in its tools: key; a <tool>.md for the same tool wins. An expect: there is a load error unless tools: lists one tool. _tools.json is a saved tools/list response so mocks carry real descriptions and schemas. Per-case mocks/ overrides suite mocks file by file.

/regex/ in expect: is a restricted dialect: literals, ., escapes like \d, classes like [a-z]; *, +, ? and {m,n} on a single item; optional leading ^ and trailing $; flags i and s only. Groups, alternation, backreferences and lookaround stop the case loading (score 0). Use a list of literals instead of alternation. Values are checked only up to a maximum length; anchor with ^ and keep quantifiers few.

Troubleshooting

SymptomFix
plugin eval is currently in early accessYour build is older than general availability. claude update, fresh session
plugin eval is currently unavailableSwitched off server-side. Update and try later
is not a trusted plugin directory, and this run cannot stop to ask you about itRun once in a terminal and accept, or pass --trust-plugin
is too old for claude plugin evalGit below 2.31 cannot honour the config that disables repo hooks and helpers. Upgrade git. (Before v2.1.283 the version was not checked)
No eval cases foundRun from the plugin root, check filters, or run init
No W/OUT column, or ablation requested but no plugin resolvedAdd plugins: ["../.."] to the case. Expected for history_file cases
Δ near zero, skill grader failingReal finding: improve the skill's description
Agent type '<plugin>:<agent>' not found in baseline runsExpected: the baseline has no plugin. Use --ablation none to skip it
Everything scores 0 though files existYou targeted files (paths) instead of { source: file, path: ... }. file_exists ignores edited or scaffolded files
Regex over trace misses visible textDefault target is last_message; trace is JSON-escaped; use flags: i
Tools denied or MCP tools missingGrant with --allow-tools; start real servers explicitly or mock them
Exit 1 but results look fineDefault threshold is 1.0; set your own. Also check stderr for load errors
--json output path must end in .jsonTarget went after --json; put it first
Grader passed: false in a 1.0 runIt is unscored by design (scored: false)
Usage or rate limit errors mid-suiteLater runs score ~0 but the suite is not partial. Check NOTES or error, re-run after reset
mock call budget exceededAgent mocks share 4 × max_turns calls (replays count). Raise max_turns
Timeouts or turn capRaise max_turns and timeout_seconds; control spend with --max-cost-usd