Test a mod
Write automated tests for Claude Code mods with claude plugin test - fire events, stub what Claude Code would answer, drive timers and press buttons, all offline.
claude plugin test runs automated tests for a mod from your shell, with no session, sign-in or network. A test loads your module, fires events through its hooks as Claude Code would, and checks what the hooks did. It is the fastest way to catch a broken mod before it reaches anyone's session, and it exits non-zero on failure so it slots straight into CI.
A first test
Test files end in .test.ts (or .test.tsx), live anywhere in the plugin folder, import the test kit from claude-code/testing, and must declare at least one test() (otherwise: declares no test(): nothing ran). They can also import your own files and sibling .ts helpers, so plain functions can be unit tested without the kit.
Using the edit-meter mod from /docs/plugins/mods/create, save edit-meter/tests/edit-meter.test.ts:
import { expect, test } from 'claude-code/testing'
test('refused edits are not counted', async ($, on) => {
// Play Claude Code: deny edits to .env, let everything else succeed
on('tool.call', ($, e) => (String(e.file_path).endsWith('.env') ? { deny: 'no' } : { result: 'ok' }))
await $.tool.call({ tool: 'Edit', file_path: 'src/a.ts', old_string: 'a', new_string: 'b' })
await $.tool.call({ tool: 'Edit', file_path: '.env', old_string: 'x', new_string: 'y' })
const reply = await $.command.run({ command: 'edits', args: '' })
expect(reply.text).toBe('1 file(s):\nsrc/a.ts')
})
Run it from the mod folder:
claude plugin test
tests/edit-meter.test.ts:
(pass) refused edits are not counted [19.40ms]
1 pass
0 fail
Ran 1 test across 1 file. [0.17s]
Each $.tool.call passed through the mod's tool.call hook and on to the stub; no real tool ran. $.command.run reached the command.run hook and returned its result.
If mods cannot load in the shell running the tests, the command prints a line starting claude plugin test: hooks modules are turned off with the reason, and exits 1.
Two kinds of $, plus stubs
The test function receives $ and on, and neither is what a hook receives:
- The test's
$plays Claude Code. Each method fires the matching event through your hooks and resolves to the result:$.tool.call(...),$.command.run(...),$.prompt.submit(...),$.session.start(...),$.turn.complete(...), and$.classic.<Event>(...)for settings hook events. You cannot fire a mods API call such asui.closedirectly; trigger it through your mod (for example, press the button that closes the pane). onregisters stubs: hooks that answer in Claude Code's place, because no model, store or tool runs in a test. Name a stub for an API call without$., soon('store.get', ...)answers your mod's$.store.get.
What a stub returns
- A stub for a mods API call returns
{ value }, the thing the call should resolve to ({ value: 7 }makes$.store.getresolve to7), or{ deny: reason }to make the call reject. A stub that throws is skipped, not treated as a rejection. - A stub for a Claude Code event such as
tool.callorturn.stepreturns that event's own result shape, like{ result: 'ok' }.
Errors that mean a stub is wrong or missing appear in a failed test's output under the engine reported::
returned neither { value } nor { deny }: an API stub returned a bare value.no implementation for <name>: your mod made a call nothing answers.
Example: stubbing a model call
A mod named naming with a /name-branch command that asks a model for a branch name:
export function register(on) {
on('command.run', { command: 'name-branch' }, async ($, e) => {
const r = await $.model.complete({
model: 'haiku',
system: 'Suggest a git branch name in kebab-case, prefixed feat/ or fix/. Reply with the name only.',
prompt: e.args,
})
const name = r.isAnswered ? r.text.trim() : ''
return { text: /^(feat|fix)\/[a-z0-9-]+$/.test(name) ? name : 'Could not suggest a name' }
})
}
import { expect, test } from 'claude-code/testing'
const usage = { input_tokens: 12, output_tokens: 6, cache_read_input_tokens: 0, cache_creation_input_tokens: 0 }
test('accepts a well-formed suggestion', async ($, on) => {
on('model.complete', () => ({ value: { isAnswered: true, text: 'fix/vat-rounding\n', usage } }))
const reply = await $.command.run({ command: 'name-branch', args: 'VAT totals round wrongly' })
expect(reply.text).toBe('fix/vat-rounding')
})
test('rejects a malformed suggestion', async ($, on) => {
on('model.complete', () => ({ value: { isAnswered: true, text: 'Sure! How about vat-fix?', usage } }))
const reply = await $.command.run({ command: 'name-branch', args: 'VAT totals round wrongly' })
expect(reply.text).toBe('Could not suggest a name')
})
As with any mod, naming also needs a manifest, a hooks.json, and (to use /name-branch in a real session) a $.command.register call.
Built-in mocks
The kit exports mock helpers that answer a whole namespace:
mock.clock(on): a controllable clock for$.clock(see timers).mock.store(on, { pins: 3 }): an in-memory$.storeseeded with entries. It returns nothing, so to inspect what was saved, write the two store stubs yourself as in Testing a drawing.mock.env(on, { CI: 'true' }): answers$.env.get.
Kit rules that trip people up
-
Register all stubs before the first call on
$. Late registration throws, for exampleon("ui.render") after the test first called $. -
session.startdoes not run on its own. Each test starts with a freshly loaded module and no hooks called. If your code depends on start-up work, fire it, and stub what that work calls:on('session.start', () => ({ cwd: '/repo' })) on('command.register', () => ({ value: undefined })) on('store.get', () => ({ value: undefined })) await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/repo' })Without the
command.registerstub, the call rejects withno implementation for command.register, your hook is skipped from that point, and the test does not fail there. It only shows underthe engine reported:if a later assertion fails. -
A
ui.renderhook that returnsnext(e)needs a stub or mounting fails withno implementation for ui.render. Return an element as plain data:on('ui.render', () => ({ type: 'Text', props: {}, children: ['engine drawing'] })).ui.find({ type: 'Text' })then finds it. -
turn.stepstubs are async generators, and the test must drain the stream:on('turn.step', async function* ($, e) { yield { kind: 'text', index: 0, text: 'done' } return { turnId: e.turnId, index: e.index, answer: 'done', toolUses: [], stopReason: 'end_turn', usage: null } }) const stream = $.turn.step({ turnId: 't1', index: 0, model: 'claude-test', messageCount: 1 }) let step = await stream.next() while (step.done !== true) step = await stream.next() expect(step.value.answer).toBe('done') -
Tool calls take the tool name and its arguments as fields:
$.tool.call({ tool: 'Bash', command: 'ls' }), with atool.callstub returning{ result }.
Stub cheat sheet
The kit answers $.ui.invalidate and $.state itself. Use mock.clock(on) for $.clock, or $.clock.now() fails with no implementation for clock.now. For everything else:
| Your mod calls or passes on | Stub |
|---|---|
$.command.register, $.tool.register, $.ui.toast, $.ui.log, $.ui.status, $.ui.close, $.store.set | () => ({ value: undefined }) (toast and log text is e.text) |
$.store.get | ($, e) => ({ value: saved.get(e.key) }) |
$.fs.read | ($, e) => ({ value: e.path.endsWith('config.json') ? '{}' : '' }) (e.path is absolute) |
$.ui.open | () => ({ value: { isPlaced: true } }) |
$.ui.ask | A tool.call stub, because it arrives as an AskUserQuestion call: ($, e) => ({ result: { answers: { [e.questions[0].question]: 'Deploy' } } }). Check e.tool if other calls pass through |
$.model.complete | () => ({ value: { isAnswered: true, text: '...', usage } }) |
$.process.run | ($, e) => ({ value: { exitCode: 0, stdout: '...', stderr: '' } }) (e.argv, and e.init with cwd and timeoutMs) |
| Any call that should fail | () => ({ deny: 'reason' }) |
session.start | () => ({ cwd: '/repo' }) |
turn.start | ($, e) => ({ turnId: e.turnId }) |
tool.call | () => ({ result: '...' }) |
turn.complete | () => ({ text: '' }); fire with $.turn.complete({ turnId, answer, durationMs, isAborted: false, usage: null }) |
prompt.submit | ($, e) => ({ text: e.text }) |
prompt.fill | () => ({ isFilled: true }) |
$.prompt.read | () => ({ value: { text: '...', cursor: 0 } }) |
$.ui.copy | () => ({ value: { isCopied: true } }) |
$.session.messages | () => ({ value: [{ role: 'assistant', text: '...', toolUses: [] }] }) |
$.session.id, $.agent.list | () => ({ value: 'abc123' }), () => ({ value: [] }) |
session.send | () => ({ isDelivered: true }) (e.to arrives as a string even if you passed { sessionId }) |
session.receive | ($, e) => ({ text: e.text }); fire with $.session.receive({ origin: { kind: 'peer-send-message' }, text }) |
ui.render | () => ({ type: 'Text', props: {}, children: ['...'] }) |
expect supports toBe, toEqual, toMatch, toMatchObject, toContain, toBeDefined, toBeUndefined and toThrow, each negatable with .not. An expect that fails inside a plain-function stub or hook fails the test, and the output names it, for example in the test's store.set hook.
Testing timers
const clock = mock.clock(on) starts at 0 (or mock.clock(on, { now: 5000 })) and only moves when you move it:
| Method | Effect |
|---|---|
await clock.advance(ms) | Move forward and run every timer that comes due |
await clock.set(ms) | Move to an absolute time, like advance |
clock.now() | Current time (what $.clock.now() resolves to) |
await clock.settle() | Run timers already due, such as chained zero-delay after calls, without moving time |
await clock.sleep(ms) | Inside a stub: answer only once the test has advanced that far, to simulate slow work |
A reminder mod whose /remind <minutes> <text> shows a toast later:
export function register(on) {
on('command.run', { command: 'remind' }, async ($, e) => {
const [mins, ...words] = e.args.split(' ')
$.clock.after(Number(mins) * 60_000, () => $.ui.toast(words.join(' ')))
return { text: 'Reminder set for ' + mins + ' min' }
})
}
import { expect, mock, test } from 'claude-code/testing'
test('fires after the delay, not before', async ($, on) => {
const clock = mock.clock(on)
const toasts: string[] = []
on('ui.toast', ($, e) => { toasts.push(e.text); return { value: undefined } })
await $.command.run({ command: 'remind', args: '25 stand up and stretch' })
await clock.advance(24 * 60_000)
expect(toasts).toEqual([])
await clock.advance(60_000)
expect(toasts).toEqual(['stand up and stretch'])
})
Each advance resolves after due timers have run, so the next line sees their effect. Twenty-five minutes, tested in milliseconds.
Testing a drawing
$.ui.mount draws a render site through your ui.render hook and returns a handle for interacting with it. Set surface to test each app. For the ctx-gauge pane from /docs/plugins/mods/interface:
import { expect, test } from 'claude-code/testing'
const PANE = {
plugin: 'ctx-gauge',
component: 'Pane',
requestId: 'ctx-gauge',
viewport: { columns: 120, rows: 40 },
props: {
title: 'Context',
isFocused: true,
bodyColumns: 70,
placement: 'inline',
scroll: { offset: 0, bodyRows: 12 },
view: {},
},
} as const
test('pins tab saves pins in both apps', async ($, on) => {
const saved = new Map<string, unknown>()
on('store.get', ($, e) => ({ value: saved.get(e.key) }))
on('store.set', ($, e) => { saved.set(e.key, e.value); return { value: undefined } })
on('turn.complete', () => ({ text: '' }))
on('session.usage', () => ({ value: { startedAt: 0, context: { tokens: 50000, window: 200000, percent: 25 }, rateLimits: [], cost: 0 } }))
// Produce a percentage by finishing one turn
await $.turn.complete({ turnId: 't1', answer: 'ok', durationMs: 10, isAborted: false, usage: null })
for (const surface of ['terminal', 'desktop'] as const) {
const ui = await $.ui.mount({ ...PANE, surface })
await ui.press({ key: 'tab-pins' })
await ui.press({ key: 'pin' })
expect(await ui.find({ type: 'Text', text: /^Pinned: / })).toBeDefined()
await ui.unmount()
}
expect(saved.get('pins')).toEqual([25, 25])
})
Both mounts share one loaded module, so state carries from the first app to the second.
| Handle method | Does |
|---|---|
press({ key }) | Press the Button with that key |
input({ key, text }) | Type into the Input and press Enter; add kind: 'change' to type without submitting |
select({ key, value }) | Pick that option in the Select |
find({ key }) or find({ type, text }) | First matching element as { type, props, children }, or undefined. text is a string or regex |
unmount() | Remove the drawing |
Each method resolves after your handler finishes. Set props to what Claude Code would pass for the site (see the render sites table). A drawing test checks the tree and its validity for the app, not how it is painted, so still eyeball new layouts in a real session.
After /clear
Tests start with all $.state at defaults, exactly as after /clear. To test your recovery path, skip session.start, fire classic.SessionStart with source: 'clear', then check the drawing:
test('pin count is restored after /clear', async ($, on) => {
on('store.get', () => ({ value: 4 }))
on('classic.SessionStart', () => ({}))
await $.classic.SessionStart({ source: 'clear' })
const ui = await $.ui.mount({ ...PANE, surface: 'terminal' })
await ui.press({ key: 'tab-pins' })
expect(await ui.find({ type: 'Text', text: 'Pins: 4' })).toBeDefined()
})
This assumes a $.state version of the pane that draws Pins: N from the pinCount atom and has the classic.SessionStart hook shown in /docs/plugins/mods/interface. Without that hook the pane draws the default Pins: 0 and the test fails.
Testing a policy mod
A mod in your organisation's prependPlugins can refuse other mods as they load. Test it by setting its tier and supplying inline mods for it to judge:
tier('prepend')at the top of the file loads your mod asprepend(orappend,builtin). Default isuser.- Pass
testan options object withplugins: inline mods, each withname,register, and optionallytier.
Suppose your policy mod acme-guard refuses any mod that calls http.fetch:
import { expect, test, tier } from 'claude-code/testing'
tier('prepend')
const phoneHome = {
name: 'phone-home',
register(on) {
on('tool.call', async ($, e, next) => {
await $.http.fetch('https://example.com/beacon')
return next(e)
})
},
}
const quiet = {
name: 'quiet',
register(on) {
on('tool.call', async () => ({ result: 'quiet answered' }))
},
}
test('refuses a mod that makes network calls', { plugins: [phoneHome] }, async ($, on) => {
on('tool.call', () => ({ result: 'engine' }))
let message = ''
try {
await $.tool.call({ tool: 'Read', file_path: 'a' }) // first call on $ loads the mods
} catch (error) {
message = error.message
}
expect(message).toMatch(/^phone-home: refused by acme-guard/)
})
test('admits a mod with no network calls', { plugins: [quiet] }, async ($, on) => {
on('tool.call', () => ({ result: 'engine' }))
expect(await $.tool.call({ tool: 'Read', file_path: 'a' })).toEqual({ result: 'quiet answered' })
})
All mods load at the test's first call on $. When your mod refuses one, that call throws with a message naming the refused mod, your mod, and your reason (for example phone-home: refused by acme-guard: ...). Writing policy mods themselves is covered in /docs/plugins/mods/admin.