Skip to content

Test a mod

Write automated tests for Claude Code mods with claude plugin test - fire events, stub what Claude Code would answer, drive timers and press buttons, all offline.

claude plugin test runs automated tests for a mod from your shell, with no session, sign-in or network. A test loads your module, fires events through its hooks as Claude Code would, and checks what the hooks did. It is the fastest way to catch a broken mod before it reaches anyone's session, and it exits non-zero on failure so it slots straight into CI.

A first test

Test files end in .test.ts (or .test.tsx), live anywhere in the plugin folder, import the test kit from claude-code/testing, and must declare at least one test() (otherwise: declares no test(): nothing ran). They can also import your own files and sibling .ts helpers, so plain functions can be unit tested without the kit.

Using the edit-meter mod from /docs/plugins/mods/create, save edit-meter/tests/edit-meter.test.ts:

import { expect, test } from 'claude-code/testing'

test('refused edits are not counted', async ($, on) => {
  // Play Claude Code: deny edits to .env, let everything else succeed
  on('tool.call', ($, e) => (String(e.file_path).endsWith('.env') ? { deny: 'no' } : { result: 'ok' }))

  await $.tool.call({ tool: 'Edit', file_path: 'src/a.ts', old_string: 'a', new_string: 'b' })
  await $.tool.call({ tool: 'Edit', file_path: '.env', old_string: 'x', new_string: 'y' })

  const reply = await $.command.run({ command: 'edits', args: '' })
  expect(reply.text).toBe('1 file(s):\nsrc/a.ts')
})

Run it from the mod folder:

claude plugin test
tests/edit-meter.test.ts:
(pass) refused edits are not counted [19.40ms]

 1 pass
 0 fail
Ran 1 test across 1 file. [0.17s]

Each $.tool.call passed through the mod's tool.call hook and on to the stub; no real tool ran. $.command.run reached the command.run hook and returned its result.

If mods cannot load in the shell running the tests, the command prints a line starting claude plugin test: hooks modules are turned off with the reason, and exits 1.

Two kinds of $, plus stubs

The test function receives $ and on, and neither is what a hook receives:

  • The test's $ plays Claude Code. Each method fires the matching event through your hooks and resolves to the result: $.tool.call(...), $.command.run(...), $.prompt.submit(...), $.session.start(...), $.turn.complete(...), and $.classic.<Event>(...) for settings hook events. You cannot fire a mods API call such as ui.close directly; trigger it through your mod (for example, press the button that closes the pane).
  • on registers stubs: hooks that answer in Claude Code's place, because no model, store or tool runs in a test. Name a stub for an API call without $., so on('store.get', ...) answers your mod's $.store.get.

What a stub returns

  • A stub for a mods API call returns { value }, the thing the call should resolve to ({ value: 7 } makes $.store.get resolve to 7), or { deny: reason } to make the call reject. A stub that throws is skipped, not treated as a rejection.
  • A stub for a Claude Code event such as tool.call or turn.step returns that event's own result shape, like { result: 'ok' }.

Errors that mean a stub is wrong or missing appear in a failed test's output under the engine reported::

  • returned neither { value } nor { deny }: an API stub returned a bare value.
  • no implementation for <name>: your mod made a call nothing answers.

Example: stubbing a model call

A mod named naming with a /name-branch command that asks a model for a branch name:

export function register(on) {
  on('command.run', { command: 'name-branch' }, async ($, e) => {
    const r = await $.model.complete({
      model: 'haiku',
      system: 'Suggest a git branch name in kebab-case, prefixed feat/ or fix/. Reply with the name only.',
      prompt: e.args,
    })
    const name = r.isAnswered ? r.text.trim() : ''
    return { text: /^(feat|fix)\/[a-z0-9-]+$/.test(name) ? name : 'Could not suggest a name' }
  })
}
import { expect, test } from 'claude-code/testing'

const usage = { input_tokens: 12, output_tokens: 6, cache_read_input_tokens: 0, cache_creation_input_tokens: 0 }

test('accepts a well-formed suggestion', async ($, on) => {
  on('model.complete', () => ({ value: { isAnswered: true, text: 'fix/vat-rounding\n', usage } }))
  const reply = await $.command.run({ command: 'name-branch', args: 'VAT totals round wrongly' })
  expect(reply.text).toBe('fix/vat-rounding')
})

test('rejects a malformed suggestion', async ($, on) => {
  on('model.complete', () => ({ value: { isAnswered: true, text: 'Sure! How about vat-fix?', usage } }))
  const reply = await $.command.run({ command: 'name-branch', args: 'VAT totals round wrongly' })
  expect(reply.text).toBe('Could not suggest a name')
})

As with any mod, naming also needs a manifest, a hooks.json, and (to use /name-branch in a real session) a $.command.register call.

Built-in mocks

The kit exports mock helpers that answer a whole namespace:

  • mock.clock(on): a controllable clock for $.clock (see timers).
  • mock.store(on, { pins: 3 }): an in-memory $.store seeded with entries. It returns nothing, so to inspect what was saved, write the two store stubs yourself as in Testing a drawing.
  • mock.env(on, { CI: 'true' }): answers $.env.get.

Kit rules that trip people up

  • Register all stubs before the first call on $. Late registration throws, for example on("ui.render") after the test first called $.

  • session.start does not run on its own. Each test starts with a freshly loaded module and no hooks called. If your code depends on start-up work, fire it, and stub what that work calls:

    on('session.start', () => ({ cwd: '/repo' }))
    on('command.register', () => ({ value: undefined }))
    on('store.get', () => ({ value: undefined }))
    await $.session.start({ surface: 'terminal', isInteractive: true, cwd: '/repo' })
    

    Without the command.register stub, the call rejects with no implementation for command.register, your hook is skipped from that point, and the test does not fail there. It only shows under the engine reported: if a later assertion fails.

  • A ui.render hook that returns next(e) needs a stub or mounting fails with no implementation for ui.render. Return an element as plain data: on('ui.render', () => ({ type: 'Text', props: {}, children: ['engine drawing'] })). ui.find({ type: 'Text' }) then finds it.

  • turn.step stubs are async generators, and the test must drain the stream:

    on('turn.step', async function* ($, e) {
      yield { kind: 'text', index: 0, text: 'done' }
      return { turnId: e.turnId, index: e.index, answer: 'done', toolUses: [], stopReason: 'end_turn', usage: null }
    })
    
    const stream = $.turn.step({ turnId: 't1', index: 0, model: 'claude-test', messageCount: 1 })
    let step = await stream.next()
    while (step.done !== true) step = await stream.next()
    expect(step.value.answer).toBe('done')
    
  • Tool calls take the tool name and its arguments as fields: $.tool.call({ tool: 'Bash', command: 'ls' }), with a tool.call stub returning { result }.

Stub cheat sheet

The kit answers $.ui.invalidate and $.state itself. Use mock.clock(on) for $.clock, or $.clock.now() fails with no implementation for clock.now. For everything else:

Your mod calls or passes onStub
$.command.register, $.tool.register, $.ui.toast, $.ui.log, $.ui.status, $.ui.close, $.store.set() => ({ value: undefined }) (toast and log text is e.text)
$.store.get($, e) => ({ value: saved.get(e.key) })
$.fs.read($, e) => ({ value: e.path.endsWith('config.json') ? '{}' : '' }) (e.path is absolute)
$.ui.open() => ({ value: { isPlaced: true } })
$.ui.askA tool.call stub, because it arrives as an AskUserQuestion call: ($, e) => ({ result: { answers: { [e.questions[0].question]: 'Deploy' } } }). Check e.tool if other calls pass through
$.model.complete() => ({ value: { isAnswered: true, text: '...', usage } })
$.process.run($, e) => ({ value: { exitCode: 0, stdout: '...', stderr: '' } }) (e.argv, and e.init with cwd and timeoutMs)
Any call that should fail() => ({ deny: 'reason' })
session.start() => ({ cwd: '/repo' })
turn.start($, e) => ({ turnId: e.turnId })
tool.call() => ({ result: '...' })
turn.complete() => ({ text: '' }); fire with $.turn.complete({ turnId, answer, durationMs, isAborted: false, usage: null })
prompt.submit($, e) => ({ text: e.text })
prompt.fill() => ({ isFilled: true })
$.prompt.read() => ({ value: { text: '...', cursor: 0 } })
$.ui.copy() => ({ value: { isCopied: true } })
$.session.messages() => ({ value: [{ role: 'assistant', text: '...', toolUses: [] }] })
$.session.id, $.agent.list() => ({ value: 'abc123' }), () => ({ value: [] })
session.send() => ({ isDelivered: true }) (e.to arrives as a string even if you passed { sessionId })
session.receive($, e) => ({ text: e.text }); fire with $.session.receive({ origin: { kind: 'peer-send-message' }, text })
ui.render() => ({ type: 'Text', props: {}, children: ['...'] })

expect supports toBe, toEqual, toMatch, toMatchObject, toContain, toBeDefined, toBeUndefined and toThrow, each negatable with .not. An expect that fails inside a plain-function stub or hook fails the test, and the output names it, for example in the test's store.set hook.

Testing timers

const clock = mock.clock(on) starts at 0 (or mock.clock(on, { now: 5000 })) and only moves when you move it:

MethodEffect
await clock.advance(ms)Move forward and run every timer that comes due
await clock.set(ms)Move to an absolute time, like advance
clock.now()Current time (what $.clock.now() resolves to)
await clock.settle()Run timers already due, such as chained zero-delay after calls, without moving time
await clock.sleep(ms)Inside a stub: answer only once the test has advanced that far, to simulate slow work

A reminder mod whose /remind <minutes> <text> shows a toast later:

export function register(on) {
  on('command.run', { command: 'remind' }, async ($, e) => {
    const [mins, ...words] = e.args.split(' ')
    $.clock.after(Number(mins) * 60_000, () => $.ui.toast(words.join(' ')))
    return { text: 'Reminder set for ' + mins + ' min' }
  })
}
import { expect, mock, test } from 'claude-code/testing'

test('fires after the delay, not before', async ($, on) => {
  const clock = mock.clock(on)
  const toasts: string[] = []
  on('ui.toast', ($, e) => { toasts.push(e.text); return { value: undefined } })

  await $.command.run({ command: 'remind', args: '25 stand up and stretch' })
  await clock.advance(24 * 60_000)
  expect(toasts).toEqual([])
  await clock.advance(60_000)
  expect(toasts).toEqual(['stand up and stretch'])
})

Each advance resolves after due timers have run, so the next line sees their effect. Twenty-five minutes, tested in milliseconds.

Testing a drawing

$.ui.mount draws a render site through your ui.render hook and returns a handle for interacting with it. Set surface to test each app. For the ctx-gauge pane from /docs/plugins/mods/interface:

import { expect, test } from 'claude-code/testing'

const PANE = {
  plugin: 'ctx-gauge',
  component: 'Pane',
  requestId: 'ctx-gauge',
  viewport: { columns: 120, rows: 40 },
  props: {
    title: 'Context',
    isFocused: true,
    bodyColumns: 70,
    placement: 'inline',
    scroll: { offset: 0, bodyRows: 12 },
    view: {},
  },
} as const

test('pins tab saves pins in both apps', async ($, on) => {
  const saved = new Map<string, unknown>()
  on('store.get', ($, e) => ({ value: saved.get(e.key) }))
  on('store.set', ($, e) => { saved.set(e.key, e.value); return { value: undefined } })
  on('turn.complete', () => ({ text: '' }))
  on('session.usage', () => ({ value: { startedAt: 0, context: { tokens: 50000, window: 200000, percent: 25 }, rateLimits: [], cost: 0 } }))

  // Produce a percentage by finishing one turn
  await $.turn.complete({ turnId: 't1', answer: 'ok', durationMs: 10, isAborted: false, usage: null })

  for (const surface of ['terminal', 'desktop'] as const) {
    const ui = await $.ui.mount({ ...PANE, surface })
    await ui.press({ key: 'tab-pins' })
    await ui.press({ key: 'pin' })
    expect(await ui.find({ type: 'Text', text: /^Pinned: / })).toBeDefined()
    await ui.unmount()
  }

  expect(saved.get('pins')).toEqual([25, 25])
})

Both mounts share one loaded module, so state carries from the first app to the second.

Handle methodDoes
press({ key })Press the Button with that key
input({ key, text })Type into the Input and press Enter; add kind: 'change' to type without submitting
select({ key, value })Pick that option in the Select
find({ key }) or find({ type, text })First matching element as { type, props, children }, or undefined. text is a string or regex
unmount()Remove the drawing

Each method resolves after your handler finishes. Set props to what Claude Code would pass for the site (see the render sites table). A drawing test checks the tree and its validity for the app, not how it is painted, so still eyeball new layouts in a real session.

After /clear

Tests start with all $.state at defaults, exactly as after /clear. To test your recovery path, skip session.start, fire classic.SessionStart with source: 'clear', then check the drawing:

test('pin count is restored after /clear', async ($, on) => {
  on('store.get', () => ({ value: 4 }))
  on('classic.SessionStart', () => ({}))

  await $.classic.SessionStart({ source: 'clear' })

  const ui = await $.ui.mount({ ...PANE, surface: 'terminal' })
  await ui.press({ key: 'tab-pins' })
  expect(await ui.find({ type: 'Text', text: 'Pins: 4' })).toBeDefined()
})

This assumes a $.state version of the pane that draws Pins: N from the pinCount atom and has the classic.SessionStart hook shown in /docs/plugins/mods/interface. Without that hook the pane draws the default Pins: 0 and the test fails.

Testing a policy mod

A mod in your organisation's prependPlugins can refuse other mods as they load. Test it by setting its tier and supplying inline mods for it to judge:

  • tier('prepend') at the top of the file loads your mod as prepend (or append, builtin). Default is user.
  • Pass test an options object with plugins: inline mods, each with name, register, and optionally tier.

Suppose your policy mod acme-guard refuses any mod that calls http.fetch:

import { expect, test, tier } from 'claude-code/testing'

tier('prepend')

const phoneHome = {
  name: 'phone-home',
  register(on) {
    on('tool.call', async ($, e, next) => {
      await $.http.fetch('https://example.com/beacon')
      return next(e)
    })
  },
}

const quiet = {
  name: 'quiet',
  register(on) {
    on('tool.call', async () => ({ result: 'quiet answered' }))
  },
}

test('refuses a mod that makes network calls', { plugins: [phoneHome] }, async ($, on) => {
  on('tool.call', () => ({ result: 'engine' }))
  let message = ''
  try {
    await $.tool.call({ tool: 'Read', file_path: 'a' })   // first call on $ loads the mods
  } catch (error) {
    message = error.message
  }
  expect(message).toMatch(/^phone-home: refused by acme-guard/)
})

test('admits a mod with no network calls', { plugins: [quiet] }, async ($, on) => {
  on('tool.call', () => ({ result: 'engine' }))
  expect(await $.tool.call({ tool: 'Read', file_path: 'a' })).toEqual({ result: 'quiet answered' })
})

All mods load at the test's first call on $. When your mod refuses one, that call throws with a message naming the refused mod, your mod, and your reason (for example phone-home: refused by acme-guard: ...). Writing policy mods themselves is covered in /docs/plugins/mods/admin.