Test self-hosted environments end to end
A CI smoke test for a new self-hosted runner image, dispatching sessions from the CLI and reading replies back through a Stop hook.
Every time I change a runner image (new toolchain, new MCP server, a tweak to the wrapper script) I want proof that a real session still works before it goes near the production environment. The cheapest proof is a scripted round trip against a throwaway test environment: start a session, read Claude's reply, send a follow-up, read that reply too. If both come back with the text you asked for, the image, git access and any custom tools all work.
This page builds that test. It assumes you already have an environment and runner from the quickstart, and that your CI job starts the runner on the same host as the test script, which is the natural shape when you are testing a freshly built image. For runners on separate infrastructure, see remote test runners.
How the test reads replies
The trick is a Claude Code Stop hook on the test runner. When Claude finishes a turn, the hook's JSON input includes the final assistant message as last_assistant_message. The hook appends that text to $E2E_REPLY_DIR/<session_id>.txt, and the test script polls the file. The only Anthropic API calls the test makes are the two dispatches.
CI script ── claude -p ... --environment ──▶ Anthropic queue ──▶ runner on this host
▲ │
└──── polls $E2E_REPLY_DIR/<session_id>.txt ◀── Stop hook ◀──────┘
Install the capture hook
The runner seeds the host's ~/.claude/ into every session, so install the hook there, exactly as you would the push nudge from customising sessions.
Merge this into ~/.claude/settings.json on the runner host:
{
"hooks": {
"Stop": [
{
"hooks": [
{ "type": "command", "timeout": 10, "command": "\"$CLAUDE_CONFIG_DIR/hooks/ci-reply-capture.sh\"" }
]
}
]
}
}
Then save this as ~/.claude/hooks/ci-reply-capture.sh and chmod +x it:
#!/bin/sh
# Test runners only. Appends each turn's final reply to $E2E_REPLY_DIR/<session_id>.txt.
dir="${E2E_REPLY_DIR:-}"
[ -n "$dir" ] && [ -d "$dir" ] || exit 0
# The session exports cse_<id>; the dispatch CLI prints session_<id>.
id="${CLAUDE_CODE_REMOTE_SESSION_ID:-}"
[ -n "$id" ] || exit 0
id="session_${id#cse_}"
# Tool-only turns have no last_assistant_message; write nothing rather than "null".
jq -r '.last_assistant_message // empty' >> "$dir/$id.txt" 2>/dev/null
exit 0
It needs jq on the runner, and it never fails the turn.
Two rules before you start the runner:
- Install first, start second. The runner snapshots
~/.claude/once at startup, so a hook added later only takes effect after a restart. - Export
E2E_REPLY_DIRto the runner process (systemd unit, pod spec, CI step). Without the variable, or if the directory does not exist, the hook quietly does nothing.
Warning: Keep this hook out of production images. It writes every session's final reply to disk whenever the directory exists, which is fine on a disposable CI runner and a data leak waiting to happen anywhere else.
The dispatch commands
Two CLI forms drive the test. Both need Claude Code v2.1.224 or later on the machine running the script.
Create a session on a specific environment:
claude -p "<prompt>" --environment <ccpool_id> --ref <branch> --output-format json
Run it from a git checkout so the CLI can work out the repository from the origin remote. --ref is optional and bases the session's checkout on a named ref instead of your local HEAD. The command creates the session, prints one JSON line containing session_id (plus a link), and exits without waiting for Claude.
Things to know about --environment:
- It overrides the
remote.defaultEnvironmentIdsetting (see the settings reference). - It does not support
--output-format stream-json. - It cannot be combined with flags that resume, attach to or preconfigure a session:
--resume,--continue,--teleport,--session-id,--init-only. --cloudalongside it is rejected when given a session ID or URL, and in non-interactive runs when given a description. A bare--cloudis ignored. From an interactive terminal you can pass the task as the--clouddescription rather than as a positional prompt.
Send a follow-up to an existing session:
claude -p "<message>" --cloud <session_id> --output-format json
This posts a user message to the running session and exits; the JSON includes "ok": true on success. More on follow-ups in Claude Code on the web.
The test script
This version checks that a project-specific tool is installed in the image as well as checking the round trip. Run it from a checkout of the repository you want the session to work in, with a runner already started on this host (hook installed, E2E_REPLY_DIR exported), and after signing in as described in authenticate from CI. Without that sign-in the first dispatch fails with something like Unable to get organization UUID for cloud session creation.
#!/usr/bin/env bash
# smoke-test-runner.sh: round trip against a self-hosted test environment.
set -euo pipefail
ENV_ID="${CLAUDE_TEST_ENVIRONMENT_ID:-${CLAUDE_TEST_POOL_ID:-}}" # POOL_ID is the old name
: "${ENV_ID:?set CLAUDE_TEST_ENVIRONMENT_ID to a ccpool_... id served by this host}"
: "${E2E_REPLY_DIR:?export E2E_REPLY_DIR here and to the runner}"
REF="${TEST_REPO_REF:-main}"
WAIT_SECS="${WAIT_SECS:-120}"
[ -d "$E2E_REPLY_DIR" ] || { echo "E2E_REPLY_DIR=$E2E_REPLY_DIR is missing" >&2; exit 1; }
wait_for() { # wait_for <session_id> <expected text>
local file="$E2E_REPLY_DIR/$1.txt" until=$(( $(date +%s) + WAIT_SECS ))
until [ -f "$file" ] && grep -qF -- "$2" "$file"; do
if [ "$(date +%s)" -ge "$until" ]; then
echo "FAIL: '$2' did not appear in $file after ${WAIT_SECS}s" >&2
ls -la "$E2E_REPLY_DIR" >&2
[ -f "$file" ] && cat "$file" >&2
exit 1
fi
sleep 2
done
}
nonce="ci-$(date +%s)-$$"
# Turn 1: prove the image has our internal CLI on PATH.
out=$(claude -p "[$nonce] Run 'acme-lint --version'. If it succeeds reply with exactly TOOLCHAIN_OK, otherwise TOOLCHAIN_MISSING." \
--environment "$ENV_ID" --ref "$REF" --output-format json)
sid=$(jq -er '.session_id' <<<"$out")
echo "created $sid"
wait_for "$sid" "TOOLCHAIN_OK"
# Turn 2: prove follow-ups reach the same session.
out=$(claude -p "[$nonce] Reply with exactly FOLLOWUP_OK." --cloud "$sid" --output-format json)
jq -e '.ok == true' <<<"$out" >/dev/null
wait_for "$sid" "FOLLOWUP_OK"
echo "PASS ($sid)"
Tune WAIT_SECS to your cold-start time; the first turn includes clone time. Swap the prompts for whatever exercises your setup: calling a custom MCP tool, reaching an internal package registry, pushing to a scratch branch.
Remote test runners
If CI cannot share a filesystem with the runner (say, a persistent Kubernetes fleet), have the hook POST the reply somewhere your driver is listening instead of writing a file:
#!/bin/sh
# Remote variant. Set E2E_REPLY_URL on the runner.
url="${E2E_REPLY_URL:-}"
[ -n "$url" ] || exit 0
id="${CLAUDE_CODE_REMOTE_SESSION_ID:-}"
[ -n "$id" ] || exit 0
jq -r '.last_assistant_message // empty' \
| curl -fsS --max-time 5 -X POST --data-binary @- "$url/session_${id#cse_}" >/dev/null 2>&1
exit 0
On the driver side, anything that accepts the POST and holds the body until the test asks for it will do: a tiny HTTP listener in the CI job or a webhook receiver you already run. The hook runs inside your network, so the endpoint only has to be reachable from the runners, and your egress rules must allow it.
Authenticate from CI
Both --environment dispatches and --cloud follow-ups need a claude.ai OAuth token. API keys (sk-ant-...) are not accepted, and neither is the environment secret, which only lets a runner register.
A long-lived CI host
Sign in once, interactively, with claude auth login on the machine that runs the script, ideally with a dedicated automation account. The token is stored in the macOS keychain, or in ~/.claude/.credentials.json on Linux and Windows (and on macOS when the keychain cannot be written, which is common over SSH because the login keychain stays locked). See authentication.
The CLI refreshes the access token on every run, but the refresh grant expires 30 days after the original login. Put a calendar reminder in to sign in again every 30 days.
Ephemeral CI runners
There is currently no long-lived token for this. The scope that controls cloud sessions, user:sessions:claude_code, is capped server-side at 30 days, so a one-year claude setup-token token (inference only) does not cover it.
To get a stored login onto a fresh runner without a browser, set CLAUDE_CODE_OAUTH_REFRESH_TOKEN and CLAUDE_CODE_OAUTH_SCOPES so claude auth login exchanges the token directly (see environment variables). The 30-day cap still applies. If you need a machine identity not tied to a person, talk to your Anthropic account team.
A fresh environment per CI run
For a fully clean run, create a test environment at the start of the job and delete it at the end. These are the same endpoints the Cloud environments admin page uses, and they need the anthropic-beta: ccr-byoc-2025-07-29 header.
The admin token
ADMIN_TOKEN is a claude.ai OAuth access token for an account with the Owner role, obtained the same way as above: sign in with claude auth login, then read the current access token from the keychain or ~/.claude/.credentials.json. Read it fresh on every run (the CLI rotates it), and pass it to curl on stdin so it never appears in the process list or build log. The -H @- trick needs curl 7.55 or later; older versions send the literal @- and the request goes out unauthenticated.
Create
pool_secret in the response is a long-lived credential that can register runners, so capture it into a masked variable and print only the ID:
api=https://api.anthropic.com/v1/code/runners/self-hosted/pools
hdrs=(-H "anthropic-beta: ccr-byoc-2025-07-29" -H "anthropic-version: 2023-06-01")
resp=$(printf 'Authorization: Bearer %s\n' "$ADMIN_TOKEN" \
| curl -fsS -X POST -H @- "${hdrs[@]}" -H "content-type: application/json" \
-d "{\"name\":\"ci-${GITHUB_RUN_ID:-local}\"}" "$api")
CLAUDE_TEST_ENVIRONMENT_ID=$(jq -er '.pool.pool_id' <<<"$resp")
SELF_HOSTED_RUNNER_ENVIRONMENT_SECRET=$(jq -er '.pool_secret' <<<"$resp")
echo "::add-mask::$SELF_HOSTED_RUNNER_ENVIRONMENT_SECRET"
echo "test environment: $CLAUDE_TEST_ENVIRONMENT_ID"
If an Owner has not yet enabled Allow self-hosted environments, the call returns 403 permission_error with self-hosted runners are disabled by your organization's policy.
Now start the runner on this host with SELF_HOSTED_RUNNER_ENVIRONMENT_SECRET set, the capture hook installed and E2E_REPLY_DIR exported, then run the smoke test.
Delete
Always tear the environment down, ideally in a step that runs even when the test fails:
printf 'Authorization: Bearer %s\n' "$ADMIN_TOKEN" \
| curl -fsS -X DELETE -H @- "${hdrs[@]}" "$api/$CLAUDE_TEST_ENVIRONMENT_ID"