Self-hosted environments reference
Every self-hosted runner and orchestrator flag, environment variable, health endpoint field and Prometheus metric in one place.
This is the lookup page for the two processes you operate in a self-hosted environment: the runner, which executes cloud sessions on your hosts, and the optional orchestrator, which starts runners on demand. Both run on Linux or macOS, which is what defaults such as /workspace and ~/.claude assume. For the version you actually have installed, claude self-hosted-runner --help is the final word.
Naming: environment versus pool
Metric names and some API fields say pool where these pages say environment. They are the same thing. The environment ID is the pool_id field and looks like ccpool_.... Flags and environment variables use environment (for example --environment-secret-file), and the older pool spellings still work with a deprecation notice.
How flags and variables relate
- Most flags have an environment variable twin. If both are set, the flag wins.
- Duration variables are always in milliseconds, marked by an
_MSsuffix, even when the flag takes minutes or seconds.--exit-if-unused-min 10equalsSELF_HOSTED_RUNNER_IDLE_SHUTDOWN_MS=600000. A Helm value ofSELF_HOSTED_RUNNER_STARTUP_TIMEOUT_MS: "15"means 15 milliseconds, not 15 minutes. This one has bitten me. - Duration caps keep timers under the runtime's 32-bit ceiling (about 24.85 days): every
--*-minflag caps at 10080 minutes (7 days),--drain-grace-secat 604800 seconds (7 days), and--drain-wait-secat 86400 seconds (24 hours).--session-stop-grace-secand--post-session-hook-timeout-sechave no cap. Exceed a cap on a flag and startup fails; exceed it in a variable and the value is clamped to the ceiling.
Runner flags
I have grouped the runner's flags by job. Every flag here belongs to claude self-hosted-runner.
Identity and registration
| Flag | Variable | Default | What it does |
|---|---|---|---|
--environment-secret-file <path> | SELF_HOSTED_RUNNER_ENVIRONMENT_SECRET | required | File holding the environment secret, or for orchestrator-spawned runners the single-use work-order JWT. Note the variable holds the value itself, not a path. Older --pool-secret-file and SELF_HOSTED_RUNNER_POOL_SECRET still work with a deprecation notice on stderr; preview builds older than 2.1.216 only know those |
--api-url <url> | none | https://api.anthropic.com | API base URL. Only change for testing |
--client-label <label> | SELF_HOSTED_RUNNER_CLIENT_LABEL | host's hostname | Label sent at registration and exposed as client_label on the info metric. v2.1.248+ |
--lock-to-account <id> | SELF_HOSTED_RUNNER_LOCK_TO_ACCOUNT | unset | Lock to one account at startup rather than on first session. Takes an email or user_... ID in the environment's organisation. A pre-locked runner never takes Claude Tag channel sessions, which have no account |
Workspace and sessions
| Flag | Variable | Default | What it does |
|---|---|---|---|
--base-dir <path> | SELF_HOSTED_RUNNER_BASE_DIR | /workspace (no default on Windows) | Where checkouts and per-session directories live. Created and write-tested at startup; failure exits with cannot create or write to base directory (before v2.1.225 this failed sessions instead). Windows is not a supported host and needs it set explicitly. Keep identical across the environment |
--capacity <n> | none | 1 | Maximum concurrent sessions, all for the same locked owner. Keep identical across the environment |
--exec-path <path> | SELF_HOSTED_RUNNER_EXEC_PATH | the runner's own binary | Binary or wrapper script spawned per session |
--hooks-dir <path> | SELF_HOSTED_RUNNER_HOOKS_DIR | unset | Directory of lifecycle hooks |
--host-config-snapshot <mode> | SELF_HOSTED_RUNNER_HOST_CONFIG_SNAPSHOT | disk | Where the startup snapshot of the host config directory lives. disk copies it under --base-dir and checks every file against an in-memory digest at each session start; a modified file fails that session and the runner refuses sessions until restarted. memory keeps it on the heap up to 64 MiB; above that, sessions start without host config and show a notice. If the disk copy cannot be written, that run falls back to memory. v2.1.271+ |
--remove-session-state [bool] | SELF_HOSTED_RUNNER_REMOVE_SESSION_STATE | off | Delete a session's directories under <base-dir>/_sessions/ when it ends, whatever the outcome. Best-effort (skipped if the runner is killed or hits its drain deadline first). Failed sessions' debug logs are not kept. v2.1.268+ |
--trust-workspace [bool] | SELF_HOSTED_RUNNER_TRUST_WORKSPACE | on | Seeds trust for each session's repository paths so repository-committed permissions.allow and additionalDirectories are honoured. Set false to ignore those grants and put allow rules in the host config's settings.json. Repository sandbox.* settings apply either way |
--confine-repo-settings <mode> | SELF_HOSTED_RUNNER_CONFINE_REPO_SETTINGS | warn | Guard against repository settings that grant access outside the workspace, set environment variables, or override the operator's sandbox or hooks posture (such as sandbox.enabled: false or disableAllHooks). warn logs and continues, enforce refuses the session, off skips the scan. See hardening |
Timeouts and session lifetime
| Flag | Variable | Default | What it does |
|---|---|---|---|
--startup-timeout-min <n> | SELF_HOSTED_RUNNER_STARTUP_TIMEOUT_MS | 15 | Release the slot if the child has not sent its init signal (on the fd 3 activity channel, not ordinary output) within N minutes. After init, idle release takes over. 0 disables |
--release-idle-session-min <n> | SELF_HOSTED_RUNNER_SESSION_IDLE_MS | 0 | Release a session after N idle minutes once a turn ends or it is waiting on the user. Mid-turn sessions (including never-ending background tasks and approvals raised inside a running tool call) are not idle. After a background task finishes, the session counts as busy until the follow-up turn starts, up to SELF_HOSTED_RUNNER_BG_RESULT_GRACE_MS. A release that leaves the runner empty follows the normal exit path governed by --drain-grace-sec (or exits immediately after a deferred first signal). 0 disables |
--kill-session-after-min <n> | SELF_HOSTED_RUNNER_MAX_LIFETIME_MS | 0 | Wall-clock backstop for stuck sessions. From v2.1.260 the session is released at the limit and only terminated if still present when SELF_HOSTED_RUNNER_MAX_LIFETIME_GRACE_MS runs out; earlier versions terminated at the limit. 0 disables |
--exit-if-unused-min <n> | SELF_HOSTED_RUNNER_IDLE_SHUTDOWN_MS | 0 | Exit after N minutes of polling if no work was ever assigned, for autoscaler scale-down. 0 disables |
--session-stop-grace-sec <n> | SELF_HOSTED_RUNNER_SESSION_STOP_GRACE_MS | 5 | How long to wait for the Claude process to exit after a session ends before force-killing. Raise if the session's own SessionEnd hooks need longer |
--post-session-hook-timeout-sec <n> | SELF_HOSTED_RUNNER_POST_SESSION_HOOK_TIMEOUT_MS | 60 | Budget for the post-session hook on every session end, shutdown included |
Shutdown and retirement
| Flag | Variable | Default | What it does |
|---|---|---|---|
--drain-grace-sec <n> | SELF_HOSTED_RUNNER_DRAIN_GRACE_MS | 0 | Before any shutdown signal or retire time: 0 exits as soon as active sessions finish; a positive value keeps polling the locked owner's queue that long first, which weakens per-session isolation. Ignored after a deferred first signal |
--drain-wait-sec <n> | SELF_HOSTED_RUNNER_DRAIN_WAIT_MS | 0 | Once a drain starts, wait up to N seconds for in-flight turns and background tasks before terminating children. A just-finished background task counts as running until its follow-up turn starts, within SELF_HOSTED_RUNNER_BG_RESULT_GRACE_MS |
--defer-shutdown-max-min <n> | SELF_HOSTED_RUNNER_DEFER_SHUTDOWN_MAX_MS | 0 | On the first SIGTERM or SIGINT, keep serving held sessions instead of draining, release what remains after N minutes, then exit. Raise the host's stop timeout first. 0 disables. v2.1.238+. See deferring the drain |
--drain-marker-file <path> | SELF_HOSTED_RUNNER_DRAIN_MARKER_FILE | unset | If this file exists when a drain starts, the runner reports the exit as a host drain rather than a plain signal. Behaviour is otherwise identical. Use a local path sessions cannot write. v2.1.271+ |
--retire-at <epoch-seconds> | SELF_HOSTED_RUNNER_RETIRE_AT | unset | Retire at an absolute Unix time, for hosts destroyed on a schedule without a usable signal. Values before 2001 or after the year 5138 are rejected by the flag and ignored in the variable |
--push-outcome-on-release | SELF_HOSTED_RUNNER_PUSH_OUTCOME_ON_RELEASE | off | On runner-initiated ends (drain, idle release), best-effort push of tracked outcome branches to origin before deleting the workspace. Adds 30 seconds to the shutdown budget; resuming needs git 2.29+. Restrict pushes to claude/* first. Checkouts made by a checkout hook are not pushed |
Git
| Flag | Variable | Default | What it does |
|---|---|---|---|
--configure-git | SELF_HOSTED_RUNNER_CONFIGURE_GIT=1 | off | Write global identity, Anthropic commit signing, push negotiation (v2.1.257+) and Co-authored-by: trailer hooks at startup. See configure git |
--use-anthropic-git-proxy | CLAUDE_RUNNER_USE_GIT_PROXY=1 | off | Clone through the Anthropic git proxy instead of your own credentials. Needs --capacity 1 and git 2.32+, otherwise the runner will not start. Makes the rewrite flags irrelevant |
--git-host-rewrite <from>=<to> | none | unset | Rewrite https://<from>/... to https://<to>/... before cloning, for split-horizon DNS. Repeatable |
--git-ssh-rewrite <host> | none | unset | Rewrite https://<host>/... to git@<host>:... for SSH-only hosts. Repeatable |
Network, health and logging
| Flag | Variable | Default | What it does |
|---|---|---|---|
--proxy-authorization-command <command> | SELF_HOSTED_RUNNER_PROXY_AUTHORIZATION_COMMAND | unset | Command run for every proxy connection; trimmed stdout becomes the Proxy-Authorization value. Needs HTTPS_PROXY or HTTP_PROXY; exclusive with the file variant. v2.1.238+ |
--proxy-authorization-file <path> | SELF_HOSTED_RUNNER_PROXY_AUTHORIZATION_FILE | unset | File read for every proxy connection; trimmed contents become the header. For tokens rotated in place. Same requirements. v2.1.238+ |
--health-port <port> | SELF_HOSTED_RUNNER_HEALTH_PORT | 8080 | Port for /healthz and /metrics. 0 disables |
--log-file <path> | SELF_HOSTED_RUNNER_LOG_FILE | unset | Mirror logs to a 0600 file as well as stdout and stderr. Needed for self-hosted-runner doctor to tail logs locally |
--log-level <level> | none | info | info or debug |
--debug-token-dir <path> | SELF_HOSTED_RUNNER_DEBUG_TOKEN_DIR | unset | Writes live tokens to disk for inspection. Never in production |
A worked example
A typical production line for a single-session, git-proxy runner on Kubernetes:
claude self-hosted-runner \
--environment-secret-file /etc/claude/environment-secret \
--capacity 1 \
--use-anthropic-git-proxy \
--configure-git \
--confine-repo-settings enforce \
--release-idle-session-min 30 \
--kill-session-after-min 480 \
--remove-session-state \
--log-file /var/log/claude-runner.log
Orchestrator flags
claude self-hosted-runner orchestrator launches on-demand runners. It shares --api-url, --environment-secret-file, --hooks-dir, --health-port and --log-level with the runner (same defaults and, where the runner has one, the same variable), except that --hooks-dir is mandatory and must contain a spawn-runner hook. It does not accept the proxy-authorization flags. Its own flags:
| Flag | Default | What it does |
|---|---|---|
--hook-concurrency <n> | 4 | Maximum spawn-runner hooks running at once, and the cap on spawn requests claimed per poll |
--hook-timeout <sec> | 60 | Kill the hook's process tree after this long. Timeout plus its 5-second kill grace must be below --expected-spawn-seconds, checked at startup |
--expected-spawn-seconds <sec> | 120 | Your p99 runner boot time, between 10 and 3600. Sent as the server-side lease; if no runner registers in time, the session is re-offered with a new order ID. Must match on every replica |
--min-idle <n> | 0 | Keep at least N idle session slots by pre-warming standby runners. Pair with the runner's --exit-if-unused-min so surplus standbys exit. 0 disables |
--debug-dir <path> | unset | Write each spawn request's work order and hook stderr to disk. Debug only |
SCM connector flags
The SCM connector is not available yet, so leave these unset. If --scm-connector-host is set, the connection never opens and the orchestrator keeps retrying, though runners still start normally.
It is designed as a standing WebSocket from the orchestrator to Anthropic so that hosted pre-session flows (repository picker, branch and ref resolution) can reach a GitHub Enterprise Server that is only routable inside your network. See GitHub Enterprise Server.
| Flag | Default | What it does |
|---|---|---|
--scm-connector-host <host[:port]> | unset | GitHub Enterprise Server host to forward to; port defaults to 443 |
--scm-connector-id <n> | required with the host | Numeric ID of your organisation's GitHub Enterprise Server connection |
--scm-connector-provider <slug> | ghe | Provider path segment matching ^[a-z0-9-]{1,32}$ |
--scm-connector-ca-file <path> | unset | Extra PEM CA bundle for TLS to the host |
--scm-connector-host-rewrite <from>=<to_host:to_port> | unset | End-to-end testing only: redirects the TCP connection but keeps Host header and SNI |
Each attempt authenticates with the environment secret and retries with exponential backoff capped at 30 seconds plus jitter.
Variables with no flag
| Variable | Default | What it does |
|---|---|---|
SELF_HOSTED_RUNNER_HOST_CONFIG_DIR | ~/.claude | Directory snapshotted at startup and seeded into every session's CLAUDE_CONFIG_DIR. Setting it (even to the default) also moves where .claude.json is read for MCP seeding. An empty directory disables seeding |
SELF_HOSTED_RUNNER_BG_RESULT_GRACE_MS | 30000 | How long a session counts as busy after a background task finishes and before the follow-up turn starts. Applies to drains, idle release and --retire-at. 0 or garbage falls back to the default. v2.1.228+ |
SELF_HOSTED_RUNNER_POST_TURN_SETTLE_MS | 7000 | How long a session counts as busy after a turn, while it reports the turn's end to Anthropic, for drain purposes. Cannot be disabled. v2.1.275+ |
SELF_HOSTED_RUNNER_MAX_LIFETIME_GRACE_MS | 900000 | Grace after --kill-session-after-min before termination |
SELF_HOSTED_RUNNER_SIGKILL_GRACE_MS | 30000 | How long to wait for the OS to deliver SIGKILL to a child stuck in uninterruptible I/O before the runner exits. Floored at the post-session timeout plus 15 seconds (plus 30 with --push-outcome-on-release), so effectively 75 seconds at defaults |
CLAUDE_RUNNER_FETCH_DEPTH | 50 | Depth for fresh clones: a positive number, or full or 0 for everything. Existing clones keep their depth |
CLAUDE_RUNNER_SKIP_GIT_VERIFY | unset | 1 skips the .git check after a checkout hook, for non-git sources |
FORCE_AUTOUPDATE_PLUGINS | unset | 1 lets plugin marketplaces auto-update while the binary stays pinned |
CLAUDE_CODE_DISABLE_ARTIFACT | unset | 1 disables the Artifact tool regardless of the organisation setting and removes the *.frame.claudeusercontent.com egress need |
Telemetry controls
Session children send operational telemetry to Anthropic by default; it never contains code or repository content. Set controls on the runner process, which reapplies them after server-provided variables so yours always win.
| Variable | Effect |
|---|---|
CLAUDE_CODE_BYOC_ENABLE_DATADOG=1 | Opt in to Datadog operational metrics, off by default for self-hosted |
DISABLE_TELEMETRY, DO_NOT_TRACK, DISABLE_ERROR_REPORTING, CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC | The usual Claude Code switches, applied to session children. See environment variables |
DISABLE_GROWTHBOOK=1 | Stops feature-flag fetching only; telemetry stays on unless DISABLE_TELEMETRY is set too |
CLAUDE_CODE_ENABLE_TELEMETRY | Unrelated to Anthropic analytics: turns on OpenTelemetry export to your own collector. See monitoring usage |
Health endpoints
Runner
GET /healthz on the health port returns 200 OK while the process is alive, regardless of the poll loop. The body tells you more:
{
"status": "ok",
"runner_id": "ccrunner_01HZX...",
"active_sessions": 1,
"last_poll_at": "2026-10-08T09:12:44.501Z",
"last_poll_age_ms": 1310
}
last_poll_at and last_poll_age_ms are null until the first poll finishes. A last_poll_age_ms that keeps climbing means the poll loop is stuck, so use it in custom probes. A quick check from the host:
curl -s localhost:8080/healthz | jq '.last_poll_age_ms < 60000'
Orchestrator
The orchestrator's /healthz always returns 200. Gate readiness and alerts on the body's connected field (whether the last poll succeeded), and read per-state spawn queue counts from queue_counts. When the SCM connector is configured the body also has scm_connector_connected and an scm_connector object (connected, last_connected_at, last_error, reconnects, requests_forwarded); both are null otherwise.
Prometheus metrics
Both processes serve GET /metrics on their health port.
Runner series
| Series | Type and meaning |
|---|---|
claude_code_self_hosted_runner_info{runner_id,version,client_label} | Always 1. Fleet inventory and version drift |
claude_code_self_hosted_runner_capacity | Configured --capacity |
claude_code_self_hosted_runner_active_sessions | Sessions running now |
claude_code_self_hosted_runner_locked_account{email} | Appears once locked to a user and issued a token with an act.email claim. Absent for Claude Tag agent locks. The label is an email, so drop or hash it at scrape time (for example with metric_relabel_configs) if your metrics store is widely readable |
claude_code_self_hosted_runner_last_poll_age_seconds | Seconds since the last successful poll. Alert above 60 |
claude_code_self_hosted_runner_poll_errors_total{error_kind} | Poll failures by transport, timeout, 5xx, 429, 4xx. All five exist from start; alert on rate(...[5m]) > 0 |
claude_code_self_hosted_runner_sessions_started_total{client_platform} | Session children spawned, labelled by origin (web_claude_ai, ios, android, desktop_app, claude_code_cli, unknown, and either claude_in_slack or claude-in-slack for Slack, so match =~"claude[-_]in[-_]slack") |
claude_code_self_hosted_runner_sessions_completed_total{client_platform} | Clean ends (see counter semantics) |
claude_code_self_hosted_runner_sessions_failed_total{client_platform} | Failed ends |
claude_code_self_hosted_runner_sessions_interrupted_total{client_platform} | Ends caused by operational termination |
claude_code_self_hosted_runner_initializing_sessions | Sessions between assignment and the child's init event |
claude_code_self_hosted_runner_session_init_duration_seconds | Histogram of init durations |
claude_code_self_hosted_runner_session_init_errors_total | Failures before init: checkout hook, git prep, token issue, early child crash |
claude_code_self_hosted_runner_session_start_hook_errors_total | SessionStart hook executions that reported an error |
claude_code_self_hosted_runner_session_idle_seconds{session_id,client_platform} | Per-session idle seconds; handy for spotting sessions stuck on an unanswered permission prompt |
Orchestrator series
| Series | Type and meaning |
|---|---|
claude_code_self_hosted_orchestrator_info{version,pool_id,orchestrator_uuid,hostname} | Always 1 |
claude_code_self_hosted_orchestrator_connected | 1 if the last poll succeeded, 0 after any failure |
claude_code_self_hosted_orchestrator_last_poll_age_seconds | Seconds since the last poll attempt, successful or not (unlike the runner's version). The loop waits on hooks, so alert around --hook-timeout plus margin, about 90 seconds at defaults |
claude_code_self_hosted_orchestrator_poll_errors_total{error_kind} | Spawn-poll failures by the same five kinds |
claude_code_self_hosted_orchestrator_queue_pending_sessions | Spawn requests claimable now |
claude_code_self_hosted_orchestrator_queue_backing_off_sessions | Requests backing off after a retryable hook failure |
claude_code_self_hosted_orchestrator_queue_circuit_broken_sessions | Requests blocked until an Owner selects Retry in the Activity tab. Alert if above zero |
claude_code_self_hosted_orchestrator_pool_pending_sessions | All sessions waiting on a runner in the environment. Identical on every replica, so aggregate with max, not sum |
claude_code_self_hosted_orchestrator_pool_active_sessions | Sessions on a live runner in the environment. Also use max |
claude_code_self_hosted_orchestrator_spawn_hooks_total{result} | Hook outcomes: ok, retryable, non_retryable. Counts hook calls, not sessions |
claude_code_self_hosted_orchestrator_spawn_hook_duration_seconds | Histogram of hook durations |
claude_code_self_hosted_orchestrator_warm_hints_dispatched_total | Standby (pre-warm) spawn requests dispatched |
claude_code_self_hosted_orchestrator_session_queue_wait_seconds | Histogram of queue wait before the orchestrator claimed a session. Pre-warms are not sampled |
claude_code_self_hosted_orchestrator_clock_skew_seconds | Local minus server clock, once measured |
claude_code_self_hosted_orchestrator_scm_connector_connected | SCM connector WebSocket state; absent without --scm-connector-host |
claude_code_self_hosted_orchestrator_scm_connector_requests_forwarded_total | Requests proxied to the SCM host; absent without --scm-connector-host |
Autoscaling signals
- Scale on queue depth with
claude_code_self_hosted_orchestrator_pool_pending_sessions, notqueue_pending_sessions. - Scale on utilisation with the runner's
active_sessionsdivided bycapacity. - Gate on
claude_code_self_hosted_orchestrator_connected == 1per instance so a disconnected replica's stale number does not drive the scaler.
During a total poll outage the gated query returns nothing. The Kubernetes HPA holds replicas on missing data, but KEDA's Prometheus scaler with its default ignoreNullValues: "true" treats empty as zero and scales in. Set ignoreNullValues: "false" on the ScaledObject, optionally with a fallback floor.
A KEDA trigger that follows that advice:
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus.monitoring:9090
query: |
max(claude_code_self_hosted_orchestrator_pool_pending_sessions
and on(pod) claude_code_self_hosted_orchestrator_connected == 1)
threshold: "1"
ignoreNullValues: "false"
Scraping and alerting
If you use the Prometheus Operator, label runner and orchestrator pods with app.kubernetes.io/part-of: claude-code-self-hosted-runner and name the port health (as the Kubernetes recipe does), then one PodMonitor covers both:
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: claude-runners
namespace: monitoring
spec:
namespaceSelector: { matchNames: [claude-runners] }
selector:
matchLabels: { app.kubernetes.io/part-of: claude-code-self-hosted-runner }
podMetricsEndpoints:
- { port: health, path: /metrics, interval: 30s }
The alerts I would start with, thresholds to taste:
| Alert | Expression | Why |
|---|---|---|
| Runner not polling | claude_code_self_hosted_runner_last_poll_age_seconds > 60 for 2m | Stuck poll loop while /healthz still says 200 |
| Runner poll errors | sum by (pod) (rate(claude_code_self_hosted_runner_poll_errors_total[5m])) > 0 | Proxy or network trouble |
| Init failures | increase(claude_code_self_hosted_runner_session_init_errors_total[10m]) > 3 | Broken hook, git credentials or image |
| SessionStart hook errors | increase(claude_code_self_hosted_runner_session_start_hook_errors_total[10m]) > 3 | A seeded Claude Code hook is failing |
| Version drift | count(count by (version) (claude_code_self_hosted_runner_info)) > 1 for 30m | Half-finished rollout |
| Orchestrator disconnected | claude_code_self_hosted_orchestrator_connected == 0 for 2m | Cannot reach the control plane |
| Orchestrator not polling | claude_code_self_hosted_orchestrator_last_poll_age_seconds > 90 | Hooks hanging |
| Circuit broken | claude_code_self_hosted_orchestrator_queue_circuit_broken_sessions > 0 | Hook returning non-retryable errors; fix, then Retry in the Activity tab |
| Spawn hook failing | sum by (pod) (increase(claude_code_self_hosted_orchestrator_spawn_hooks_total{result!="ok"}[5m])) > 3 | Provisioning trouble |
Passing through session metrics
Each session child has its own OpenTelemetry metrics. With OTEL_METRICS_EXPORTER=prometheus on the runner host and CLAUDE_CODE_ENABLE_TELEMETRY=1 in the session environment (from the wrapper or inherited from the runner), and --capacity above one, the runner re-exposes each child's counters and gauges on its own /metrics. It does this by pointing the child's exporter at a loopback-only OTLP receiver on the health port, adding session_id and client_platform labels, and dropping a session's series when it ends. Histograms are not passed through, and any child metric that would clash with the runner's own prefix is dropped. At --capacity 1 none of this happens; the child binds its own Prometheus endpoint on port 9464.
How the session counters classify endings
Every spawned child increments sessions_started_total at spawn and exactly one of the other three at exit, so started minus the other three equals children running now.
| Counter | Counts |
|---|---|
completed | Child exited with code 0; session archived or deleted while connected; or the runner handed the slot back cleanly (idle release, retire time, --kill-session-after-min release, startup timeout, or a server-side deassign noticed before the child exited) |
failed | Child exited non-zero on its own: a crash or a setup failure after spawn |
interrupted | Runner terminated the child for operational reasons: a drain (such as a Kubernetes rolling restart's SIGTERM), or a session still present when the lifetime grace window closed |
Before v2.1.260 every session hitting --kill-session-after-min was terminated and counted as interrupted.
The post-session hook's CLAUDE_RUNNER_EXIT_REASON reports releases, startup timeouts and server deassigns as interrupted, whereas these counters call them completed. Reconciling hook receipts against sessions_completed_total therefore undercounts completions. Use the hook for per-session guarantees and the counters for rates.
One-shot fleets. At --capacity 1 with --drain-grace-sec 0, a runner exits moments after its only session ends. The three terminal counters only tick at that moment, so a 15 to 60 second scrape rarely sees them before the series disappears. sessions_started_total is visible for the session's life, but on such a fleet it behaves more like "sessions running now". Use these instead:
| Goal | Series |
|---|---|
| Throughput | claude_code_self_hosted_orchestrator_spawn_hooks_total{result="ok"} under rate(). Counts hooks, so pre-warms and respawns diverge from session counts |
| Utilisation | sum(claude_code_self_hosted_runner_active_sessions) against sum(claude_code_self_hosted_runner_capacity) |
| Backlog | claude_code_self_hosted_orchestrator_pool_pending_sessions, plus queue_circuit_broken_sessions above zero |
| Failures | sessions_failed_total is best effort here; treat any non-zero sighting as worth a look. Pre-spawn failures only appear in session_init_errors_total |
The orchestrator series only exist when you run the orchestrator. On a fixed fleet whose runners outlive sessions (--drain-grace-sec above 0), sum(rate(claude_code_self_hosted_runner_sessions_started_total[5m])) works for throughput. Runners export no queue depth, so check backlog in the environment's Activity tab on the Cloud environments admin page. For per-session outcomes, rely on the post-session hook, which fires on every end where a child was spawned except abrupt host loss.