Skip to content

Self-hosted environments reference

Every self-hosted runner and orchestrator flag, environment variable, health endpoint field and Prometheus metric in one place.

This is the lookup page for the two processes you operate in a self-hosted environment: the runner, which executes cloud sessions on your hosts, and the optional orchestrator, which starts runners on demand. Both run on Linux or macOS, which is what defaults such as /workspace and ~/.claude assume. For the version you actually have installed, claude self-hosted-runner --help is the final word.

Naming: environment versus pool

Metric names and some API fields say pool where these pages say environment. They are the same thing. The environment ID is the pool_id field and looks like ccpool_.... Flags and environment variables use environment (for example --environment-secret-file), and the older pool spellings still work with a deprecation notice.

How flags and variables relate

  • Most flags have an environment variable twin. If both are set, the flag wins.
  • Duration variables are always in milliseconds, marked by an _MS suffix, even when the flag takes minutes or seconds. --exit-if-unused-min 10 equals SELF_HOSTED_RUNNER_IDLE_SHUTDOWN_MS=600000. A Helm value of SELF_HOSTED_RUNNER_STARTUP_TIMEOUT_MS: "15" means 15 milliseconds, not 15 minutes. This one has bitten me.
  • Duration caps keep timers under the runtime's 32-bit ceiling (about 24.85 days): every --*-min flag caps at 10080 minutes (7 days), --drain-grace-sec at 604800 seconds (7 days), and --drain-wait-sec at 86400 seconds (24 hours). --session-stop-grace-sec and --post-session-hook-timeout-sec have no cap. Exceed a cap on a flag and startup fails; exceed it in a variable and the value is clamped to the ceiling.

Runner flags

I have grouped the runner's flags by job. Every flag here belongs to claude self-hosted-runner.

Identity and registration

FlagVariableDefaultWhat it does
--environment-secret-file <path>SELF_HOSTED_RUNNER_ENVIRONMENT_SECRETrequiredFile holding the environment secret, or for orchestrator-spawned runners the single-use work-order JWT. Note the variable holds the value itself, not a path. Older --pool-secret-file and SELF_HOSTED_RUNNER_POOL_SECRET still work with a deprecation notice on stderr; preview builds older than 2.1.216 only know those
--api-url <url>nonehttps://api.anthropic.comAPI base URL. Only change for testing
--client-label <label>SELF_HOSTED_RUNNER_CLIENT_LABELhost's hostnameLabel sent at registration and exposed as client_label on the info metric. v2.1.248+
--lock-to-account <id>SELF_HOSTED_RUNNER_LOCK_TO_ACCOUNTunsetLock to one account at startup rather than on first session. Takes an email or user_... ID in the environment's organisation. A pre-locked runner never takes Claude Tag channel sessions, which have no account

Workspace and sessions

FlagVariableDefaultWhat it does
--base-dir <path>SELF_HOSTED_RUNNER_BASE_DIR/workspace (no default on Windows)Where checkouts and per-session directories live. Created and write-tested at startup; failure exits with cannot create or write to base directory (before v2.1.225 this failed sessions instead). Windows is not a supported host and needs it set explicitly. Keep identical across the environment
--capacity <n>none1Maximum concurrent sessions, all for the same locked owner. Keep identical across the environment
--exec-path <path>SELF_HOSTED_RUNNER_EXEC_PATHthe runner's own binaryBinary or wrapper script spawned per session
--hooks-dir <path>SELF_HOSTED_RUNNER_HOOKS_DIRunsetDirectory of lifecycle hooks
--host-config-snapshot <mode>SELF_HOSTED_RUNNER_HOST_CONFIG_SNAPSHOTdiskWhere the startup snapshot of the host config directory lives. disk copies it under --base-dir and checks every file against an in-memory digest at each session start; a modified file fails that session and the runner refuses sessions until restarted. memory keeps it on the heap up to 64 MiB; above that, sessions start without host config and show a notice. If the disk copy cannot be written, that run falls back to memory. v2.1.271+
--remove-session-state [bool]SELF_HOSTED_RUNNER_REMOVE_SESSION_STATEoffDelete a session's directories under <base-dir>/_sessions/ when it ends, whatever the outcome. Best-effort (skipped if the runner is killed or hits its drain deadline first). Failed sessions' debug logs are not kept. v2.1.268+
--trust-workspace [bool]SELF_HOSTED_RUNNER_TRUST_WORKSPACEonSeeds trust for each session's repository paths so repository-committed permissions.allow and additionalDirectories are honoured. Set false to ignore those grants and put allow rules in the host config's settings.json. Repository sandbox.* settings apply either way
--confine-repo-settings <mode>SELF_HOSTED_RUNNER_CONFINE_REPO_SETTINGSwarnGuard against repository settings that grant access outside the workspace, set environment variables, or override the operator's sandbox or hooks posture (such as sandbox.enabled: false or disableAllHooks). warn logs and continues, enforce refuses the session, off skips the scan. See hardening

Timeouts and session lifetime

FlagVariableDefaultWhat it does
--startup-timeout-min <n>SELF_HOSTED_RUNNER_STARTUP_TIMEOUT_MS15Release the slot if the child has not sent its init signal (on the fd 3 activity channel, not ordinary output) within N minutes. After init, idle release takes over. 0 disables
--release-idle-session-min <n>SELF_HOSTED_RUNNER_SESSION_IDLE_MS0Release a session after N idle minutes once a turn ends or it is waiting on the user. Mid-turn sessions (including never-ending background tasks and approvals raised inside a running tool call) are not idle. After a background task finishes, the session counts as busy until the follow-up turn starts, up to SELF_HOSTED_RUNNER_BG_RESULT_GRACE_MS. A release that leaves the runner empty follows the normal exit path governed by --drain-grace-sec (or exits immediately after a deferred first signal). 0 disables
--kill-session-after-min <n>SELF_HOSTED_RUNNER_MAX_LIFETIME_MS0Wall-clock backstop for stuck sessions. From v2.1.260 the session is released at the limit and only terminated if still present when SELF_HOSTED_RUNNER_MAX_LIFETIME_GRACE_MS runs out; earlier versions terminated at the limit. 0 disables
--exit-if-unused-min <n>SELF_HOSTED_RUNNER_IDLE_SHUTDOWN_MS0Exit after N minutes of polling if no work was ever assigned, for autoscaler scale-down. 0 disables
--session-stop-grace-sec <n>SELF_HOSTED_RUNNER_SESSION_STOP_GRACE_MS5How long to wait for the Claude process to exit after a session ends before force-killing. Raise if the session's own SessionEnd hooks need longer
--post-session-hook-timeout-sec <n>SELF_HOSTED_RUNNER_POST_SESSION_HOOK_TIMEOUT_MS60Budget for the post-session hook on every session end, shutdown included

Shutdown and retirement

FlagVariableDefaultWhat it does
--drain-grace-sec <n>SELF_HOSTED_RUNNER_DRAIN_GRACE_MS0Before any shutdown signal or retire time: 0 exits as soon as active sessions finish; a positive value keeps polling the locked owner's queue that long first, which weakens per-session isolation. Ignored after a deferred first signal
--drain-wait-sec <n>SELF_HOSTED_RUNNER_DRAIN_WAIT_MS0Once a drain starts, wait up to N seconds for in-flight turns and background tasks before terminating children. A just-finished background task counts as running until its follow-up turn starts, within SELF_HOSTED_RUNNER_BG_RESULT_GRACE_MS
--defer-shutdown-max-min <n>SELF_HOSTED_RUNNER_DEFER_SHUTDOWN_MAX_MS0On the first SIGTERM or SIGINT, keep serving held sessions instead of draining, release what remains after N minutes, then exit. Raise the host's stop timeout first. 0 disables. v2.1.238+. See deferring the drain
--drain-marker-file <path>SELF_HOSTED_RUNNER_DRAIN_MARKER_FILEunsetIf this file exists when a drain starts, the runner reports the exit as a host drain rather than a plain signal. Behaviour is otherwise identical. Use a local path sessions cannot write. v2.1.271+
--retire-at <epoch-seconds>SELF_HOSTED_RUNNER_RETIRE_ATunsetRetire at an absolute Unix time, for hosts destroyed on a schedule without a usable signal. Values before 2001 or after the year 5138 are rejected by the flag and ignored in the variable
--push-outcome-on-releaseSELF_HOSTED_RUNNER_PUSH_OUTCOME_ON_RELEASEoffOn runner-initiated ends (drain, idle release), best-effort push of tracked outcome branches to origin before deleting the workspace. Adds 30 seconds to the shutdown budget; resuming needs git 2.29+. Restrict pushes to claude/* first. Checkouts made by a checkout hook are not pushed

Git

FlagVariableDefaultWhat it does
--configure-gitSELF_HOSTED_RUNNER_CONFIGURE_GIT=1offWrite global identity, Anthropic commit signing, push negotiation (v2.1.257+) and Co-authored-by: trailer hooks at startup. See configure git
--use-anthropic-git-proxyCLAUDE_RUNNER_USE_GIT_PROXY=1offClone through the Anthropic git proxy instead of your own credentials. Needs --capacity 1 and git 2.32+, otherwise the runner will not start. Makes the rewrite flags irrelevant
--git-host-rewrite <from>=<to>noneunsetRewrite https://<from>/... to https://<to>/... before cloning, for split-horizon DNS. Repeatable
--git-ssh-rewrite <host>noneunsetRewrite https://<host>/... to git@<host>:... for SSH-only hosts. Repeatable

Network, health and logging

FlagVariableDefaultWhat it does
--proxy-authorization-command <command>SELF_HOSTED_RUNNER_PROXY_AUTHORIZATION_COMMANDunsetCommand run for every proxy connection; trimmed stdout becomes the Proxy-Authorization value. Needs HTTPS_PROXY or HTTP_PROXY; exclusive with the file variant. v2.1.238+
--proxy-authorization-file <path>SELF_HOSTED_RUNNER_PROXY_AUTHORIZATION_FILEunsetFile read for every proxy connection; trimmed contents become the header. For tokens rotated in place. Same requirements. v2.1.238+
--health-port <port>SELF_HOSTED_RUNNER_HEALTH_PORT8080Port for /healthz and /metrics. 0 disables
--log-file <path>SELF_HOSTED_RUNNER_LOG_FILEunsetMirror logs to a 0600 file as well as stdout and stderr. Needed for self-hosted-runner doctor to tail logs locally
--log-level <level>noneinfoinfo or debug
--debug-token-dir <path>SELF_HOSTED_RUNNER_DEBUG_TOKEN_DIRunsetWrites live tokens to disk for inspection. Never in production

A worked example

A typical production line for a single-session, git-proxy runner on Kubernetes:

claude self-hosted-runner \
  --environment-secret-file /etc/claude/environment-secret \
  --capacity 1 \
  --use-anthropic-git-proxy \
  --configure-git \
  --confine-repo-settings enforce \
  --release-idle-session-min 30 \
  --kill-session-after-min 480 \
  --remove-session-state \
  --log-file /var/log/claude-runner.log

Orchestrator flags

claude self-hosted-runner orchestrator launches on-demand runners. It shares --api-url, --environment-secret-file, --hooks-dir, --health-port and --log-level with the runner (same defaults and, where the runner has one, the same variable), except that --hooks-dir is mandatory and must contain a spawn-runner hook. It does not accept the proxy-authorization flags. Its own flags:

FlagDefaultWhat it does
--hook-concurrency <n>4Maximum spawn-runner hooks running at once, and the cap on spawn requests claimed per poll
--hook-timeout <sec>60Kill the hook's process tree after this long. Timeout plus its 5-second kill grace must be below --expected-spawn-seconds, checked at startup
--expected-spawn-seconds <sec>120Your p99 runner boot time, between 10 and 3600. Sent as the server-side lease; if no runner registers in time, the session is re-offered with a new order ID. Must match on every replica
--min-idle <n>0Keep at least N idle session slots by pre-warming standby runners. Pair with the runner's --exit-if-unused-min so surplus standbys exit. 0 disables
--debug-dir <path>unsetWrite each spawn request's work order and hook stderr to disk. Debug only

SCM connector flags

The SCM connector is not available yet, so leave these unset. If --scm-connector-host is set, the connection never opens and the orchestrator keeps retrying, though runners still start normally.

It is designed as a standing WebSocket from the orchestrator to Anthropic so that hosted pre-session flows (repository picker, branch and ref resolution) can reach a GitHub Enterprise Server that is only routable inside your network. See GitHub Enterprise Server.

FlagDefaultWhat it does
--scm-connector-host <host[:port]>unsetGitHub Enterprise Server host to forward to; port defaults to 443
--scm-connector-id <n>required with the hostNumeric ID of your organisation's GitHub Enterprise Server connection
--scm-connector-provider <slug>gheProvider path segment matching ^[a-z0-9-]{1,32}$
--scm-connector-ca-file <path>unsetExtra PEM CA bundle for TLS to the host
--scm-connector-host-rewrite <from>=<to_host:to_port>unsetEnd-to-end testing only: redirects the TCP connection but keeps Host header and SNI

Each attempt authenticates with the environment secret and retries with exponential backoff capped at 30 seconds plus jitter.

Variables with no flag

VariableDefaultWhat it does
SELF_HOSTED_RUNNER_HOST_CONFIG_DIR~/.claudeDirectory snapshotted at startup and seeded into every session's CLAUDE_CONFIG_DIR. Setting it (even to the default) also moves where .claude.json is read for MCP seeding. An empty directory disables seeding
SELF_HOSTED_RUNNER_BG_RESULT_GRACE_MS30000How long a session counts as busy after a background task finishes and before the follow-up turn starts. Applies to drains, idle release and --retire-at. 0 or garbage falls back to the default. v2.1.228+
SELF_HOSTED_RUNNER_POST_TURN_SETTLE_MS7000How long a session counts as busy after a turn, while it reports the turn's end to Anthropic, for drain purposes. Cannot be disabled. v2.1.275+
SELF_HOSTED_RUNNER_MAX_LIFETIME_GRACE_MS900000Grace after --kill-session-after-min before termination
SELF_HOSTED_RUNNER_SIGKILL_GRACE_MS30000How long to wait for the OS to deliver SIGKILL to a child stuck in uninterruptible I/O before the runner exits. Floored at the post-session timeout plus 15 seconds (plus 30 with --push-outcome-on-release), so effectively 75 seconds at defaults
CLAUDE_RUNNER_FETCH_DEPTH50Depth for fresh clones: a positive number, or full or 0 for everything. Existing clones keep their depth
CLAUDE_RUNNER_SKIP_GIT_VERIFYunset1 skips the .git check after a checkout hook, for non-git sources
FORCE_AUTOUPDATE_PLUGINSunset1 lets plugin marketplaces auto-update while the binary stays pinned
CLAUDE_CODE_DISABLE_ARTIFACTunset1 disables the Artifact tool regardless of the organisation setting and removes the *.frame.claudeusercontent.com egress need

Telemetry controls

Session children send operational telemetry to Anthropic by default; it never contains code or repository content. Set controls on the runner process, which reapplies them after server-provided variables so yours always win.

VariableEffect
CLAUDE_CODE_BYOC_ENABLE_DATADOG=1Opt in to Datadog operational metrics, off by default for self-hosted
DISABLE_TELEMETRY, DO_NOT_TRACK, DISABLE_ERROR_REPORTING, CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFICThe usual Claude Code switches, applied to session children. See environment variables
DISABLE_GROWTHBOOK=1Stops feature-flag fetching only; telemetry stays on unless DISABLE_TELEMETRY is set too
CLAUDE_CODE_ENABLE_TELEMETRYUnrelated to Anthropic analytics: turns on OpenTelemetry export to your own collector. See monitoring usage

Health endpoints

Runner

GET /healthz on the health port returns 200 OK while the process is alive, regardless of the poll loop. The body tells you more:

{
  "status": "ok",
  "runner_id": "ccrunner_01HZX...",
  "active_sessions": 1,
  "last_poll_at": "2026-10-08T09:12:44.501Z",
  "last_poll_age_ms": 1310
}

last_poll_at and last_poll_age_ms are null until the first poll finishes. A last_poll_age_ms that keeps climbing means the poll loop is stuck, so use it in custom probes. A quick check from the host:

curl -s localhost:8080/healthz | jq '.last_poll_age_ms < 60000'

Orchestrator

The orchestrator's /healthz always returns 200. Gate readiness and alerts on the body's connected field (whether the last poll succeeded), and read per-state spawn queue counts from queue_counts. When the SCM connector is configured the body also has scm_connector_connected and an scm_connector object (connected, last_connected_at, last_error, reconnects, requests_forwarded); both are null otherwise.

Prometheus metrics

Both processes serve GET /metrics on their health port.

Runner series

SeriesType and meaning
claude_code_self_hosted_runner_info{runner_id,version,client_label}Always 1. Fleet inventory and version drift
claude_code_self_hosted_runner_capacityConfigured --capacity
claude_code_self_hosted_runner_active_sessionsSessions running now
claude_code_self_hosted_runner_locked_account{email}Appears once locked to a user and issued a token with an act.email claim. Absent for Claude Tag agent locks. The label is an email, so drop or hash it at scrape time (for example with metric_relabel_configs) if your metrics store is widely readable
claude_code_self_hosted_runner_last_poll_age_secondsSeconds since the last successful poll. Alert above 60
claude_code_self_hosted_runner_poll_errors_total{error_kind}Poll failures by transport, timeout, 5xx, 429, 4xx. All five exist from start; alert on rate(...[5m]) > 0
claude_code_self_hosted_runner_sessions_started_total{client_platform}Session children spawned, labelled by origin (web_claude_ai, ios, android, desktop_app, claude_code_cli, unknown, and either claude_in_slack or claude-in-slack for Slack, so match =~"claude[-_]in[-_]slack")
claude_code_self_hosted_runner_sessions_completed_total{client_platform}Clean ends (see counter semantics)
claude_code_self_hosted_runner_sessions_failed_total{client_platform}Failed ends
claude_code_self_hosted_runner_sessions_interrupted_total{client_platform}Ends caused by operational termination
claude_code_self_hosted_runner_initializing_sessionsSessions between assignment and the child's init event
claude_code_self_hosted_runner_session_init_duration_secondsHistogram of init durations
claude_code_self_hosted_runner_session_init_errors_totalFailures before init: checkout hook, git prep, token issue, early child crash
claude_code_self_hosted_runner_session_start_hook_errors_totalSessionStart hook executions that reported an error
claude_code_self_hosted_runner_session_idle_seconds{session_id,client_platform}Per-session idle seconds; handy for spotting sessions stuck on an unanswered permission prompt

Orchestrator series

SeriesType and meaning
claude_code_self_hosted_orchestrator_info{version,pool_id,orchestrator_uuid,hostname}Always 1
claude_code_self_hosted_orchestrator_connected1 if the last poll succeeded, 0 after any failure
claude_code_self_hosted_orchestrator_last_poll_age_secondsSeconds since the last poll attempt, successful or not (unlike the runner's version). The loop waits on hooks, so alert around --hook-timeout plus margin, about 90 seconds at defaults
claude_code_self_hosted_orchestrator_poll_errors_total{error_kind}Spawn-poll failures by the same five kinds
claude_code_self_hosted_orchestrator_queue_pending_sessionsSpawn requests claimable now
claude_code_self_hosted_orchestrator_queue_backing_off_sessionsRequests backing off after a retryable hook failure
claude_code_self_hosted_orchestrator_queue_circuit_broken_sessionsRequests blocked until an Owner selects Retry in the Activity tab. Alert if above zero
claude_code_self_hosted_orchestrator_pool_pending_sessionsAll sessions waiting on a runner in the environment. Identical on every replica, so aggregate with max, not sum
claude_code_self_hosted_orchestrator_pool_active_sessionsSessions on a live runner in the environment. Also use max
claude_code_self_hosted_orchestrator_spawn_hooks_total{result}Hook outcomes: ok, retryable, non_retryable. Counts hook calls, not sessions
claude_code_self_hosted_orchestrator_spawn_hook_duration_secondsHistogram of hook durations
claude_code_self_hosted_orchestrator_warm_hints_dispatched_totalStandby (pre-warm) spawn requests dispatched
claude_code_self_hosted_orchestrator_session_queue_wait_secondsHistogram of queue wait before the orchestrator claimed a session. Pre-warms are not sampled
claude_code_self_hosted_orchestrator_clock_skew_secondsLocal minus server clock, once measured
claude_code_self_hosted_orchestrator_scm_connector_connectedSCM connector WebSocket state; absent without --scm-connector-host
claude_code_self_hosted_orchestrator_scm_connector_requests_forwarded_totalRequests proxied to the SCM host; absent without --scm-connector-host

Autoscaling signals

  • Scale on queue depth with claude_code_self_hosted_orchestrator_pool_pending_sessions, not queue_pending_sessions.
  • Scale on utilisation with the runner's active_sessions divided by capacity.
  • Gate on claude_code_self_hosted_orchestrator_connected == 1 per instance so a disconnected replica's stale number does not drive the scaler.

During a total poll outage the gated query returns nothing. The Kubernetes HPA holds replicas on missing data, but KEDA's Prometheus scaler with its default ignoreNullValues: "true" treats empty as zero and scales in. Set ignoreNullValues: "false" on the ScaledObject, optionally with a fallback floor.

A KEDA trigger that follows that advice:

triggers:
  - type: prometheus
    metadata:
      serverAddress: http://prometheus.monitoring:9090
      query: |
        max(claude_code_self_hosted_orchestrator_pool_pending_sessions
          and on(pod) claude_code_self_hosted_orchestrator_connected == 1)
      threshold: "1"
      ignoreNullValues: "false"

Scraping and alerting

If you use the Prometheus Operator, label runner and orchestrator pods with app.kubernetes.io/part-of: claude-code-self-hosted-runner and name the port health (as the Kubernetes recipe does), then one PodMonitor covers both:

apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
  name: claude-runners
  namespace: monitoring
spec:
  namespaceSelector: { matchNames: [claude-runners] }
  selector:
    matchLabels: { app.kubernetes.io/part-of: claude-code-self-hosted-runner }
  podMetricsEndpoints:
    - { port: health, path: /metrics, interval: 30s }

The alerts I would start with, thresholds to taste:

AlertExpressionWhy
Runner not pollingclaude_code_self_hosted_runner_last_poll_age_seconds > 60 for 2mStuck poll loop while /healthz still says 200
Runner poll errorssum by (pod) (rate(claude_code_self_hosted_runner_poll_errors_total[5m])) > 0Proxy or network trouble
Init failuresincrease(claude_code_self_hosted_runner_session_init_errors_total[10m]) > 3Broken hook, git credentials or image
SessionStart hook errorsincrease(claude_code_self_hosted_runner_session_start_hook_errors_total[10m]) > 3A seeded Claude Code hook is failing
Version driftcount(count by (version) (claude_code_self_hosted_runner_info)) > 1 for 30mHalf-finished rollout
Orchestrator disconnectedclaude_code_self_hosted_orchestrator_connected == 0 for 2mCannot reach the control plane
Orchestrator not pollingclaude_code_self_hosted_orchestrator_last_poll_age_seconds > 90Hooks hanging
Circuit brokenclaude_code_self_hosted_orchestrator_queue_circuit_broken_sessions > 0Hook returning non-retryable errors; fix, then Retry in the Activity tab
Spawn hook failingsum by (pod) (increase(claude_code_self_hosted_orchestrator_spawn_hooks_total{result!="ok"}[5m])) > 3Provisioning trouble

Passing through session metrics

Each session child has its own OpenTelemetry metrics. With OTEL_METRICS_EXPORTER=prometheus on the runner host and CLAUDE_CODE_ENABLE_TELEMETRY=1 in the session environment (from the wrapper or inherited from the runner), and --capacity above one, the runner re-exposes each child's counters and gauges on its own /metrics. It does this by pointing the child's exporter at a loopback-only OTLP receiver on the health port, adding session_id and client_platform labels, and dropping a session's series when it ends. Histograms are not passed through, and any child metric that would clash with the runner's own prefix is dropped. At --capacity 1 none of this happens; the child binds its own Prometheus endpoint on port 9464.

How the session counters classify endings

Every spawned child increments sessions_started_total at spawn and exactly one of the other three at exit, so started minus the other three equals children running now.

CounterCounts
completedChild exited with code 0; session archived or deleted while connected; or the runner handed the slot back cleanly (idle release, retire time, --kill-session-after-min release, startup timeout, or a server-side deassign noticed before the child exited)
failedChild exited non-zero on its own: a crash or a setup failure after spawn
interruptedRunner terminated the child for operational reasons: a drain (such as a Kubernetes rolling restart's SIGTERM), or a session still present when the lifetime grace window closed

Before v2.1.260 every session hitting --kill-session-after-min was terminated and counted as interrupted.

The post-session hook's CLAUDE_RUNNER_EXIT_REASON reports releases, startup timeouts and server deassigns as interrupted, whereas these counters call them completed. Reconciling hook receipts against sessions_completed_total therefore undercounts completions. Use the hook for per-session guarantees and the counters for rates.

One-shot fleets. At --capacity 1 with --drain-grace-sec 0, a runner exits moments after its only session ends. The three terminal counters only tick at that moment, so a 15 to 60 second scrape rarely sees them before the series disappears. sessions_started_total is visible for the session's life, but on such a fleet it behaves more like "sessions running now". Use these instead:

GoalSeries
Throughputclaude_code_self_hosted_orchestrator_spawn_hooks_total{result="ok"} under rate(). Counts hooks, so pre-warms and respawns diverge from session counts
Utilisationsum(claude_code_self_hosted_runner_active_sessions) against sum(claude_code_self_hosted_runner_capacity)
Backlogclaude_code_self_hosted_orchestrator_pool_pending_sessions, plus queue_circuit_broken_sessions above zero
Failuressessions_failed_total is best effort here; treat any non-zero sighting as worth a look. Pre-spawn failures only appear in session_init_errors_total

The orchestrator series only exist when you run the orchestrator. On a fixed fleet whose runners outlive sessions (--drain-grace-sec above 0), sum(rate(claude_code_self_hosted_runner_sessions_started_total[5m])) works for throughput. Runners export no queue depth, so check backlog in the environment's Activity tab on the Cloud environments admin page. For per-session outcomes, rely on the post-session hook, which fires on every end where a child was spawned except abrupt host loss.