Skip to content

Deploy self-hosted environments to production

Harden, network, image, size, run and troubleshoot a fleet of self-hosted Claude Code runners for production use.

Getting one runner to pick up one session is the easy part. Running a fleet that every developer in your organisation can send model-directed code to is a different job, and this page is the checklist I work through before an environment touches anything real. It assumes you already have an environment and a working runner from the quickstart, and that you understand the environment, runner and session model from self-hosted environments.

Note: Self-hosted environments are in public beta on Team and Enterprise plans. An Owner has to switch them on before any of this applies.

The order below is the order I would do it in: lock things down, open only the egress you need, sort out git, build the image, size it, pick a recipe, then learn how it fails.

Hardening checklist

Treat a runner host as a machine that anyone in your Anthropic organisation can run arbitrary code on. That is literally true: any organisation member can dispatch a session to any environment, and if an Owner has routed Claude Tag channels to the environment, anyone that Claude Tag's access setting admits (by default, anyone in the connected Slack workspace, account or not) can start sessions there too. There is no per-environment access control on dispatch.

With that framing, each item below follows naturally.

ControlWhat to doWhy
One container per sessionRun each runner in a fresh container or VM destroyed on exit, with --capacity 1 and the default --drain-grace-sec 0Nothing from one session survives into the next. Higher capacity or a positive drain grace lets one container serve several sessions for the same locked owner
Kill leftovers by destroying the boxDo not rely on the runner to clean up daemonsWhen a session stops, the runner does not signal processes that outlived their shell command (a daemonised service, say). Only tearing down the container ends them
No broad credentials baked inKeep long-lived SSH keys, cloud credentials and wide-scope tokens out of the imageEvery session the image runs can read them. Mint per-session credentials in a wrapper script; handle the first clone with a checkout hook or the Anthropic git proxy
Keep the environment secret away from session hostsPrefer on-demand runnersThe secret can register runners and claim any queued session. On a fixed fleet it sits on every host where session code can read it, so rotate it after any suspected compromise
Default-deny egressRestrict outbound traffic at your own network edgeSee Default-deny egress
Least-privilege host identityThe instance profile or node service account grants only what the runner process needsSessions should get their own credentials, not inherit the host's
Block the metadata endpointIMDSv2 with hop limit 1, GKE Workload Identity with metadata concealment, or an explicit deny for 169.254.169.254 inside the session's network namespaceSubnet egress rules do not catch link-local metadata traffic. The block also applies to wrapper scripts and hooks, so do token exchange via the session JWT against your own token service, or use file-based web identity such as IRSA on EKS
Per-runner filesystem isolationGive each runner its own private working directory. Make --hooks-dir, the wrapper script and the host's ~/.claude/ read-only to sessionsA session must not be able to rewrite the code that sets up the next session
Repository settings guardPick a mode with --confine-repo-settings: warn (default, logs and continues), enforce (refuses the session) or offStops a repo's committed settings from widening a session's reach

The repo-settings guard scans committed settings for three things:

  • Any grant that resolves outside the session's own workspace: an additionalDirectories entry, an Edit, Write or NotebookEdit rule under permissions.allow, or a sandbox.filesystem.allowWrite or allowRead path.
  • An env block that is not empty.
  • An override of the operator's posture, such as sandbox.enabled: false.

It runs whether or not --trust-workspace is set. It does not look at repository hooks, .mcp.json or Bash rules; permissions and tool approval explains where those grants should live instead.

--lock-to-account limits which account's sessions a particular host will run, but it does not narrow who can dispatch to the environment. If you want self-hosted environments to be the only option in the picker, an Owner can hide the Anthropic-hosted ones for the whole organisation on the Cloud environments admin page.

Warning: Your organisation's IP allowlist does not apply to self-hosted runner traffic by default. Do not count it as a control; enforce egress at your own boundary, and talk to your Anthropic account team if you need IP allowlist enforcement.

Network requirements

All traffic is outbound. Allow these and nothing else beyond the internal services your sessions genuinely need.

Always needed:

  • api.anthropic.com on 443. HTTPS for the control plane, session streaming, model inference, feature flags, analytics, JWKS fetches, commit signing and the git proxy (when enabled). It also carries the orchestrator's SCM connector tunnel over WSS when --scm-connector-host is set.
  • Your git host (for example github.com or your GitHub Enterprise hostname) on 443 or 22, for clone and push. Not needed when the runner uses --use-anthropic-git-proxy.

Needed only in certain setups:

HostWhen you need it
downloads.claude.aiInstalling or updating Claude Code with the native installer, and at runtime only if sessions install plugins from the official Anthropic marketplace. The install.sh script itself comes from claude.ai
storage.googleapis.comPlugin install counts and metadata shown in /plugin
claude.com and its code documentation subdomainDocumentation lookups by the built-in guide agent and pre-approved WebFetch calls. Blocking them only breaks doc lookups
*.frame.claudeusercontent.comOnly when the Artifact tool is available to sessions in your organisation. CLAUDE_CODE_DISABLE_ARTIFACT=1 on the runner keeps it off
registry.npmjs.orgPlugin installs (npm-sourced packages and plugin Node dependencies) and npx-launched MCP servers
http-intake.logs.us5.datadoghq.comAnthropic operational metrics, only with CLAUDE_CODE_BYOC_ENABLE_DATADOG=1 (off by default here)
browser-intake-us5-datadoghq.comError report uploads when error reporting is on for the account. DISABLE_ERROR_REPORTING=1 or DISABLE_TELEMETRY=1 suppress it
Cloud provider endpoints such as bedrock-runtime.<region>.amazonaws.com or aiplatform.googleapis.comOnly when the runner sends model requests to Bedrock or Agent Platform

Some older enterprise checklists list statsig.anthropic.com, *.sentry.io, claude.ai, platform.claude.com and mcp-proxy.anthropic.com. Runners and sessions do not need any of them. Feature flags come from api.anthropic.com, the runner authenticates with the environment secret rather than OAuth, and organisation connectors are delivered through api.anthropic.com.

Two host-level tasks do talk to claude.ai: fetching install.sh, and interactive claude auth login (used by guided setup, signed-in doctor and CI dispatch, which also reaches claude.com and platform.claude.com). Do those from a host with wider egress rather than opening session containers up.

Default-deny egress

Put runners and session containers in a segment or namespace whose outbound rules allow only the hosts above, your git host, and named internal services. Claude Code cannot check this for you. It matters regardless of permission mode, because Bash is in the default pre-approved tool set: a session can run curl without asking anyone.

Telemetry details and how to turn each stream off are in the reference.

Proxies that need a Proxy-Authorization header

If your corporate proxy wants a rotating token in Proxy-Authorization, keep HTTPS_PROXY or HTTP_PROXY pointing at the proxy as normal and add one of these (Claude Code v2.1.238 or later):

  • --proxy-authorization-command <command>: the runner runs the command and uses trimmed stdout as the header value, such as Bearer abc123. Good for tokens you generate on demand.
  • --proxy-authorization-file <path>: the runner reads the file and uses its trimmed contents. Good when another process rewrites the token in place.

The runner refuses to start if you set both (a flag plus the other flag's environment variable counts as both), if neither HTTPS_PROXY nor HTTP_PROXY (either case) holds an http:// or https:// URL (ALL_PROXY is ignored), or if you pass either to the orchestrator subcommand. Give it to each runner the orchestrator launches instead.

When one of these flags is active, the runner:

  • Starts a local forward proxy on 127.0.0.1 before registering, and exits if it cannot.
  • Rewrites whichever of HTTPS_PROXY and HTTP_PROXY you set so it points at that listener, for itself, its hooks and its sessions.
  • Fetches the header afresh for every upstream connection, so rotated tokens apply without a restart.
  • Strips ALL_PROXY and any proxy variable spelling you did not set from session environments, and pins NO_PROXY to its own value.
  • Never logs the header.

Here is the shape I use when a sidecar refreshes the token:

export HTTPS_PROXY=http://egress.corp.internal:3128
claude self-hosted-runner \
  --environment-secret-file /etc/claude/env-secret \
  --proxy-authorization-file /var/run/proxy-token/header

Configure git

The runner checks out repositories but leaves identity and credentials to you. Pick one of two approaches.

Git version floors on the host: --configure-git commit signing needs 2.34+, --use-anthropic-git-proxy needs 2.32+, and resuming from branches pushed by --push-outcome-on-release needs 2.29+. Without any of those, 2.24 is enough.

Option 1: let the runner write the config

Pass --configure-git (or SELF_HOSTED_RUNNER_CONFIGURE_GIT=1). At startup the runner writes global git config containing:

  • user.name = Claude and user.email = noreply@anthropic.com, the same identity as Anthropic-hosted sessions.
  • SSH-format commit and tag signing via a runner-managed shim that signs through Anthropic's signing service with the session's own credentials. GitHub verifies the signatures against Anthropic's published SSH signing key.
  • push.negotiate = true (v2.1.257+).
  • core.hooksPath pointing at runner-managed commit-msg and prepare-commit-msg hooks that append a Co-authored-by: trailer for the session creator, built from CCR_SESSION_ACCOUNT_EMAIL (skipped when that variable is unset). If your image already sets core.hooksPath, the runner leaves yours alone, skips its hooks and prints a [runner:git] warning.

The runner exits at startup if git is older than 2.34. This option does not set up push credentials. From v2.1.280, commits made inside checkout or post-session hooks are signed as the session too, minus the trailer.

Option 2: ship your own config in the image

Any commit needs an identity, so set one at system level so it applies whatever user runs the runner. I usually use a bot identity:

RUN git config --system user.name "acme-claude-bot" \
 && git config --system user.email "claude-bot@acme.example"

Without it, git commit fails with Please tell me who you are. The runner never overrides these values.

For push credentials, the right pattern is a short-lived, narrowly scoped token minted per session by your wrapper script, using the session creator's identity from the session JWT, inside a --capacity 1 throwaway container. If you must put something in the image (a read-only deploy key, for example), keep it tight:

  • An SSH deploy key for a single repository combined with a url.<base>.insteadOf rewrite.
  • A credential.helper that hands back a minimally scoped token.
  • GIT_SSH_COMMAND pointing at a narrow key.

Whatever you use must work without prompting, because the runner's own clone and fetch disable every prompt: it sets GIT_TERMINAL_PROMPT=0, appends BatchMode=yes to SSH (including your GIT_SSH_COMMAND), sets GCM_INTERACTIVE=never, and clears core.askPass (so use the GIT_ASKPASS variable instead). None of these settings are passed into the session's environment. Keep any program or key referenced by GIT_SSH_COMMAND or GIT_ASKPASS somewhere sessions cannot write.

If credentials are missing or rejected, the runner retries a few times and then fails repository preparation for the repository the session pushes to. Read-only repositories may be skipped instead (see troubleshooting).

When checkout directories belong to a different uid than the runner, add git config --system --add safe.directory '*'.

The Anthropic git proxy

--use-anthropic-git-proxy (or CLAUDE_RUNNER_USE_GIT_PROXY=1) clones through Anthropic's git proxy using the session's short-lived token: the creator's stored GitHub or GitHub Enterprise OAuth token for normal sessions, or your organisation's GitHub App installation token for bot and agent sessions. The image then needs no git credentials at all. This is the same path Anthropic-hosted sessions use.

Constraints:

  • --capacity 1 is mandatory because the proxy URL is per session. The runner refuses to start otherwise, printing a [runner:fatal] line.
  • Git 2.32 or later.
  • Anthropic must be able to reach your git host. For a host routable only inside your network, use a checkout hook.
  • --git-host-rewrite and --git-ssh-rewrite do nothing, since the URL points at api.anthropic.com.

From v2.1.267 the runner reports the opt-in at registration and prints Registering as opted in to Anthropic-managed git (--use-anthropic-git-proxy). Each session then uses either Anthropic-managed git or the per-session proxy URL; the latter produces one [runner:warn] line.

Warning: The Kubernetes and Compose recipes below use --capacity 4. Add the git proxy without dropping to --capacity 1 and the runner will exit on every restart.

A built-in gh without the GitHub CLI. On runners using Anthropic-managed git with v2.1.287 or later, Claude Code can supply a built-in gh that supports only gh api (GitHub's REST API). It fills in {owner} and {repo} for the current repository and sends requests through Anthropic-managed git, so the image needs no GitHub token. Anthropic decides per session whether it is available; the runner's [runner:session] governed git ACTIVE line shows gh_path_shim=true when it is. Install jq if you want --jq. Setting CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC removes it. If the real GitHub CLI is in the image, sessions use that. For example, to list open PRs:

gh api repos/{owner}/{repo}/pulls?state=open

Private certificate authorities (v2.1.283+). If a TLS-inspecting proxy re-signs traffic, either install your CA in the host's system store or set GIT_SSL_CAINFO to a PEM bundle. GIT_SSL_NO_VERIFY does not help: the runner's own clone through Anthropic-managed git still verifies certificates. With GIT_SSL_CAINFO set, the runner builds a per-session certificate file combining the host's system bundle (from /etc/ssl/certs/ca-certificates.crt or /etc/pki/tls/certs/ca-bundle.crt) and your file, which must be readable, contain PEM CERTIFICATE blocks and be at most 1 MiB. Git inside the session gets http.sslCAInfo and per-URL sslCAInfo/sslVerify entries in place of the variables; checkout and post-session hooks inherit GIT_SSL_CAINFO unchanged. If the file cannot be built, a [runner:warn] line containing did not build the certificate file says why. Each session also logs a governed git: GIT_SSL_CAINFO is set (or GIT_SSL_NO_VERIFY is set) line explaining what was done and whether you need to act.

Rewriting URLs on private networks

Repository URLs arrive as HTTPS using your git host's public name (for GitHub Enterprise, the hostname configured in the GitHub Enterprise integration). Two repeatable flags adjust them:

  • --git-host-rewrite <from>=<to> for split-horizon DNS, for example --git-host-rewrite ghe.acme.com=ghe.internal.acme.net.
  • --git-ssh-rewrite <host> for SSH-only hosts, turning https://<host>/owner/repo into git@<host>:owner/repo.

Host rewriting happens first, so give --git-ssh-rewrite the internal name when you use both.

Build the runner image

There is no published runner image; you build one around the claude binary plus whatever your repositories need. Here is a starting Dockerfile for a Node and Python shop:

FROM debian:bookworm-slim
ARG CLAUDE_CODE_VERSION
RUN apt-get update \
 && apt-get install -y --no-install-recommends git openssh-client curl ca-certificates nodejs npm python3 python3-venv jq \
 && rm -rf /var/lib/apt/lists/*
RUN curl -fsSL -o /usr/local/bin/claude \
      "https://downloads.claude.ai/claude-code-releases/${CLAUDE_CODE_VERSION:?pass --build-arg CLAUDE_CODE_VERSION}/linux-x64/claude" \
 && chmod 0755 /usr/local/bin/claude
RUN git config --system user.name "acme-claude-bot" \
 && git config --system user.email "claude-bot@acme.example" \
 && git config --system --add safe.directory '*'
ENTRYPOINT ["claude"]

Notes:

  • Platform segments are linux-x64, linux-arm64, linux-x64-musl and linux-arm64-musl. Musl images such as Alpine need extra packages; see setup.
  • You can verify the binary against the release's signed manifest, as described in setup.
  • The runner needs Claude Code 2.1.224 or later.

Build with the current stable version looked up at build time, or pin a number for reproducibility:

VERSION="$(curl -fsSL https://downloads.claude.ai/claude-code-releases/stable)"
docker build --build-arg CLAUDE_CODE_VERSION="$VERSION" -t registry.acme.example/claude-runner:"$VERSION" .

Swap stable for latest in that URL when you need something newer than stable, for instance a release a new model requires.

Size CPU and memory

Size for the sessions, not the runner. The runner only polls, prepares checkouts, runs hooks and supervises children. The weight is each session: a Claude Code process plus every build, test run, package install and MCP server it starts.

Per session, a sensible starting point:

  • Memory: 4 GiB request and 4 GiB limit (Claude Code's minimum). Keep them equal; hitting the limit gets processes killed, which can end a session mid-task.
  • CPU: 2 CPU request, 4 CPU limit. CPU limits throttle rather than kill, so bursts during builds just slow down.
resources:
  requests: { cpu: "2", memory: 4Gi }
  limits:   { cpu: "4", memory: 4Gi }

Then measure a representative build of your biggest repository and raise anything that leaves no headroom for Claude Code on top of the build's peak.

--capacity caps concurrent sessions but does not split resources between them. At --capacity 1 (and for on-demand runners, where you put the values on whatever your spawn-runner hook submits) use one session's values. Above 1, multiply by the capacity, or cap individual sessions from your wrapper script.

Kubernetes recipe

The runner serves GET /healthz on port 8080 (change with --health-port). It returns 200 whenever the process is alive, so probes only catch a dead process. To catch a runner that has stopped polling, alert on the last_poll_age_seconds metric from /metrics (see the reference).

This manifest uses --capacity 1, one session per pod, and adds resources and a grace period long enough for a drain:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: claude-runner
  namespace: claude-runners
spec:
  replicas: 6
  selector:
    matchLabels: { app: claude-runner }
  template:
    metadata:
      labels:
        app: claude-runner
        app.kubernetes.io/part-of: claude-code-self-hosted-runner
    spec:
      terminationGracePeriodSeconds: 120
      containers:
        - name: runner
          image: registry.acme.example/claude-runner:2.1.290
          args: ["self-hosted-runner",
                 "--environment-secret-file", "/etc/claude/environment-secret",
                 "--capacity", "1",
                 "--configure-git"]
          resources:
            requests: { cpu: "2", memory: 4Gi }
            limits:   { cpu: "4", memory: 4Gi }
          ports:
            - { name: health, containerPort: 8080 }
          readinessProbe:
            httpGet: { path: /healthz, port: 8080 }
            periodSeconds: 10
          livenessProbe:
            httpGet: { path: /healthz, port: 8080 }
            initialDelaySeconds: 30
            periodSeconds: 30
          volumeMounts:
            - { name: env-secret, mountPath: /etc/claude, readOnly: true }
      volumes:
        - name: env-secret
          secret: { secretName: claude-runner-environment-secret }

Create the namespace, then the Secret from a file so the value never lands in shell history. Run (umask 077 && cat > ./environment-secret), paste the key copied from the admin UI, press Enter then Ctrl-D:

kubectl create namespace claude-runners
kubectl -n claude-runners create secret generic claude-runner-environment-secret \
  --from-file=environment-secret=./environment-secret
rm ./environment-secret

With --capacity 1 and --drain-grace-sec 0 the runner exits after each session and the kubelet starts a new container. If you would rather have a brand new pod (and its own disposable volumes) per session, on-demand runners creating a Job per session are the cleaner fit, and they keep the environment secret off session hosts as well.

Docker Compose recipe

Compose is fine for evaluation. restart: always restarts the same container with its writable layer intact, so the runner comes back on a reused filesystem, which is not the per-session isolation you want in production.

services:
  runner:
    image: registry.acme.example/claude-runner:2.1.290
    command: ["self-hosted-runner",
              "--environment-secret-file", "/run/secrets/env_secret",
              "--capacity", "1"]
    secrets: [env_secret]
    restart: always
    stop_grace_period: 90s
secrets:
  env_secret:
    file: ./environment-secret

Docker backs off between restarts of a container that keeps exiting, so a broken runner does not spin in a tight loop.

Shutdown timing

At startup the runner logs how long it needs to shut down cleanly, in a line containing This runner needs up to 80s at defaults. Set your stop timeout (terminationGracePeriodSeconds, stop_grace_period or equivalent) to at least that. Kubernetes defaults to 30 seconds, which is too short.

On SIGTERM the runner stops accepting work and drains (unless deferred, see below):

  1. Waits up to --drain-wait-sec (default 0) for running turns.
  2. Terminates each session's process tree, allowing --session-stop-grace-sec (default 5).
  3. Runs the post-session hook, allowing --post-session-hook-timeout-sec (default 60).

It keeps polling throughout, so its sessions stay assigned to it while the hook saves work. The logged figure is those three values plus 15 seconds of headroom, plus 30 more when --push-outcome-on-release is set: 0 + 5 + 60 + 15 = 80 at defaults. Capacity does not add to it because sessions drain in parallel.

Because --drain-wait-sec defaults to 0, a rolling restart cuts off running turns, and the session resumes elsewhere without unpushed work. Raise it (and the stop timeout) if you want turns to finish.

If you use --retire-at, leave room between the retire time and the host's end of life for typical turns, the background-task wait, and the logged figure. Compute the epoch value at each launch, for example $(( $(date +%s) + 3*3600 )) for a three-hour host.

Deferring the drain

--defer-shutdown-max-min <n> (v2.1.238+) lets a runner keep serving the sessions it holds for up to n minutes after the first SIGTERM or SIGINT, while taking no new work and polling to keep its lease:

  • First n minutes: sessions run normally; --startup-timeout-min and --kill-session-after-min still apply. With --release-idle-session-min set, idle sessions are released; otherwise they stay.
  • When n runs out: every remaining session is released, after waiting for any mid-turn session's turn to end plus up to 60 seconds for its background tasks.
  • When the post-release grace runs out (75 seconds by default, or --drain-wait-sec + 15 if that is above 60): whatever remains is drained and requeued immediately.

The runner exits 0 as soon as it holds nothing. A second signal skips straight to a drain, and a signal during a drain force-exits. Size the stop timeout to n minutes plus 155 seconds at defaults (the runner prints the figure). If the host kills it early, remaining sessions get no post-session hook and requeue within a few minutes; if you cannot afford that timeout, leave the flag off.

How signals reach a running post-session hook

The post-session hook and each session child live in their own process groups, separate from the runner:

  • Signal during a drain: force-exits the runner. Nothing signals a running hook, so on a bare host it carries on unsupervised (its timeout no longer applies, and writing to the closed log pipe can kill it with SIGPIPE, so redirect its output to a file). In containers where the runner is PID 1, or under systemd's default KillMode=control-group, treat a forced exit as fatal to the hook.
  • Process-group signals (kill -- -<pid>, job control): reach the runner and an in-flight checkout hook, but not a running post-session hook or the session.
  • Cgroup-wide kills (systemd's default, Kubernetes' SIGKILL at the end of the grace period): reach everything. This is why the grace period must cover the whole drain.
  • Hook timeout: the runner sends SIGTERM to the hook's process group, then SIGKILL two seconds later, so forked workers like tar or rsync die with it. Supervision ends when the hook's stdio closes.

The runner logs how many post-session hooks are still running when a drain starts and on a forced exit.

Keep base directory and capacity identical

If a runner dies, another runner picks the session up and derives the checkout path from its own --base-dir and --capacity: capacity 1 checks out directly under the base directory, higher capacity uses per-session worktrees. Mismatched values change the working directory, and every absolute path the agent remembered breaks. Use the same values on every runner in an environment and never a per-host value like the hostname.

--base-dir defaults to /workspace (with one exception noted in the reference). The runner creates and write-tests it before registering, and exits with cannot create or write to base directory if it cannot. Root creates /workspace itself; for a non-root runner, create it and chown it first, or point --base-dir somewhere that user owns.

Pre-warmed checkouts for big repositories

At --capacity 1 with no checkout hook, the runner keeps one canonical clone per repository at <base-dir>/<repo-owner>/<repo>. Each session fetches the ref, detaches HEAD and hard-resets, which is near-instant when little has changed. You can skip the cold clone entirely by:

  • Baking the clone into the image at that path, so every fresh container starts warm without disk reuse.
  • Using a persistent volume for --base-dir on runners pre-locked with --lock-to-account, so the disk only ever serves one account. Pre-locked runners never take Claude Tag channel sessions.

What to expect:

  • Any clone shape works and is kept as-is; the runner never passes --depth to an existing clone. CLAUDE_RUNNER_FETCH_DEPTH (full, 0 or a number, default 50) only affects the cold clone.
  • Tracked changes are reset; untracked files survive because the runner never runs git clean.
  • Per-session data under <base-dir>/_sessions/ (the session's Claude config directory with a transcript copy, uploaded files, worktrees and checkout hook checkouts) is left in place by default and readable by later sessions on the same disk. Size persistent volumes for that growth.
  • --remove-session-state deletes those per-session directories as each session ends, best-effort. The canonical clone and files written elsewhere remain.
  • With the git proxy, .git/ is sanitised before each session (objects, refs and shallow state kept, index deleted), so you pay a full working-tree checkout but never a re-clone. Submodule pre-warms are not supported with the proxy.
  • Git operations have a 120-second no-progress watchdog and a 30-minute hard cap, so long clones that keep progressing complete.

Pin the Claude Code version

Sessions run the runner's own binary with auto-update turned off, so the version you install is the version every session uses until the runner restarts.

  • To hold a version, build the image pinned, or on a bare host install a specific version and disable auto-updates (see setup).
  • To upgrade, rebuild or reinstall, then restart runners.
  • Plugin marketplaces do not auto-update either; set FORCE_AUTOUPDATE_PLUGINS=1 to let plugins update while the binary stays pinned.

Check that every model your sessions use supports the pinned version (see model configuration), otherwise requests fail with "Claude Code does not support this model" (see errors).

Scale the fleet

Because each runner locks to one owner, the minimum replica count is the number of users and Claude Tag agents you expect to be active at once. --capacity adds parallelism within an owner, not across owners. Two ways to scale:

  • Fixed fleet: a static set of replicas, scaled on the Prometheus metrics each runner exposes.
  • On-demand: claude self-hosted-runner orchestrator polls for queued sessions with no runner and calls your spawn-runner hook to start one per session. See on-demand runners.

Known limitations

Connector traffic does not originate in your network. Claude.ai connectors (GitHub, Slack, Linear and so on) are called from Anthropic's infrastructure via api.anthropic.com. To keep one out of self-hosted sessions, use allowedMcpServers and deniedMcpServers (see managed MCP). Those policies apply to delivered connectors too, so a URL allowlist blocks them unless you also allow these paths:

  • https://api.anthropic.com/v2/ccr-sessions/*
  • https://api.anthropic.com/v1/code/sessions/*
  • https://api.anthropic.com/v1/code/mcp/*

If tool traffic must stay internal, run local MCP servers in the image instead.

Some sessions never count as idle. A session with a never-ending background task, or one waiting on an approval raised inside a running tool call, is not idle, so --release-idle-session-min will not free it. Always pair it with --kill-session-after-min (for example 480 for eight hours). From v2.1.260 that limit is soft: the runner gives a grace window (15 minutes by default, set with SELF_HOSTED_RUNNER_MAX_LIFETIME_GRACE_MS), releasing the session when it next waits on its user or when its turn ends (waiting up to 60 seconds for background tasks), and only terminates it if it is still there when the window closes. Earlier versions terminated at the limit.

Resumed sessions lose unpushed work. A fresh runner re-clones from the starting branch. --push-outcome-on-release makes a best-effort push of the session's outcome branches before release so committed work survives; uncommitted changes are still lost. Before turning it on, restrict who can push claude/* refs, because the runner fetches the previously pushed branch on resume without checking who pushed it.

Repositories added mid-session can fail to clone. Claude uses plain git clone over HTTPS, so without the git proxy the clone fails if nothing on the host can read the repository. Select every repository you need when creating the session.

Unconnected connectors are invisible. A connector you have not connected in claude.ai Settings does not appear and you are not prompted. Connectors added during a session only show up in a new session.

For issues, contact your Anthropic account team.

Troubleshooting

Start with the doctor, which opens an interactive Claude Code session with the runner's logs and state attached:

claude self-hosted-runner doctor

Sign in on that host with claude auth login first so it can query the environment, runners and queue. Without that (for example with API key auth) it can only see the local health endpoint, metrics and the runner log, and reads the log only if the runner was started with --log-file.

SymptomLikely cause and fix
Runner never appears in the environmentCheck HTTPS to api.anthropic.com, that the secret is current, and that the clock is within five minutes. Auth failures log [runner:fatal] with the reason
Exits with cannot create or write to base directoryFix ownership of --base-dir or point it somewhere writable. A [runner:fatal] saying the check timed out means a hung NFS or CSI mount. Both print to stderr before --log-file opens, so read container logs
Sessions stay queuedEvery runner is locked to someone else. Check the claude_code_self_hosted_runner_locked_account metric or the locked_account field in [runner:health] lines. The email appears only if the session token had an act.email claim (Claude Tag agent sessions never do; you will see locked_account=yes). Add replicas or wait for drains. With on-demand runners, check the orchestrator
Sessions fail right after pickupOpen the session in claude.ai/code for the error. Usually missing git credentials or missing build tools
No network through an authenticating proxyThe header source failed, timed out after 30 seconds, or returned empty; the runner answers 502 Bad Gateway and logs why. Run your command by hand to check it prints the full header. could not start the proxy-authorization listener means the loopback listener failed
Poll failed with rejecting the malformed poll responseSomething (an intercepting proxy, captive portal) replaced the response. Counted under transport in claude_code_self_hosted_runner_poll_errors_total. Make the proxy pass api.anthropic.com responses through untouched
Branch no longer existsFor read-only sources the runner skips it. For the push target (often a merged and auto-deleted branch) the session fails asking you to restore the branch
Session starts without one of its repositoriesWith no checkout hook, a clear refusal (not found, no credentials, auth failure) on a read-only repository is skipped with a [runner:warn] could not access context source line. Network errors, timeouts and HTTP 403 still fail the start. The check reruns on every start
Sessions take minutes to startConfirm with claude_code_self_hosted_runner_session_init_duration_seconds, then pre-warm or lower CLAUDE_RUNNER_FETCH_DEPTH
Turns fail with 401The runner refreshes CLAUDE_CODE_OAUTH_TOKEN over the session's stdin after a 401 or 403 (the failed turn is not retried), and logs inference_token refresh failed while retrying. Failures starting about 30 minutes in mean a wrapper cut stdin; see keep stdin and fd 3 attached
Pod killed mid-drainRaise terminationGracePeriodSeconds to at least the logged figure

Logging: once initialised, the lifecycle log (including [runner:fatal]) goes to stdout and debug output to stderr, all plain text. Capture both with --log-file or your platform's log collector. Each session child writes its own debug log; on failure its tail appears in claude.ai/code, and unless --remove-session-state is set the full log stays on disk with its path printed.

When the runner exits

Never restart an on-demand runner; its work order is single-use. For the rest, tell two cases apart:

  • Normal exit: sessions finished, retire time reached, or told to stop. Restart it.
  • Failed start: it exits within seconds, every time. Restarting faster does nothing; someone needs to read the output.

A failed start prints a reason, usually a [runner:fatal] line, or an error: line followed by a pointer to --help for flag parsing, unreadable secrets and base directory problems. For example:

[runner:fatal] --use-anthropic-git-proxy requires --capacity 1 (the proxy URL is per-session and linked worktrees share origin). Omit --use-anthropic-git-proxy or set --capacity 1.

Other clues:

  • No line at all usually means the host killed it, for example for exceeding a memory limit.
  • Exit codes do not distinguish permanent configuration errors from transient ones. Base backoff on how quickly it exited.
  • A healthy-looking environment can still list a runner for a few minutes after a post-registration step (such as --configure-git or git proxy setup) failed. If sessions queue while the page says Healthy, check whether your supervisor is restart-looping.
  • RegisterRunner auth failed means a revoked or mistyped secret.
  • The environment's Activity tab filling with new runners that never take work also points to a restart loop.
  • If it starts by hand but not under your supervisor, compare user, home directory, PATH and memory limit. --configure-git and the git proxy need git on PATH and a writable ~/.gitconfig.

Supervisors and backoff:

  • Kubernetes: the kubelet's growing restart delay already applies. Short-lived normal exits also trigger it, so a frequently draining runner can show CrashLoopBackOff. Read the last run with kubectl logs --previous -n claude-runners deploy/claude-runner.
  • Docker / Compose: restart: always backs off too. docker inspect --format '{{.RestartCount}}' <container> climbing means a loop.
  • systemd: RestartSec is constant and also delays normal restarts; five starts in 10 seconds hits the default rate limit and the unit stays stopped until its interval passes or you run systemctl reset-failed. Pick a balanced value and alert on restart count.
  • Your own loop: start at five seconds, double after each run shorter than a minute up to five minutes, and reset to five seconds after a run of a minute or more.