Skip to content

Hosting the Agent SDK

Run Agent SDK agents in production containers, choose a session pattern, persist state, size hosts and isolate tenants that share a machine.

Hosting an Agent SDK app is not like hosting a thin API wrapper. Every query() starts a claude CLI subprocess that owns a shell, a working directory and session files on local disk. Each running agent is a stateful, long-lived process, and that shapes everything about resourcing, persistence and scaling.

This page is about running it on your own infrastructure. If you do not actually need the agent loop on your own machines, Anthropic's Managed Agents product hosts the loop for you, with tools running in an Anthropic-managed sandbox or a self-hosted sandbox. For ready-made Dockerfiles and Kubernetes manifests, Anthropic's claude-cookbooks repository on GitHub has a hosting folder with examples for Docker, Modal and Kubernetes.

The process model

client  ->  your app  --stdio-->  claude CLI subprocess  -->  api.anthropic.com
                                     |
                                     +-- shell, cwd, ~/.claude/projects/*.jsonl
  • One session is one subprocess. Fifty concurrent sessions means fifty process trees and fifty transcript files.
  • Every subprocess inherits your app's working directory unless you pass cwd. If sessions must not see each other's files, give each a distinct cwd.
  • The subprocess never listens on the network. Your app takes inbound requests and calls the SDK.
from claude_agent_sdk import query, ClaudeAgentOptions

async for msg in query(prompt="Extract line items from invoice.pdf",
                       options=ClaudeAgentOptions(cwd=f"/work/jobs/{job_id}")):
    ...

TypeScript examples on this page use top-level await, so save them as .mts or set "type": "module".

State on local disk

None of this survives a container restart, scale-down or reschedule:

StateDefault location
Session transcripts~/.claude/projects/, or projects/ under CLAUDE_CONFIG_DIR
CLAUDE.md memory~/.claude/CLAUDE.md (user) and the working directory (project)
Files the agent madeThe working directory

Transcripts can be mirrored to durable storage with a SessionStore. Memory files and working-directory output need their own plan: a mounted volume or an object-store sync. Sessions explains resume and fork.

Pick a session pattern

PatternContainer lifetimeSuits
EphemeralOne task, then destroyedBug fixes, invoice extraction, translation, media conversion
Long-runningPersistent, many sessions per containerEmail triage agents, chat bots on Slack, per-user live sites
HybridSpun up on demand, rehydrated from a storeIntermittent project assistants, research that pauses for hours, support agents that reload ticket history
Multi-agentSeveral SDK processes cooperating in one containerSimulations where agents share an environment

Ephemeral

The container reads the task from the environment, runs it and exits:

import { query } from "@anthropic-ai/claude-agent-sdk";

try {
  for await (const msg of query({ prompt: process.env.JOB_PROMPT!, options: { maxTurns: 25 } })) {
    if (msg.type === "result") console.log(JSON.stringify({ outcome: msg.subtype, cost: msg.total_cost_usd }));
  }
} catch {
  process.exitCode = 1; // e.g. error_max_turns: query() throws after yielding the result
}

If the turn cap is hit, the result subtype is error_max_turns and query() throws after yielding it, so catch it if the container must exit cleanly. See the agent loop for every subtype.

Long-running

The container exposes HTTP or WebSocket and maps each active session to a long-lived query:

  • TypeScript: add turns with streamInput() on the query object. Pre-warm subprocesses before traffic arrives with startup(), or with prewarm() if you do not know a session's cwd until its first request.
  • Python: hold a session open with ClaudeSDKClient.

Size the container to hold your peak number of concurrent sessions in memory.

Hybrid

Containers spin down when idle and rehydrate from a SessionStore when the user returns. The store is mandatory here: shutting down without one throws away the transcript. Tune your provider's idle timeout to how often users come back.

async for msg in query(prompt=user_message,
                       options=ClaudeAgentOptions(resume=session_id_from_db, session_store=store)):
    ...

In TypeScript the options are resume and sessionStore.

Multi-agent

Give each agent its own working directory and isolate settings loading so one agent's CLAUDE.md does not leak into another. The tenant isolation options below apply equally here.

Provisioning

Choosing a sandbox

Run the SDK inside a container for process isolation, resource limits, network control and a throwaway filesystem. When comparing providers, consider:

  • Managed or self-run. Sandbox-as-a-service versus software you operate.
  • Cold start. Ephemeral patterns need sub-second starts; long-running ones can wait.
  • Durable storage. Hybrid sessions need it somewhere.
  • Billing. Per-second suits bursty ephemeral work; hourly suits always-on sessions.
  • Networking. Custom egress rules, outbound proxies, private VPC peering.

Self-hosted options such as Docker, gVisor and Firecracker are compared in secure deployment.

Runtime

  • Python 3.10+ for the Python SDK, or Node.js 18+ for TypeScript.
  • Both SDKs bundle a native Claude Code binary for most installs; the CLI needs no separate Node.js. The quickstart lists the cases that need a separate native install.
  • The bundled binary is pinned to the SDK version, so upgrading the SDK upgrades the CLI. The SDKs follow semver: take patches freely, read the changelog (on the claude-agent-sdk-typescript and claude-agent-sdk-python GitHub repos) before a minor.

Resources

Start from 1 GiB RAM, 5 GiB disk and 1 CPU per agent. That is an idle baseline: memory grows with session length and tool use, so measure.

Network

  • Outbound HTTPS to api.anthropic.com, or your provider's regional endpoint on Bedrock or Google Cloud's Agent Platform.
  • Outbound access to any MCP servers or external tools.
  • In production, route egress through a proxy that enforces a domain allowlist, injects credentials and logs requests.
  • Inbound, expose one HTTP or WebSocket port for your app.

Production checklist

Persistence

Mirror transcripts with a SessionStore adapter. The session storage page has reference adapters for an object store, a key-value store and a database, plus a conformance suite for your own. Three behaviours to know:

  • Transcripts only. CLAUDE.md and working files are not mirrored.
  • A mirror, not a replacement. The subprocess writes locally first and the SDK forwards each batch. A fresh session's local transcript outlives the run; a run resumed from the store deletes its local copy at the end, leaving the store as the only durable copy.
  • mirror_error. If a batch cannot be delivered, it is dropped, a { type: "system", subtype: "mirror_error" } message is emitted, and the query continues. Alert on these if durability matters.

Observability

Set OpenTelemetry variables at the container or orchestrator level and every query() exports:

CLAUDE_CODE_ENABLE_TELEMETRY=1
CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1   # traces only
OTEL_TRACES_EXPORTER=otlp
OTEL_METRICS_EXPORTER=otlp
OTEL_LOGS_EXPORTER=otlp
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector.observability:4318

Prompt text and tool inputs are excluded by default. See observability.

Credentials

  • Model API. The subprocess reads ANTHROPIC_API_KEY. Supply it from your secret manager, or set ANTHROPIC_BASE_URL to a proxy that injects the key outside the container (see secure deployment).
  • Inbound auth. Authenticate at a gateway in front of the agent. The agent should receive pre-authenticated requests, not validate user tokens.
  • Tool credentials. Keep them out of the agent's environment entirely; a proxy adds them after the request leaves the container.

Scaling

Concurrency per host is bounded by memory:

agents per host = (host RAM - overhead) / per-session peak RAM

Measure the peak by running a realistic session to your target length under realistic tool load and recording peak RSS. The 1 GiB starting point is a floor.

For long-running containers behind a load balancer, pin each session to one container with consistent hashing on sessionId, so it keeps hitting the same warm subprocess until evicted or restarted.

Cost

Token spend usually outweighs container cost by an order of magnitude or more. A minimal container is roughly $0.05 an hour; a single long session can spend dollars. Track it with cost tracking.

Tenant isolation

By default the SDK reads settings and CLAUDE.md files from disk, which in a shared container can leak one tenant's context into another's session. For each tenant:

StepHow
Skip filesystem settingssettingSources: [] / setting_sources=[]
Disable auto memoryCLAUDE_CODE_DISABLE_AUTO_MEMORY=1 in env. Auto memory under ~/.claude/projects/<project>/memory/ loads regardless of setting sources
Separate global configPoint CLAUDE_CONFIG_DIR at a per-tenant directory
Short transcript pathsOptionally set CLAUDE_CODE_PROJECT_DIR_NAME when each config directory serves one working directory and you use no SessionStore (TypeScript SDK v0.3.234+, Python SDK v0.2.140+)
Separate filesA per-tenant cwd on every call
Separate egressPer-tenant outbound IPs, credentials or allowlists at the proxy
for await (const msg of query({
  prompt,
  options: {
    cwd: `/tenants/${tenantId}/work`,
    settingSources: [],
    env: {
      ...process.env,
      CLAUDE_CONFIG_DIR: `/tenants/${tenantId}/config`,
      CLAUDE_CODE_DISABLE_AUTO_MEMORY: "1"
    }
  }
})) { /* ... */ }

Claude Code features in the SDK lists the other inputs that load whatever your setting sources say.

Known limitations

LimitationWorkaround
Sessions never time out on their ownBound them with maxTurns / max_turns
Memory grows over long sessionsCap session length or recycle subprocesses
Wide parallel subagent fan-outs can hit rate limitsDispatch in smaller batches
No wall-clock limit per subagentSet maxTurns in each AgentDefinition. CLAUDE_ASYNC_AGENT_STALL_TIMEOUT_MS is a stall watchdog for subagents that stop producing output, not a total runtime limit

When it works locally but not deployed

  • CLI not found at start-up. In Python, the service manager's PATH differs from your shell's. In TypeScript, the image build skipped optional dependencies, or pathToClaudeCodeExecutable points at a missing file.
  • CLI present but will not launch. Wrong architecture or libc for the container, or the execute bit was lost during the build.
  • CLI exits mid-run. The error you see depends on language and on whether an error result arrived first.

All three are covered in detail on the troubleshooting page.