Hosting the Agent SDK
Run Agent SDK agents in production containers, choose a session pattern, persist state, size hosts and isolate tenants that share a machine.
Hosting an Agent SDK app is not like hosting a thin API wrapper. Every query() starts a claude CLI subprocess that owns a shell, a working directory and session files on local disk. Each running agent is a stateful, long-lived process, and that shapes everything about resourcing, persistence and scaling.
This page is about running it on your own infrastructure. If you do not actually need the agent loop on your own machines, Anthropic's Managed Agents product hosts the loop for you, with tools running in an Anthropic-managed sandbox or a self-hosted sandbox. For ready-made Dockerfiles and Kubernetes manifests, Anthropic's claude-cookbooks repository on GitHub has a hosting folder with examples for Docker, Modal and Kubernetes.
The process model
client -> your app --stdio--> claude CLI subprocess --> api.anthropic.com
|
+-- shell, cwd, ~/.claude/projects/*.jsonl
- One session is one subprocess. Fifty concurrent sessions means fifty process trees and fifty transcript files.
- Every subprocess inherits your app's working directory unless you pass
cwd. If sessions must not see each other's files, give each a distinctcwd. - The subprocess never listens on the network. Your app takes inbound requests and calls the SDK.
from claude_agent_sdk import query, ClaudeAgentOptions
async for msg in query(prompt="Extract line items from invoice.pdf",
options=ClaudeAgentOptions(cwd=f"/work/jobs/{job_id}")):
...
TypeScript examples on this page use top-level await, so save them as .mts or set "type": "module".
State on local disk
None of this survives a container restart, scale-down or reschedule:
| State | Default location |
|---|---|
| Session transcripts | ~/.claude/projects/, or projects/ under CLAUDE_CONFIG_DIR |
| CLAUDE.md memory | ~/.claude/CLAUDE.md (user) and the working directory (project) |
| Files the agent made | The working directory |
Transcripts can be mirrored to durable storage with a SessionStore. Memory files and working-directory output need their own plan: a mounted volume or an object-store sync. Sessions explains resume and fork.
Pick a session pattern
| Pattern | Container lifetime | Suits |
|---|---|---|
| Ephemeral | One task, then destroyed | Bug fixes, invoice extraction, translation, media conversion |
| Long-running | Persistent, many sessions per container | Email triage agents, chat bots on Slack, per-user live sites |
| Hybrid | Spun up on demand, rehydrated from a store | Intermittent project assistants, research that pauses for hours, support agents that reload ticket history |
| Multi-agent | Several SDK processes cooperating in one container | Simulations where agents share an environment |
Ephemeral
The container reads the task from the environment, runs it and exits:
import { query } from "@anthropic-ai/claude-agent-sdk";
try {
for await (const msg of query({ prompt: process.env.JOB_PROMPT!, options: { maxTurns: 25 } })) {
if (msg.type === "result") console.log(JSON.stringify({ outcome: msg.subtype, cost: msg.total_cost_usd }));
}
} catch {
process.exitCode = 1; // e.g. error_max_turns: query() throws after yielding the result
}
If the turn cap is hit, the result subtype is error_max_turns and query() throws after yielding it, so catch it if the container must exit cleanly. See the agent loop for every subtype.
Long-running
The container exposes HTTP or WebSocket and maps each active session to a long-lived query:
- TypeScript: add turns with
streamInput()on the query object. Pre-warm subprocesses before traffic arrives withstartup(), or withprewarm()if you do not know a session'scwduntil its first request. - Python: hold a session open with
ClaudeSDKClient.
Size the container to hold your peak number of concurrent sessions in memory.
Hybrid
Containers spin down when idle and rehydrate from a SessionStore when the user returns. The store is mandatory here: shutting down without one throws away the transcript. Tune your provider's idle timeout to how often users come back.
async for msg in query(prompt=user_message,
options=ClaudeAgentOptions(resume=session_id_from_db, session_store=store)):
...
In TypeScript the options are resume and sessionStore.
Multi-agent
Give each agent its own working directory and isolate settings loading so one agent's CLAUDE.md does not leak into another. The tenant isolation options below apply equally here.
Provisioning
Choosing a sandbox
Run the SDK inside a container for process isolation, resource limits, network control and a throwaway filesystem. When comparing providers, consider:
- Managed or self-run. Sandbox-as-a-service versus software you operate.
- Cold start. Ephemeral patterns need sub-second starts; long-running ones can wait.
- Durable storage. Hybrid sessions need it somewhere.
- Billing. Per-second suits bursty ephemeral work; hourly suits always-on sessions.
- Networking. Custom egress rules, outbound proxies, private VPC peering.
Self-hosted options such as Docker, gVisor and Firecracker are compared in secure deployment.
Runtime
- Python 3.10+ for the Python SDK, or Node.js 18+ for TypeScript.
- Both SDKs bundle a native Claude Code binary for most installs; the CLI needs no separate Node.js. The quickstart lists the cases that need a separate native install.
- The bundled binary is pinned to the SDK version, so upgrading the SDK upgrades the CLI. The SDKs follow semver: take patches freely, read the changelog (on the
claude-agent-sdk-typescriptandclaude-agent-sdk-pythonGitHub repos) before a minor.
Resources
Start from 1 GiB RAM, 5 GiB disk and 1 CPU per agent. That is an idle baseline: memory grows with session length and tool use, so measure.
Network
- Outbound HTTPS to
api.anthropic.com, or your provider's regional endpoint on Bedrock or Google Cloud's Agent Platform. - Outbound access to any MCP servers or external tools.
- In production, route egress through a proxy that enforces a domain allowlist, injects credentials and logs requests.
- Inbound, expose one HTTP or WebSocket port for your app.
Production checklist
Persistence
Mirror transcripts with a SessionStore adapter. The session storage page has reference adapters for an object store, a key-value store and a database, plus a conformance suite for your own. Three behaviours to know:
- Transcripts only. CLAUDE.md and working files are not mirrored.
- A mirror, not a replacement. The subprocess writes locally first and the SDK forwards each batch. A fresh session's local transcript outlives the run; a run resumed from the store deletes its local copy at the end, leaving the store as the only durable copy.
mirror_error. If a batch cannot be delivered, it is dropped, a{ type: "system", subtype: "mirror_error" }message is emitted, and the query continues. Alert on these if durability matters.
Observability
Set OpenTelemetry variables at the container or orchestrator level and every query() exports:
CLAUDE_CODE_ENABLE_TELEMETRY=1
CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 # traces only
OTEL_TRACES_EXPORTER=otlp
OTEL_METRICS_EXPORTER=otlp
OTEL_LOGS_EXPORTER=otlp
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector.observability:4318
Prompt text and tool inputs are excluded by default. See observability.
Credentials
- Model API. The subprocess reads
ANTHROPIC_API_KEY. Supply it from your secret manager, or setANTHROPIC_BASE_URLto a proxy that injects the key outside the container (see secure deployment). - Inbound auth. Authenticate at a gateway in front of the agent. The agent should receive pre-authenticated requests, not validate user tokens.
- Tool credentials. Keep them out of the agent's environment entirely; a proxy adds them after the request leaves the container.
Scaling
Concurrency per host is bounded by memory:
agents per host = (host RAM - overhead) / per-session peak RAM
Measure the peak by running a realistic session to your target length under realistic tool load and recording peak RSS. The 1 GiB starting point is a floor.
For long-running containers behind a load balancer, pin each session to one container with consistent hashing on sessionId, so it keeps hitting the same warm subprocess until evicted or restarted.
Cost
Token spend usually outweighs container cost by an order of magnitude or more. A minimal container is roughly $0.05 an hour; a single long session can spend dollars. Track it with cost tracking.
Tenant isolation
By default the SDK reads settings and CLAUDE.md files from disk, which in a shared container can leak one tenant's context into another's session. For each tenant:
| Step | How |
|---|---|
| Skip filesystem settings | settingSources: [] / setting_sources=[] |
| Disable auto memory | CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 in env. Auto memory under ~/.claude/projects/<project>/memory/ loads regardless of setting sources |
| Separate global config | Point CLAUDE_CONFIG_DIR at a per-tenant directory |
| Short transcript paths | Optionally set CLAUDE_CODE_PROJECT_DIR_NAME when each config directory serves one working directory and you use no SessionStore (TypeScript SDK v0.3.234+, Python SDK v0.2.140+) |
| Separate files | A per-tenant cwd on every call |
| Separate egress | Per-tenant outbound IPs, credentials or allowlists at the proxy |
for await (const msg of query({
prompt,
options: {
cwd: `/tenants/${tenantId}/work`,
settingSources: [],
env: {
...process.env,
CLAUDE_CONFIG_DIR: `/tenants/${tenantId}/config`,
CLAUDE_CODE_DISABLE_AUTO_MEMORY: "1"
}
}
})) { /* ... */ }
Claude Code features in the SDK lists the other inputs that load whatever your setting sources say.
Known limitations
| Limitation | Workaround |
|---|---|
| Sessions never time out on their own | Bound them with maxTurns / max_turns |
| Memory grows over long sessions | Cap session length or recycle subprocesses |
| Wide parallel subagent fan-outs can hit rate limits | Dispatch in smaller batches |
| No wall-clock limit per subagent | Set maxTurns in each AgentDefinition. CLAUDE_ASYNC_AGENT_STALL_TIMEOUT_MS is a stall watchdog for subagents that stop producing output, not a total runtime limit |
When it works locally but not deployed
- CLI not found at start-up. In Python, the service manager's
PATHdiffers from your shell's. In TypeScript, the image build skipped optional dependencies, orpathToClaudeCodeExecutablepoints at a missing file. - CLI present but will not launch. Wrong architecture or libc for the container, or the execute bit was lost during the build.
- CLI exits mid-run. The error you see depends on language and on whether an error result arrived first.
All three are covered in detail on the troubleshooting page.