How the agent loop works
What happens between your prompt and the final result in an Agent SDK session, from turns and message types to tools, limits, context and compaction.
When you call query(), the SDK runs the same loop Claude Code uses in the terminal: Claude looks at the task, calls some tools, reads the results, and goes round again until it has nothing left to do. Knowing the shape of that loop makes it far easier to read the message stream, set sensible limits and work out why an agent did something odd.
Both SDKs bundle a native Claude Code binary, so the loop runs locally in a child process your code controls. For the conceptual picture outside the SDK, see how Claude Code works.
The cycle
- Start. Claude receives your prompt along with the system prompt, tool definitions and any history. The SDK emits a system message with subtype
initthat carries session metadata. - Think and act. Claude decides what to do next. It may write text, ask for one or more tool calls, or both. You get an assistant message per content block.
- Run tools. The SDK executes the requested tools and sends the results back to Claude. Hooks can inspect, change or block calls at this point. Results come back to you as user messages.
- Repeat. Steps 2 and 3 form one turn. They repeat until Claude replies with no tool calls.
- Finish. You get a final assistant message with the answer, then a result message with the final text, usage, cost and session ID.
A small question ("which files are in this folder?") might take a turn or two. A real task ("migrate the auth module to the new client and fix the tests") might run dozens of tool calls across many turns, adjusting as results come in.
A worked trace
Here is how a run for "the invoice export test is failing, sort it out" might play out:
| Turn | Claude does | You receive |
|---|---|---|
| (start) | system / init | |
| 1 | Runs Bash with pytest tests/test_export.py | Assistant message (tool call), then user message (two failures) |
| 2 | Calls Read on export.py and the test file | An assistant message per call, then the file contents |
| 3 | Calls Edit on export.py, then Bash to re-run the tests | Assistant messages for each call, results showing a pass |
| 4 | Replies in text only: "The date formatter used the US locale; fixed and tests pass." | Final assistant message, then the result message |
Three tool-using turns, then one text-only reply. Your code sees every step but does not need to do anything for the loop to continue: tool results feed back to Claude automatically.
Message types
The stream contains five core kinds of message.
| Message | TypeScript type | What it tells you |
|---|---|---|
| System | "system" | Lifecycle events. Subtype init carries session metadata; compact_boundary follows a compaction; informational is a plain status banner; worker_shutting_down means the host is exiting or Remote Control disconnected. |
| Assistant | "assistant" | One content block from Claude: text, a tool call, or thinking. Blocks from the same response share a message ID. |
| User | "user" | Tool results going back to Claude, and any input you stream in mid-run. |
| Stream event | "stream_event" | Raw API streaming events. Only sent when partial messages are switched on. |
| Result | "result" | The end of the loop: final text, usage, cost, session ID and a subtype saying how it ended. |
A few details worth knowing:
- If a
SessionStartorSetuphook runs at startup, its hook lifecycle messages arrive beforeinit. - In TypeScript, system events other than
initare separate members of theSDKMessageunion rather than subtypes ofSDKSystemMessage. - A small number of system events (such as
prompt_suggestion) can arrive after the result. Iterate to the end of the stream rather than breaking as soon as you see the result. - Both SDKs also emit observability events (rate-limit status, task notifications). You do not need them to drive the loop. The TypeScript and Python references list them all.
Reading messages in each language
In Python you test with isinstance() against classes from claude_agent_sdk. In TypeScript you compare the type string. Watch out for one TypeScript quirk: assistant and user messages wrap the raw API message, so content lives at message.message.content, not message.content.
What to handle depends on the job:
- Just the answer: watch for the result message.
- Progress: watch assistant messages to see text and tool calls as they happen.
- Live typing effect: turn on partial messages (
includePartialMessages/include_partial_messages) and handle stream events. See streaming output.
from claude_agent_sdk import query, AssistantMessage, ResultMessage, TextBlock, ToolUseBlock
async def watch(task: str):
try:
async for msg in query(prompt=task):
if isinstance(msg, AssistantMessage):
for block in msg.content:
if isinstance(block, ToolUseBlock):
print(f"[tool] {block.name}")
elif isinstance(block, TextBlock):
print(block.text)
elif isinstance(msg, ResultMessage):
print("OK" if msg.subtype == "success" else f"Ended early: {msg.subtype}")
except Exception as exc:
print(f"Run failed: {exc}")
Tools
What is built in
| Group | Tools | Purpose |
|---|---|---|
| Files | Read, Write, Edit | Read, create and change files |
| Search | Glob, Grep | Find files by pattern and search contents |
| Shell | Bash | Run commands, scripts and git |
| Web | WebSearch, WebFetch | Search the web and fetch pages |
| Discovery | ToolSearch | Load tool definitions on demand instead of up front |
| Orchestration | Agent, Skill, AskUserQuestion, TaskCreate, TaskUpdate | Spawn subagents, run skills, ask the user, track tasks |
On models that do not get the task-tracking tools by default, TaskCreate and TaskUpdate appear only if you opt in; see todo tracking. Beyond the built-ins you can add MCP servers, your own custom tools and project skills.
Who decides whether a tool runs
Claude chooses which tools to call. You decide whether each call is allowed. Three options interact:
allowedTools/allowed_toolspre-approves the listed tools. Unlisted tools are still available; calls to them that need approval go on to the permission mode and yourcanUseToolcallback.disallowedTools/disallowed_toolsblocks the listed tools whatever else is configured.permissionMode/permission_modesets the overall level of oversight (see permission modes below).
Rules can be scoped, for example "Bash(npm run *)" to allow only npm scripts. SDK permissions has the full syntax and the order in which rules are checked. When a call is refused, Claude receives the refusal as the tool result and usually tries another approach or explains that it could not continue.
Parallel calls
If Claude asks for several tools in one turn, read-only ones (Read, Glob, Grep, and MCP tools marked read-only) can run at the same time. Tools that change state (Edit, Write, Bash) run one after another to avoid stepping on each other. Custom tools run sequentially unless you set readOnlyHint in their annotations.
Controlling the loop
All of these are fields on the options object. Configure your agent shows how to compose them.
Turns and budget
| Option | Limits | Default |
|---|---|---|
maxTurns / max_turns | Tool-using round trips | Unlimited |
maxBudgetUsd / max_budget_usd | Total spend in US dollars | Unlimited |
maxTurns counts only turns that call tools, so in the trace above maxTurns: 2 would have stopped the run before the edit. Hitting either cap ends the run with result subtype error_max_turns or error_max_budget_usd.
The budget includes subagent spend. Once the cap is reached, spawning another subagent fails with Budget limit reached and any background subagents are stopped. That enforcement needs Claude Code v2.1.217 or later.
With streaming input, a message still waiting when a turn hits the turn cap stays queued and starts a fresh turn with a fresh turn count. The budget keeps accumulating across messages; once exhausted, every later message in the conversation ends with error_max_budget_usd until a /clear resets it.
Unbounded runs are fine for tight, well-specified tasks. For anything open-ended or in production, I always set a budget.
Effort
effort controls how much reasoning Claude puts into each response. Lower settings use fewer tokens and finish faster. Not every model supports it.
| Level | Character | Fits |
|---|---|---|
"low" | Quick, minimal reasoning | Lookups, listing files, single greps |
"medium" | Balanced | Routine edits |
"high" | Thorough | Debugging, refactors |
"xhigh" | Deeper still | Coding and agentic work on models that support it |
"max" | Deepest | Hard multi-step problems |
If you leave it unset, Claude Code resolves a level itself using the order described on the model configuration page. You can set it for the whole session or per subagent through the effort field on an agent definition.
Effort is separate from extended thinking. Thinking produces thinking blocks, and the display field on the thinking config controls whether you receive their text. You can combine any effort level with thinking on or off.
Permission modes
| Mode | What happens | When I use it |
|---|---|---|
"default" | Calls needing approval that no allow rule covers go to your canUseTool callback. No callback means they are denied. | Interactive apps with an approval UI |
"acceptEdits" | File edits and common filesystem commands (mkdir, touch, mv, cp and so on) are approved; other Bash follows normal rules. | Prototyping, or work in a scratch directory |
"plan" | Claude explores and plans but does not edit source files. Edits are never auto-approved and go through canUseTool. | Reviews, or when changes need sign-off first |
"dontAsk" | Never prompts. Pre-approved tools and calls that need no approval (such as reads inside the working directories) run; anything that would prompt is denied. AskUserQuestion, connector tools your organisation set to ask, and MCP tools marked requiresUserInteraction are denied even if allowed. | Headless agents with a fixed, explicit tool surface |
"auto" | A classifier model reviews actions such as shell commands and network requests, allowing or blocking each. | Autonomous agents that still want guardrails |
"bypassPermissions" | Every allowed tool runs without asking, except tools hit by an explicit ask rule, connector tools your organisation set to ask, and tools needing user interaction. Cross-session messaging safeguards still apply. TypeScript also requires allowDangerouslySkipPermissions: true. Refused when running as root on Unix. | Throwaway containers and CI only |
See permission modes for auto mode availability and SDK permissions for precedence.
Model
Set model to choose what runs the session. Configure your agent covers aliases and fallbacks.
Context
Within a session, context only grows. The system prompt, tool definitions, every message, every tool input and every tool output all stay in the window. Content that is identical from request to request (system prompt, tool definitions, CLAUDE.md) is prompt-cached automatically, so you pay full price for it once. Modifying system prompts explains how custom prompts affect cache reuse.
What fills it
| Source | Loaded | Cost |
|---|---|---|
| System prompt | Every request | Small and fixed |
| CLAUDE.md | Session start, through settingSources | Whole file every request, but cached after the first |
| Tool definitions | Every request | Built-in schemas always; MCP schemas are deferred by tool search by default, with upfront loading as a fallback on unsupported models and some platforms |
| Conversation history | Builds up turn by turn | Grows with every prompt, reply and tool result |
| Skill descriptions | Session start | Short summaries; full skill text loads only when used |
Big tool outputs are the usual culprit for runaway context. One verbose test run or a large file read can cost thousands of tokens, and it stays there for the rest of the session.
Automatic compaction
As the window nears its limit, the SDK summarises older history to make room, keeping recent exchanges and key decisions. You will see a system message with subtype compact_boundary (a SystemMessage in Python, an SDKCompactBoundaryMessage in TypeScript).
Because compaction replaces early messages with a summary, instructions you gave only in the first prompt can get lost. Put rules that must survive into CLAUDE.md, which is re-sent with every request.
You can steer compaction three ways:
-
Tell it what to keep in CLAUDE.md. The compactor reads CLAUDE.md like anything else, and any clearly labelled section works:
## When summarising this session Keep: the ticket number and acceptance criteria, every file path touched, the last test output, and any decision the user made about scope. -
PreCompacthook. Runs before compaction with atriggerofmanualorauto. Handy for archiving the full transcript first. -
Compact on demand. Send
/compactas a prompt. Commands sent this way are ordinary inputs.
Keeping context lean
- Delegate to subagents. A subagent starts with no message history (it still gets its own system prompt and project context such as CLAUDE.md), and only its final answer comes back to the parent. The parent grows by a summary, not a whole transcript.
- Trim tools. Each definition takes space. Give subagents only the tools they need with the
toolsfield on their definition. - Watch MCP. With tool search on, MCP schemas load on demand. Where it is off or has fallen back, every server's full tool list ships with every request.
- Lower the effort for simple jobs.
The features overview breaks down context costs feature by feature.
Sessions
Every run creates or continues a session. Grab the ID from the result's session_id (both SDKs); TypeScript also has it directly on the init message, while Python nests it in SystemMessage.data. Resuming restores everything: files read, analysis done, actions taken. Forking lets you try a different direction without touching the original. In Python, ClaudeSDKClient threads the session through multiple calls for you.
Sessions covers resume, continue and fork. To resume on a different machine or in a stateless container, give the SDK a session store so transcripts are mirrored to your own backend.
Reading the result
subtype | Meaning | result text present |
|---|---|---|
success | Finished normally | Yes |
error_max_turns | Turn cap reached | No |
error_max_budget_usd | Budget cap reached | No |
error_during_execution | Something interrupted the loop, such as a cancelled request | No |
error_max_structured_output_retries | No output passed schema validation within the retry limit, or a model fallback retracted a completed output and no retry succeeded | No |
Always check subtype before reading result. Every variant carries total_cost_usd, usage, num_turns and session_id, so you can record cost and resume even after failures. Some caveats:
- After a session crash you get
error_during_executionwith cost fields possibly zeroed andstop_reasonset to null, and the process exits. - In Python,
total_cost_usd,usageandmodel_usageare optional; check forNone. usagecovers the main loop only. For the whole tree including subagents, usemodelUsage/model_usage. See cost tracking.
stop_reason says why the model stopped on its final turn: commonly end_turn, max_tokens or refusal. Check for "refusal" to detect declined requests.
Note: On an error result, a single-message
query()yields the result and then raises (for example withReached maximum number of turns), and the Claude Code process exits non-zero. This is deliberate, so wrap the loop intryif you need to carry on. A streaming input session stays alive after an error result, unless the session itself crashed.
Hooks in the loop
Hooks let your code run at fixed points. The ones I use most:
| Hook | Fires | Typical use |
|---|---|---|
PreToolUse | Before a tool runs | Validate input, block risky commands |
PostToolUse | After a tool returns | Audit logs, side effects |
UserPromptSubmit | When a prompt is sent | Add context |
Stop | When the agent finishes | Check the result, persist state |
SubagentStart / SubagentStop | Subagent begins or ends | Track parallel work |
PreCompact | Before compaction | Archive the transcript |
Hooks run in your process, not in Claude's context, so they cost no tokens. A PreToolUse hook that rejects a call stops it running and Claude sees the rejection. TypeScript supports some events Python does not yet have.
Putting it together
A TypeScript agent that fixes a failing test suite with sensible guard rails:
import { query } from "@anthropic-ai/claude-agent-sdk";
let sessionId: string | undefined;
try {
for await (const msg of query({
prompt: "The billing tests are red. Find out why and fix the code, not the tests.",
options: {
allowedTools: ["Read", "Edit", "Bash", "Glob", "Grep"],
settingSources: ["project"],
maxTurns: 25,
maxBudgetUsd: 2,
effort: "high",
},
})) {
if (msg.type === "system" && msg.subtype === "init") sessionId = msg.session_id;
if (msg.type === "result") {
switch (msg.subtype) {
case "success":
console.log(msg.result);
break;
case "error_max_turns":
case "error_max_budget_usd":
console.log(`Stopped at a limit. Resume ${sessionId} with a higher cap.`);
break;
default:
console.log(`Ended: ${msg.subtype}`);
}
console.log(`Spent $${msg.total_cost_usd.toFixed(3)} over ${msg.num_turns} turns`);
}
}
} catch (err) {
console.error("Agent run failed:", err);
}