Skip to content

How the agent loop works

What happens between your prompt and the final result in an Agent SDK session, from turns and message types to tools, limits, context and compaction.

When you call query(), the SDK runs the same loop Claude Code uses in the terminal: Claude looks at the task, calls some tools, reads the results, and goes round again until it has nothing left to do. Knowing the shape of that loop makes it far easier to read the message stream, set sensible limits and work out why an agent did something odd.

Both SDKs bundle a native Claude Code binary, so the loop runs locally in a child process your code controls. For the conceptual picture outside the SDK, see how Claude Code works.

The cycle

  1. Start. Claude receives your prompt along with the system prompt, tool definitions and any history. The SDK emits a system message with subtype init that carries session metadata.
  2. Think and act. Claude decides what to do next. It may write text, ask for one or more tool calls, or both. You get an assistant message per content block.
  3. Run tools. The SDK executes the requested tools and sends the results back to Claude. Hooks can inspect, change or block calls at this point. Results come back to you as user messages.
  4. Repeat. Steps 2 and 3 form one turn. They repeat until Claude replies with no tool calls.
  5. Finish. You get a final assistant message with the answer, then a result message with the final text, usage, cost and session ID.

A small question ("which files are in this folder?") might take a turn or two. A real task ("migrate the auth module to the new client and fix the tests") might run dozens of tool calls across many turns, adjusting as results come in.

A worked trace

Here is how a run for "the invoice export test is failing, sort it out" might play out:

TurnClaude doesYou receive
(start)system / init
1Runs Bash with pytest tests/test_export.pyAssistant message (tool call), then user message (two failures)
2Calls Read on export.py and the test fileAn assistant message per call, then the file contents
3Calls Edit on export.py, then Bash to re-run the testsAssistant messages for each call, results showing a pass
4Replies in text only: "The date formatter used the US locale; fixed and tests pass."Final assistant message, then the result message

Three tool-using turns, then one text-only reply. Your code sees every step but does not need to do anything for the loop to continue: tool results feed back to Claude automatically.

Message types

The stream contains five core kinds of message.

MessageTypeScript typeWhat it tells you
System"system"Lifecycle events. Subtype init carries session metadata; compact_boundary follows a compaction; informational is a plain status banner; worker_shutting_down means the host is exiting or Remote Control disconnected.
Assistant"assistant"One content block from Claude: text, a tool call, or thinking. Blocks from the same response share a message ID.
User"user"Tool results going back to Claude, and any input you stream in mid-run.
Stream event"stream_event"Raw API streaming events. Only sent when partial messages are switched on.
Result"result"The end of the loop: final text, usage, cost, session ID and a subtype saying how it ended.

A few details worth knowing:

  • If a SessionStart or Setup hook runs at startup, its hook lifecycle messages arrive before init.
  • In TypeScript, system events other than init are separate members of the SDKMessage union rather than subtypes of SDKSystemMessage.
  • A small number of system events (such as prompt_suggestion) can arrive after the result. Iterate to the end of the stream rather than breaking as soon as you see the result.
  • Both SDKs also emit observability events (rate-limit status, task notifications). You do not need them to drive the loop. The TypeScript and Python references list them all.

Reading messages in each language

In Python you test with isinstance() against classes from claude_agent_sdk. In TypeScript you compare the type string. Watch out for one TypeScript quirk: assistant and user messages wrap the raw API message, so content lives at message.message.content, not message.content.

What to handle depends on the job:

  • Just the answer: watch for the result message.
  • Progress: watch assistant messages to see text and tool calls as they happen.
  • Live typing effect: turn on partial messages (includePartialMessages / include_partial_messages) and handle stream events. See streaming output.
from claude_agent_sdk import query, AssistantMessage, ResultMessage, TextBlock, ToolUseBlock

async def watch(task: str):
    try:
        async for msg in query(prompt=task):
            if isinstance(msg, AssistantMessage):
                for block in msg.content:
                    if isinstance(block, ToolUseBlock):
                        print(f"[tool] {block.name}")
                    elif isinstance(block, TextBlock):
                        print(block.text)
            elif isinstance(msg, ResultMessage):
                print("OK" if msg.subtype == "success" else f"Ended early: {msg.subtype}")
    except Exception as exc:
        print(f"Run failed: {exc}")

Tools

What is built in

GroupToolsPurpose
FilesRead, Write, EditRead, create and change files
SearchGlob, GrepFind files by pattern and search contents
ShellBashRun commands, scripts and git
WebWebSearch, WebFetchSearch the web and fetch pages
DiscoveryToolSearchLoad tool definitions on demand instead of up front
OrchestrationAgent, Skill, AskUserQuestion, TaskCreate, TaskUpdateSpawn subagents, run skills, ask the user, track tasks

On models that do not get the task-tracking tools by default, TaskCreate and TaskUpdate appear only if you opt in; see todo tracking. Beyond the built-ins you can add MCP servers, your own custom tools and project skills.

Who decides whether a tool runs

Claude chooses which tools to call. You decide whether each call is allowed. Three options interact:

  • allowedTools / allowed_tools pre-approves the listed tools. Unlisted tools are still available; calls to them that need approval go on to the permission mode and your canUseTool callback.
  • disallowedTools / disallowed_tools blocks the listed tools whatever else is configured.
  • permissionMode / permission_mode sets the overall level of oversight (see permission modes below).

Rules can be scoped, for example "Bash(npm run *)" to allow only npm scripts. SDK permissions has the full syntax and the order in which rules are checked. When a call is refused, Claude receives the refusal as the tool result and usually tries another approach or explains that it could not continue.

Parallel calls

If Claude asks for several tools in one turn, read-only ones (Read, Glob, Grep, and MCP tools marked read-only) can run at the same time. Tools that change state (Edit, Write, Bash) run one after another to avoid stepping on each other. Custom tools run sequentially unless you set readOnlyHint in their annotations.

Controlling the loop

All of these are fields on the options object. Configure your agent shows how to compose them.

Turns and budget

OptionLimitsDefault
maxTurns / max_turnsTool-using round tripsUnlimited
maxBudgetUsd / max_budget_usdTotal spend in US dollarsUnlimited

maxTurns counts only turns that call tools, so in the trace above maxTurns: 2 would have stopped the run before the edit. Hitting either cap ends the run with result subtype error_max_turns or error_max_budget_usd.

The budget includes subagent spend. Once the cap is reached, spawning another subagent fails with Budget limit reached and any background subagents are stopped. That enforcement needs Claude Code v2.1.217 or later.

With streaming input, a message still waiting when a turn hits the turn cap stays queued and starts a fresh turn with a fresh turn count. The budget keeps accumulating across messages; once exhausted, every later message in the conversation ends with error_max_budget_usd until a /clear resets it.

Unbounded runs are fine for tight, well-specified tasks. For anything open-ended or in production, I always set a budget.

Effort

effort controls how much reasoning Claude puts into each response. Lower settings use fewer tokens and finish faster. Not every model supports it.

LevelCharacterFits
"low"Quick, minimal reasoningLookups, listing files, single greps
"medium"BalancedRoutine edits
"high"ThoroughDebugging, refactors
"xhigh"Deeper stillCoding and agentic work on models that support it
"max"DeepestHard multi-step problems

If you leave it unset, Claude Code resolves a level itself using the order described on the model configuration page. You can set it for the whole session or per subagent through the effort field on an agent definition.

Effort is separate from extended thinking. Thinking produces thinking blocks, and the display field on the thinking config controls whether you receive their text. You can combine any effort level with thinking on or off.

Permission modes

ModeWhat happensWhen I use it
"default"Calls needing approval that no allow rule covers go to your canUseTool callback. No callback means they are denied.Interactive apps with an approval UI
"acceptEdits"File edits and common filesystem commands (mkdir, touch, mv, cp and so on) are approved; other Bash follows normal rules.Prototyping, or work in a scratch directory
"plan"Claude explores and plans but does not edit source files. Edits are never auto-approved and go through canUseTool.Reviews, or when changes need sign-off first
"dontAsk"Never prompts. Pre-approved tools and calls that need no approval (such as reads inside the working directories) run; anything that would prompt is denied. AskUserQuestion, connector tools your organisation set to ask, and MCP tools marked requiresUserInteraction are denied even if allowed.Headless agents with a fixed, explicit tool surface
"auto"A classifier model reviews actions such as shell commands and network requests, allowing or blocking each.Autonomous agents that still want guardrails
"bypassPermissions"Every allowed tool runs without asking, except tools hit by an explicit ask rule, connector tools your organisation set to ask, and tools needing user interaction. Cross-session messaging safeguards still apply. TypeScript also requires allowDangerouslySkipPermissions: true. Refused when running as root on Unix.Throwaway containers and CI only

See permission modes for auto mode availability and SDK permissions for precedence.

Model

Set model to choose what runs the session. Configure your agent covers aliases and fallbacks.

Context

Within a session, context only grows. The system prompt, tool definitions, every message, every tool input and every tool output all stay in the window. Content that is identical from request to request (system prompt, tool definitions, CLAUDE.md) is prompt-cached automatically, so you pay full price for it once. Modifying system prompts explains how custom prompts affect cache reuse.

What fills it

SourceLoadedCost
System promptEvery requestSmall and fixed
CLAUDE.mdSession start, through settingSourcesWhole file every request, but cached after the first
Tool definitionsEvery requestBuilt-in schemas always; MCP schemas are deferred by tool search by default, with upfront loading as a fallback on unsupported models and some platforms
Conversation historyBuilds up turn by turnGrows with every prompt, reply and tool result
Skill descriptionsSession startShort summaries; full skill text loads only when used

Big tool outputs are the usual culprit for runaway context. One verbose test run or a large file read can cost thousands of tokens, and it stays there for the rest of the session.

Automatic compaction

As the window nears its limit, the SDK summarises older history to make room, keeping recent exchanges and key decisions. You will see a system message with subtype compact_boundary (a SystemMessage in Python, an SDKCompactBoundaryMessage in TypeScript).

Because compaction replaces early messages with a summary, instructions you gave only in the first prompt can get lost. Put rules that must survive into CLAUDE.md, which is re-sent with every request.

You can steer compaction three ways:

  • Tell it what to keep in CLAUDE.md. The compactor reads CLAUDE.md like anything else, and any clearly labelled section works:

    ## When summarising this session
    Keep: the ticket number and acceptance criteria, every file path touched,
    the last test output, and any decision the user made about scope.
    
  • PreCompact hook. Runs before compaction with a trigger of manual or auto. Handy for archiving the full transcript first.

  • Compact on demand. Send /compact as a prompt. Commands sent this way are ordinary inputs.

Keeping context lean

  • Delegate to subagents. A subagent starts with no message history (it still gets its own system prompt and project context such as CLAUDE.md), and only its final answer comes back to the parent. The parent grows by a summary, not a whole transcript.
  • Trim tools. Each definition takes space. Give subagents only the tools they need with the tools field on their definition.
  • Watch MCP. With tool search on, MCP schemas load on demand. Where it is off or has fallen back, every server's full tool list ships with every request.
  • Lower the effort for simple jobs.

The features overview breaks down context costs feature by feature.

Sessions

Every run creates or continues a session. Grab the ID from the result's session_id (both SDKs); TypeScript also has it directly on the init message, while Python nests it in SystemMessage.data. Resuming restores everything: files read, analysis done, actions taken. Forking lets you try a different direction without touching the original. In Python, ClaudeSDKClient threads the session through multiple calls for you.

Sessions covers resume, continue and fork. To resume on a different machine or in a stateless container, give the SDK a session store so transcripts are mirrored to your own backend.

Reading the result

subtypeMeaningresult text present
successFinished normallyYes
error_max_turnsTurn cap reachedNo
error_max_budget_usdBudget cap reachedNo
error_during_executionSomething interrupted the loop, such as a cancelled requestNo
error_max_structured_output_retriesNo output passed schema validation within the retry limit, or a model fallback retracted a completed output and no retry succeededNo

Always check subtype before reading result. Every variant carries total_cost_usd, usage, num_turns and session_id, so you can record cost and resume even after failures. Some caveats:

  • After a session crash you get error_during_execution with cost fields possibly zeroed and stop_reason set to null, and the process exits.
  • In Python, total_cost_usd, usage and model_usage are optional; check for None.
  • usage covers the main loop only. For the whole tree including subagents, use modelUsage / model_usage. See cost tracking.

stop_reason says why the model stopped on its final turn: commonly end_turn, max_tokens or refusal. Check for "refusal" to detect declined requests.

Note: On an error result, a single-message query() yields the result and then raises (for example with Reached maximum number of turns), and the Claude Code process exits non-zero. This is deliberate, so wrap the loop in try if you need to carry on. A streaming input session stays alive after an error result, unless the session itself crashed.

Hooks in the loop

Hooks let your code run at fixed points. The ones I use most:

HookFiresTypical use
PreToolUseBefore a tool runsValidate input, block risky commands
PostToolUseAfter a tool returnsAudit logs, side effects
UserPromptSubmitWhen a prompt is sentAdd context
StopWhen the agent finishesCheck the result, persist state
SubagentStart / SubagentStopSubagent begins or endsTrack parallel work
PreCompactBefore compactionArchive the transcript

Hooks run in your process, not in Claude's context, so they cost no tokens. A PreToolUse hook that rejects a call stops it running and Claude sees the rejection. TypeScript supports some events Python does not yet have.

Putting it together

A TypeScript agent that fixes a failing test suite with sensible guard rails:

import { query } from "@anthropic-ai/claude-agent-sdk";

let sessionId: string | undefined;

try {
  for await (const msg of query({
    prompt: "The billing tests are red. Find out why and fix the code, not the tests.",
    options: {
      allowedTools: ["Read", "Edit", "Bash", "Glob", "Grep"],
      settingSources: ["project"],
      maxTurns: 25,
      maxBudgetUsd: 2,
      effort: "high",
    },
  })) {
    if (msg.type === "system" && msg.subtype === "init") sessionId = msg.session_id;

    if (msg.type === "result") {
      switch (msg.subtype) {
        case "success":
          console.log(msg.result);
          break;
        case "error_max_turns":
        case "error_max_budget_usd":
          console.log(`Stopped at a limit. Resume ${sessionId} with a higher cap.`);
          break;
        default:
          console.log(`Ended: ${msg.subtype}`);
      }
      console.log(`Spent $${msg.total_cost_usd.toFixed(3)} over ${msg.num_turns} turns`);
    }
  }
} catch (err) {
  console.error("Agent run failed:", err);
}