Skip to content

Streaming output

Turn on partial messages in the Agent SDK to show text and tool calls token by token, and build a responsive progress display.

By default the SDK hands you each piece of Claude's response once it is finished: a complete text block, a complete tool call. That is fine for logs and batch jobs, but in a chat interface it means the user stares at nothing for several seconds. Partial message streaming fixes that by passing through the raw API events as they arrive.

This page covers output. For how you send messages in, see streaming input versus single messages. The CLI can stream too; see headless mode.

Switching it on

Set includePartialMessages: true (TypeScript) or include_partial_messages=True (Python). You then receive stream event messages alongside the usual assistant and result messages.

To print text as it is generated you look for three things, nested:

  1. a stream event message,
  2. whose event.type is content_block_delta,
  3. whose delta.type is text_delta.

TypeScript:

import { query } from "@anthropic-ai/claude-agent-sdk";

for await (const msg of query({
  prompt: "Explain what the scripts in package.json do",
  options: { includePartialMessages: true, allowedTools: ["Read"] },
})) {
  if (msg.type !== "stream_event") continue;
  const ev = msg.event;
  if (ev.type === "content_block_delta" && ev.delta.type === "text_delta") {
    process.stdout.write(ev.delta.text);
  }
}

Python:

from claude_agent_sdk import query, ClaudeAgentOptions
from claude_agent_sdk.types import StreamEvent

async def explain_scripts():
    opts = ClaudeAgentOptions(include_partial_messages=True, allowed_tools=["Read"])
    async for msg in query(prompt="Explain what the scripts in package.json do", options=opts):
        if isinstance(msg, StreamEvent):
            ev = msg.event
            if ev.get("type") == "content_block_delta" and ev.get("delta", {}).get("type") == "text_delta":
                print(ev["delta"]["text"], end="", flush=True)

Note that in Python the event is a plain dict, so use .get().

What a stream event contains

PythonTypeScript
TypeStreamEvent (import from claude_agent_sdk.types)SDKPartialAssistantMessage, with type: "stream_event"
eventThe raw Claude API streaming eventSame
parent_tool_use_idAlways NoneAlways null
user_message_uuidNot exposedPresent on some events (see below)

The events are raw API events, not running totals. If you want the full text so far, accumulate the deltas yourself.

Stream events come from the main session only. Subagents do not forward token-level deltas, which is why parent_tool_use_id is always empty. To attribute output to a subagent, use the complete assistant messages, which do carry parent_tool_use_id; see subagents.

In TypeScript, Claude Code sets user_message_uuid on the first non-ping stream event of a turn, and again whenever the user message that turn is answering changes. The TypeScript reference details the conditions.

Event types you will see

event.typeSignals
message_startA new response from Claude begins
content_block_startA new block (text or tool use) begins; check content_block.type
content_block_deltaA chunk of a block: text_delta for text, input_json_delta for tool input
content_block_stopThe current block is finished
message_deltaResponse-level updates such as stop reason and usage
message_stopThe response is finished

Order of messages

Claude Code still emits a complete assistant message for each non-empty block, with partial streaming on or off. A response containing some text and then a tool call produces two assistant messages that share a message ID (message.message.id in TypeScript, message.message_id in Python), each holding one block.

With partial messages on, each complete assistant message arrives just before that block's content_block_stop. A single turn looks like this:

stream_event   message_start
stream_event   content_block_start      (text)
stream_event   content_block_delta      (text_delta, repeated)
assistant      [text block, complete]
stream_event   content_block_stop
stream_event   content_block_start      (tool_use)
stream_event   content_block_delta      (input_json_delta, repeated)
assistant      [tool_use block, complete]
stream_event   content_block_stop
stream_event   message_delta
stream_event   message_stop
               ...tool runs, next turn streams the same way...
result

With partial messages off you get everything except the stream events: the system init message, complete assistant messages, the result, and a compaction marker when history is compacted (SDKCompactBoundaryMessage in TypeScript, a SystemMessage with subtype compact_boundary in Python).

Streaming tool calls

Tool calls stream too. A tool call starts with a content_block_start whose block is tool_use (carrying the tool name), continues with input_json_delta chunks containing fragments of the JSON input in partial_json, and ends with content_block_stop.

The fragments are not valid JSON on their own. Concatenate them and parse once the block stops:

current = None
buffer = ""

async for msg in query(prompt="Find where VAT is calculated", options=opts):
    if not isinstance(msg, StreamEvent):
        continue
    ev = msg.event
    kind = ev.get("type")
    if kind == "content_block_start" and ev["content_block"].get("type") == "tool_use":
        current, buffer = ev["content_block"]["name"], ""
    elif kind == "content_block_delta" and ev["delta"].get("type") == "input_json_delta":
        buffer += ev["delta"].get("partial_json", "")
    elif kind == "content_block_stop" and current:
        print(f"{current} -> {buffer}")
        current = None

A progress display

Combining text and tool streaming gives the sort of interface people expect from an agent: prose appears as it is written, and while a tool runs you show a status line instead. This TypeScript version tracks which tool is active and how long it took:

import { query } from "@anthropic-ai/claude-agent-sdk";

let activeTool: string | null = null;
let startedAt = 0;

for await (const msg of query({
  prompt: "Audit src/ for hard-coded API URLs and list them",
  options: { includePartialMessages: true, allowedTools: ["Read", "Grep", "Glob"] },
})) {
  if (msg.type === "stream_event") {
    const ev = msg.event;
    if (ev.type === "content_block_start" && ev.content_block.type === "tool_use") {
      activeTool = ev.content_block.name;
      startedAt = Date.now();
      process.stdout.write(`\n  ... ${activeTool}`);
    } else if (ev.type === "content_block_delta" && ev.delta.type === "text_delta" && !activeTool) {
      process.stdout.write(ev.delta.text);
    } else if (ev.type === "content_block_stop" && activeTool) {
      process.stdout.write(` (${Date.now() - startedAt} ms)\n`);
      activeTool = null;
    }
  } else if (msg.type === "result") {
    console.log(`\n\nFinished in ${msg.num_turns} turns.`);
  }
}

Bear in mind the timer measures how long Claude took to write the tool input, not how long the tool ran: the block stops before the tool executes. For real tool timings, use PreToolUse and PostToolUse hooks.

Limitations

  • Structured output. With partial messages on, structured output streams as unvalidated input_json_delta chunks of a tool call. Only the validated object reaches structured_output on the final result, so do not trust the streamed fragments. See structured outputs.
  • Subagents. As above, their token-level output is not streamed.