Skip to content

Streaming input versus single messages

The two ways to send prompts to an Agent SDK agent, what each one supports, and when a long-lived streaming session is worth the extra code.

There are two ways to feed prompts into an Agent SDK agent. You can pass a single string and let the run finish, or you can pass an async generator that keeps yielding messages for as long as the session should live. The second is called streaming input, and it is the recommended default for anything interactive.

This page is about input. If you want tokens to appear on screen as Claude writes them, that is streaming output, a separate option.

Side by side

Streaming inputSingle message
What you pass as promptAn async iterable of user messagesA string
LifetimeLong-lived session you keep feedingOne run, then done
Images in messagesYesNo
Queue several messagesYes, processed in orderNo
Interrupt mid-runYesNo
Change model or permission mode mid-sessionYesNo
Natural multi-turn chatYesOnly by calling again with continue or resume
After an error resultSession stays aliveRun yields the result, then raises
Good forChat UIs, IDE plugins, assistantsCron jobs, CI steps, serverless handlers

Streaming input

In streaming mode the agent behaves like a long-running process: it accepts new user messages as you produce them, surfaces permission requests, handles interruptions and manages the session for you. You get:

  • Image attachments on any message.
  • Queued messages that run one after another, with the option to interrupt.
  • Every tool and MCP server available throughout.
  • Live feedback as responses are produced.
  • Context that carries forward between turns without any session bookkeeping.

TypeScript

Pass an async generator of SDKUserMessage objects as prompt. Here a support agent receives a written complaint, and a moment later a screenshot of the error the customer saw:

import { query, type SDKUserMessage } from "@anthropic-ai/claude-agent-sdk";
import { readFile } from "node:fs/promises";

function text(content: string): SDKUserMessage {
  return { type: "user", message: { role: "user", content }, parent_tool_use_id: null };
}

async function* conversation(): AsyncGenerator<SDKUserMessage> {
  yield text("A customer says invoices fail to send after they changed their company name. Investigate src/mailer.");

  const screenshot = await readFile("screenshots/send-error.png", "base64");
  yield {
    type: "user",
    message: {
      role: "user",
      content: [
        { type: "text", text: "Here is what they saw. Does it match what you found?" },
        { type: "image", source: { type: "base64", media_type: "image/png", data: screenshot } },
      ],
    },
    parent_tool_use_id: null,
  };
}

for await (const msg of query({
  prompt: conversation(),
  options: { allowedTools: ["Read", "Grep", "Glob"], maxTurns: 12 },
})) {
  if (msg.type === "result" && msg.subtype === "success") console.log(msg.result);
}

Each message gets its own result as it completes. In a real app the generator would await input from a websocket or a queue rather than read a file.

Python

Use ClaudeSDKClient. You can pass an async generator to client.query(), but bear in mind that receive_response() stops at the first result message. For a conversation where you want each answer, the cleaner pattern is one query() and receive_response() pair per message:

import asyncio, base64
from claude_agent_sdk import ClaudeSDKClient, ClaudeAgentOptions, AssistantMessage, TextBlock

async def print_reply(client):
    async for msg in client.receive_response():
        if isinstance(msg, AssistantMessage):
            for block in msg.content:
                if isinstance(block, TextBlock):
                    print(block.text)

async def main():
    opts = ClaudeAgentOptions(allowed_tools=["Read", "Grep", "Glob"], max_turns=12)
    async with ClaudeSDKClient(opts) as client:
        await client.query("Invoices fail to send after a company rename. Investigate src/mailer.")
        await print_reply(client)

        with open("screenshots/send-error.png", "rb") as f:
            data = base64.b64encode(f.read()).decode()

        async def with_image():
            yield {
                "type": "user",
                "message": {
                    "role": "user",
                    "content": [
                        {"type": "text", "text": "Here is the error they saw. Does it match?"},
                        {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": data}},
                    ],
                },
            }

        await client.query(with_image())
        await print_reply(client)

asyncio.run(main())

Debugging a streaming session

  • Bad image blocks fail quietly. If an image block's source is missing or not an object, the SDK does not raise. Claude receives a text note such as [Image could not be processed: image block has no source object] and carries on.
  • TypeScript generator errors are disguised. If your generator throws (a missing file is the usual cause), the stream ends with Claude Code process aborted by user, not your original error. A long line of minified SDK source may come first, so scroll to the end. When you see that message, look inside your generator.
  • Python generator errors hang. An exception in your generator is logged at debug level and the session simply stalls. If a streaming session sits silent, turn on debug logging and check the generator.

Single message input

Passing a plain string is simpler, and right when:

  • you want one answer and are done,
  • you do not need images, interruption or mid-session control, or
  • you are in a stateless environment such as a Lambda or Cloud Run job.

Warning: Single message mode cannot attach images, queue messages dynamically, interrupt a run or hold a natural multi-turn conversation.

You can still chain calls by passing continue: true (Python: continue_conversation=True) or resume with a session ID; see sessions.

When a single-message run ends on an error result such as error_max_turns, the SDK yields the final result and then raises an error containing the failure text (in Python, a ResultError). Wrap the loop if your code should keep going:

from claude_agent_sdk import query, ClaudeAgentOptions, ResultMessage

async def nightly_report():
    opts = ClaudeAgentOptions(allowed_tools=["Read", "Grep"], max_turns=6)
    try:
        async for msg in query(prompt="Summarise yesterday's errors in logs/app.log", options=opts):
            if isinstance(msg, ResultMessage) and msg.subtype == "success":
                post_to_slack(msg.result)
    except Exception as exc:
        post_to_slack(f"Report failed: {exc}")

    try:
        async for msg in query(prompt="Which of those errors are new this week?",
                               options=ClaudeAgentOptions(continue_conversation=True, max_turns=6)):
            if isinstance(msg, ResultMessage) and msg.subtype == "success":
                post_to_slack(msg.result)
    except Exception as exc:
        post_to_slack(f"Follow-up failed: {exc}")

The TypeScript version is identical in shape, using continue: true. The result subtypes are listed in the agent loop.

Which should I pick?

My rule of thumb: if a human is on the other end, use streaming input. If a scheduler or a webhook is on the other end and the job fits in one prompt, use a single message. Streaming input is also required for changing model or permission mode mid-session and makes approval flows feel natural.