Streaming input versus single messages
The two ways to send prompts to an Agent SDK agent, what each one supports, and when a long-lived streaming session is worth the extra code.
There are two ways to feed prompts into an Agent SDK agent. You can pass a single string and let the run finish, or you can pass an async generator that keeps yielding messages for as long as the session should live. The second is called streaming input, and it is the recommended default for anything interactive.
This page is about input. If you want tokens to appear on screen as Claude writes them, that is streaming output, a separate option.
Side by side
| Streaming input | Single message | |
|---|---|---|
What you pass as prompt | An async iterable of user messages | A string |
| Lifetime | Long-lived session you keep feeding | One run, then done |
| Images in messages | Yes | No |
| Queue several messages | Yes, processed in order | No |
| Interrupt mid-run | Yes | No |
| Change model or permission mode mid-session | Yes | No |
| Natural multi-turn chat | Yes | Only by calling again with continue or resume |
| After an error result | Session stays alive | Run yields the result, then raises |
| Good for | Chat UIs, IDE plugins, assistants | Cron jobs, CI steps, serverless handlers |
Streaming input
In streaming mode the agent behaves like a long-running process: it accepts new user messages as you produce them, surfaces permission requests, handles interruptions and manages the session for you. You get:
- Image attachments on any message.
- Queued messages that run one after another, with the option to interrupt.
- Every tool and MCP server available throughout.
- Live feedback as responses are produced.
- Context that carries forward between turns without any session bookkeeping.
TypeScript
Pass an async generator of SDKUserMessage objects as prompt. Here a support agent receives a written complaint, and a moment later a screenshot of the error the customer saw:
import { query, type SDKUserMessage } from "@anthropic-ai/claude-agent-sdk";
import { readFile } from "node:fs/promises";
function text(content: string): SDKUserMessage {
return { type: "user", message: { role: "user", content }, parent_tool_use_id: null };
}
async function* conversation(): AsyncGenerator<SDKUserMessage> {
yield text("A customer says invoices fail to send after they changed their company name. Investigate src/mailer.");
const screenshot = await readFile("screenshots/send-error.png", "base64");
yield {
type: "user",
message: {
role: "user",
content: [
{ type: "text", text: "Here is what they saw. Does it match what you found?" },
{ type: "image", source: { type: "base64", media_type: "image/png", data: screenshot } },
],
},
parent_tool_use_id: null,
};
}
for await (const msg of query({
prompt: conversation(),
options: { allowedTools: ["Read", "Grep", "Glob"], maxTurns: 12 },
})) {
if (msg.type === "result" && msg.subtype === "success") console.log(msg.result);
}
Each message gets its own result as it completes. In a real app the generator would await input from a websocket or a queue rather than read a file.
Python
Use ClaudeSDKClient. You can pass an async generator to client.query(), but bear in mind that receive_response() stops at the first result message. For a conversation where you want each answer, the cleaner pattern is one query() and receive_response() pair per message:
import asyncio, base64
from claude_agent_sdk import ClaudeSDKClient, ClaudeAgentOptions, AssistantMessage, TextBlock
async def print_reply(client):
async for msg in client.receive_response():
if isinstance(msg, AssistantMessage):
for block in msg.content:
if isinstance(block, TextBlock):
print(block.text)
async def main():
opts = ClaudeAgentOptions(allowed_tools=["Read", "Grep", "Glob"], max_turns=12)
async with ClaudeSDKClient(opts) as client:
await client.query("Invoices fail to send after a company rename. Investigate src/mailer.")
await print_reply(client)
with open("screenshots/send-error.png", "rb") as f:
data = base64.b64encode(f.read()).decode()
async def with_image():
yield {
"type": "user",
"message": {
"role": "user",
"content": [
{"type": "text", "text": "Here is the error they saw. Does it match?"},
{"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": data}},
],
},
}
await client.query(with_image())
await print_reply(client)
asyncio.run(main())
Debugging a streaming session
- Bad image blocks fail quietly. If an image block's
sourceis missing or not an object, the SDK does not raise. Claude receives a text note such as[Image could not be processed: image block has no source object]and carries on. - TypeScript generator errors are disguised. If your generator throws (a missing file is the usual cause), the stream ends with
Claude Code process aborted by user, not your original error. A long line of minified SDK source may come first, so scroll to the end. When you see that message, look inside your generator. - Python generator errors hang. An exception in your generator is logged at debug level and the session simply stalls. If a streaming session sits silent, turn on debug logging and check the generator.
Single message input
Passing a plain string is simpler, and right when:
- you want one answer and are done,
- you do not need images, interruption or mid-session control, or
- you are in a stateless environment such as a Lambda or Cloud Run job.
Warning: Single message mode cannot attach images, queue messages dynamically, interrupt a run or hold a natural multi-turn conversation.
You can still chain calls by passing continue: true (Python: continue_conversation=True) or resume with a session ID; see sessions.
When a single-message run ends on an error result such as error_max_turns, the SDK yields the final result and then raises an error containing the failure text (in Python, a ResultError). Wrap the loop if your code should keep going:
from claude_agent_sdk import query, ClaudeAgentOptions, ResultMessage
async def nightly_report():
opts = ClaudeAgentOptions(allowed_tools=["Read", "Grep"], max_turns=6)
try:
async for msg in query(prompt="Summarise yesterday's errors in logs/app.log", options=opts):
if isinstance(msg, ResultMessage) and msg.subtype == "success":
post_to_slack(msg.result)
except Exception as exc:
post_to_slack(f"Report failed: {exc}")
try:
async for msg in query(prompt="Which of those errors are new this week?",
options=ClaudeAgentOptions(continue_conversation=True, max_turns=6)):
if isinstance(msg, ResultMessage) and msg.subtype == "success":
post_to_slack(msg.result)
except Exception as exc:
post_to_slack(f"Follow-up failed: {exc}")
The TypeScript version is identical in shape, using continue: true. The result subtypes are listed in the agent loop.
Which should I pick?
My rule of thumb: if a human is on the other end, use streaming input. If a scheduler or a webhook is on the other end and the job fits in one prompt, use a single message. Streaming input is also required for changing model or permission mode mid-session and makes approval flows feel natural.