Tool search
Let an Agent SDK agent work with hundreds or thousands of tools by discovering and loading definitions only when a task needs them.
Once an agent is wired up to a few MCP servers, the tool list grows fast. A CRM server, a ticketing server and a warehouse server can easily expose 150 tools between them, and every one of those definitions costs tokens on every turn. Tool search fixes this by keeping the definitions out of the context window until the agent actually asks for them.
This page covers how it works, how to tune or disable it, and how to name and describe tools so the agent finds them.
Why it matters
Two things go wrong as a tool catalogue grows:
- Context bloat. Fifty tool definitions can run to somewhere between 10,000 and 20,000 tokens. That is space the agent cannot use for files, conversation or reasoning.
- Worse choices. Once more than roughly 30 to 50 tools are loaded at the same time, the model gets measurably worse at picking the right one.
Tool search addresses both. The agent sees a compact summary of what exists, searches when it needs a capability it does not have loaded, and only the matching definitions enter context.
What happens at runtime
When tool search is active:
- Deferrable tool definitions are held back. The agent gets a summary instead.
- When a task needs something that is not loaded, the agent runs a search against the catalogue.
- By default, the five most relevant matches are loaded into context.
- Loaded tools stay available on later turns until the SDK compacts the messages in which they were discovered. After compaction, the agent simply searches again the next time it needs them.
The trade-off is one extra model round trip per search. For a large catalogue that is a bargain, because every turn is lighter. For a small set (fewer than about ten tools that fit comfortably in context) loading everything upfront is usually quicker, and the auto mode below handles that case for you.
Tool search covers every registered tool: remote MCP servers, in-process SDK MCP servers built with custom tools, and the built-in tools that already load on demand. Core built-ins such as Bash, Read and Edit are always loaded upfront and are not deferred.
When it is on
Tool search is on by default. The SDK falls back to loading every definition upfront in these situations:
| Situation | Behaviour | Can ENABLE_TOOL_SEARCH override it? |
|---|---|---|
| Model on the SDK's unsupported-model list | Upfront loading | No |
| Google Cloud Agent Platform with a model older than the Claude 4.5 generation | Upfront loading, because those serving stacks reject the beta header | No |
| Google Cloud Agent Platform with Opus 4.5, Sonnet 4.5, Haiku 4.5 or later | Tool search on | Not needed |
| Microsoft Foundry deployment hosted on Azure | Upfront loading, because the deployment rejects tool search server-side and the SDK detects that | No |
ANTHROPIC_BASE_URL points at a host that is not first-party | Tool search off, as most proxies drop tool_reference blocks | Yes |
CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS is set | Tool search off | No |
Note: Versions of Claude Code before v2.1.221 turned tool search off for every model on Google Cloud's Agent Platform unless you set
ENABLE_TOOL_SEARCHexplicitly. From v2.1.227, an organisation can keep tool search on through managed settings.
Supported models are Claude Sonnet 4.5, Haiku 4.5, Opus 4.5 and anything newer. The same minimum applies on Google Cloud.
The ENABLE_TOOL_SEARCH variable
This environment variable is the main control. Pass it to the agent's subprocess through the env option.
| Value | What it does |
|---|---|
| unset | Default. Tool search on, with the automatic fallbacks listed above. |
true | Force it on, including through a custom ANTHROPIC_BASE_URL. The beta header is sent to the proxy, so requests fail if the proxy cannot handle tool_reference blocks. Still cannot beat the Azure Foundry or older Agent Platform fallbacks. |
auto | Add up the tokens in every deferrable definition and compare with the model's context window. Tool search switches on at 10% of the window. Below that, everything loads upfront. |
auto:N | Same as auto, with your own percentage. auto:3 switches on at 3%. Smaller numbers kick in earlier. |
false | Off. Every definition is sent on every turn. |
In auto mode the threshold is a single combined total: every MCP tool from every server that is not marked alwaysLoad, plus the on-demand built-ins. A mod in one of your plugins can also mark a tool as deferred or upfront, which shifts the count.
Passing it in TypeScript and Python
The two SDKs treat env differently, and this catches people out:
- TypeScript:
envreplaces the subprocess environment entirely. Spreadprocess.envfirst or you lose your API key andPATH. - Python:
envis merged over the inherited environment, so you only pass what you are adding.
Here is an agent for a support desk that talks to a large internal MCP server and only wants tool search once the definitions pass 3% of context:
import { query } from "@anthropic-ai/claude-agent-sdk";
const run = query({
prompt: "Refund order 88412 and post a note on the customer's ticket",
options: {
mcpServers: {
helpdesk: { type: "http", url: "https://mcp.internal.example/helpdesk" }
},
allowedTools: ["mcp__helpdesk__*"],
env: { ...process.env, ENABLE_TOOL_SEARCH: "auto:3" }
}
});
try {
for await (const msg of run) {
if (msg.type === "result") {
console.log(msg.subtype === "success" ? msg.result : `Stopped: ${msg.subtype}`);
}
}
} catch (err) {
console.error("Run failed", err);
}
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions, ResultMessage
async def main():
opts = ClaudeAgentOptions(
mcp_servers={"helpdesk": {"type": "http", "url": "https://mcp.internal.example/helpdesk"}},
allowed_tools=["mcp__helpdesk__*"],
env={"ENABLE_TOOL_SEARCH": "auto:3"},
)
try:
async for msg in query(prompt="Refund order 88412 and note it on the ticket", options=opts):
if isinstance(msg, ResultMessage):
print(msg.result if msg.subtype == "success" else f"Stopped: {msg.subtype}")
except Exception as err:
print(f"Run failed: {err}")
asyncio.run(main())
A one-shot query() raises after it has yielded an error result, which is why both examples wrap the loop. Check subtype on the result message (for example error_during_execution) to see why a run ended. The agent loop page explains result handling in more detail.
Help the agent find the right tool
Search matches against tool names and descriptions, so the quality of both decides whether the agent finds what it needs.
- Descriptive names win.
list_overdue_invoicesmatches far more phrasings thaninv_q. - Descriptions should carry keywords. "List unpaid invoices filtered by customer, due date or amount" beats "Invoice query".
- Tell the agent what categories exist. A line in the system prompt saying what kinds of tools are available gives it something to search for. Append it to the
claude_codepreset rather than replacing the whole prompt:
systemPrompt: {
type: "preset",
preset: "claude_code",
append: "Searchable tools cover billing, the helpdesk and the stock database."
}
In Python the option is system_prompt and takes the same dictionary shape. See modifying system prompts for every way to shape the prompt.
If you build your own MCP servers, I find it worth treating tool descriptions like search-engine copy: front-load the nouns users will actually type.
Limits
- Up to 10,000 tools in a catalogue.
- Each search returns the five most relevant tools by default.
- Requires Claude Sonnet 4.5, Haiku 4.5, Opus 4.5 or a later model.