Context window
What fills Claude Code's context window during a real session, what you see versus what Claude sees, what survives compaction, and how to keep it lean.
The context window is Claude's working memory for a session. Everything it knows about your task at a given moment, from the system prompt to the last test run, has to fit in there. When people tell me Claude Code "got worse" halfway through a long session, nine times out of ten the context window is full of old file reads and the important instructions have been summarised away.
This page follows a realistic session from start to finish so you can see what goes in, when, and what you can do about it. The token numbers are illustrative; your own will depend on your CLAUDE.md, MCP servers and file sizes. Run /context at any point to see the real breakdown for your session.
Three levels of visibility
A useful mental model: every item in context is one of three kinds.
| Visibility | What it means | Examples |
|---|---|---|
| Hidden | Claude sees it, you never do | System prompt, CLAUDE.md, auto memory, hook output |
| One-liner | You see a short notice, Claude sees the full content | "Read auth.ts", "Loaded .claude/rules/api.md", "Running tests..." |
| Full | Both of you see it | Your prompts, Claude's replies, diffs |
The gap between a one-liner and its real size is where context disappears. A file read that shows as five words in your terminal may be several thousand tokens for Claude.
Before you type anything
A surprising amount loads before your first prompt:
| Item | Rough size | Notes |
|---|---|---|
| System prompt | Several thousand tokens | Core behaviour, tool use and formatting rules. Always first, never shown |
Auto memory (MEMORY.md) | Hundreds of tokens | Claude's notes from earlier sessions. The first 200 lines or 25 KB, whichever comes first |
| Environment details | A few hundred tokens | Working directory, platform, shell, OS version, whether it is a git repo. Branch, status and recent commits arrive as a separate block |
| MCP tool names | Small | Just the names. Full schemas are deferred and fetched through tool search when needed |
| Skill descriptions | Hundreds of tokens | One line per skill Claude may invoke. Skills with disable-model-invocation: true are left out entirely |
~/.claude/CLAUDE.md | Varies | Your personal preferences for every project |
Project CLAUDE.md | Often the largest item you control | Conventions, commands, architecture |
Depending on your setup there may be more: an output style, text from --append-system-prompt, or an AGENTS.md read in place of CLAUDE.md.
The lesson is that your first prompt is tiny next to what is already loaded. That is fine as long as what is loaded is useful. Keep CLAUDE.md under about 200 lines and move reference material into skills or path-scoped rules.
Note: You can change how MCP schemas load with the
ENABLE_TOOL_SEARCHenvironment variable.autoloads them up front if they fit within 10% of the window;falseloads every schema immediately. The default defers them all.
While Claude works
Suppose you ask: "users on the mobile app get logged out after exactly an hour; find out why". A typical sequence:
- Claude reads the session service. You see a one-line "Read" notice. Claude receives the whole file, perhaps a couple of thousand tokens.
- It follows imports into the token helper and the middleware. More one-liners, more tokens.
- A path-scoped rule loads. Because
.claude/rules/backend.mdhaspaths:matchingsrc/server/**, it attaches itself the moment Claude reads a file there. You see "Loaded .claude/rules/backend.md", not its contents. - It reads the existing tests, which trigger a second rule scoped to
*.test.ts. - It greps for
refreshToken. You see that the search ran; Claude sees every matching line. - It explains the bug (refresh tokens were issued with the access token's expiry). This text is visible to both of you.
- It edits the file. You see the diff.
- A
PostToolUsehook runs the formatter. Silent for you. Only output returned throughhookSpecificOutput.additionalContextreaches Claude; plain stdout from a hook that exits 0 goes to the debug log, not into context. - It adds a regression test, and the hook fires again.
- It runs the test suite. You see a pass count; Claude sees all the output.
- It summarises what changed.
File reads dominate. The easiest saving is a specific prompt: "the bug is in src/server/session.ts" means fewer files opened while hunting.
Tip: For a
PostToolUsehook, exiting with code 2 shows stderr to Claude as an error but cannot block anything, because the tool has already run. Hook output over 10,000 characters is saved to a file and Claude gets a preview and the path.
Delegating to a subagent
Now you follow up: "use a subagent to check how idle timeouts are configured across environments, then fix any mismatch".
Claude writes a task and hands it to a subagent, which starts with its own fresh context window:
- Its own system prompt, shorter than the main one. The main session's auto memory is not included; a custom agent with
memory:in its frontmatter loads its ownMEMORY.mdinstead. - Its own copy of
CLAUDE.md, which counts against its window, not yours. The built-in Explore and Plan agents skip it. - The same MCP servers and skills, and most of the parent's tools, minus a few that make no sense when nested (plan mode controls, background task tools and, by default, the Agent tool itself to stop recursion).
- The task Claude wrote for it, in place of a user prompt.
The subagent then reads config files, environment overrides and the timeout module. Thousands of tokens, none of them in your window. When it finishes, you receive only its final answer plus a small metadata trailer with token counts and duration. Reading six thousand tokens of files to return a four-hundred-token summary is exactly the trade you want.
Compaction
Eventually, either because you run /compact or because the window is nearly full, Claude Code replaces the conversation with a structured summary. The summary keeps your requests and intent, key technical concepts, files examined or changed with important snippets, errors and their fixes, pending tasks and the current piece of work. Verbatim tool output and intermediate reasoning are gone.
From v2.1.198, the summarisation request uses the same extended thinking setting as your session. Your settings are not changed afterwards.
What comes back afterwards
| Item | After compaction |
|---|---|
| System prompt and output style | Still in force |
Project-root CLAUDE.md and rules without paths: | Re-read from disk |
| Auto memory | Re-read from disk |
| Git status | A fresh snapshot is taken |
| A plan written in plan mode | Re-read from disk |
Rules with paths: | Reload only when a matching file is touched again |
Nested CLAUDE.md files in subfolders | Reload only when Claude works in that folder again |
| Files Claude read or edited | Up to five are re-read, most recently modified first |
| Bodies of skills you invoked | Re-injected, up to 5,000 tokens per skill and 25,000 in total; the oldest are dropped first |
| Skill description listing | Not re-injected; only skills you actually used come back |
| Background shell commands and background subagents | Keep running; Claude is reminded which are active so it does not start duplicates |
| Context earlier hooks added | Summarised along with everything else |
SessionStart hooks matching the compact source | Run again, and their output is added |
A re-read file larger than 5,000 tokens comes back as a path only, labelled Referenced file rather than Read.
Two practical consequences:
- Path-scoped rules and nested
CLAUDE.mdfiles are summarised away. They live in message history from the moment they load. If a rule must survive compaction, remove itspaths:frontmatter or move it into the rootCLAUDE.md. - Long skills get truncated from the bottom. Put the instructions that matter most near the top of
SKILL.md.
In the terminal, compaction shows as a "Conversation compacted" message followed by one-liners for each re-read file and the restored skills.
Staying ahead of a full window
A full context window does not end a session, because auto-compaction kicks in. You will usually get better results by acting first:
- Compact with a focus. Before a big new chunk of work, run something like
/compact keep the timeout investigation and the list of affected environments. You choose what survives rather than leaving it to the automatic pass. - Summarise part of the conversation. Run
/rewind, pick a message and choose Summarize from here or Summarize up to here. See checkpointing. - Compact earlier.
/autocompactwith a token count, for example/autocompact 400k, sets how full the window gets before the automatic pass. Model configuration lists accepted values. - Clear between unrelated tasks.
/clearcosts nothing and stops yesterday's exploration being re-sent with every message. - Delegate heavy reading to subagents, as above.
Bigger windows
If you need more room rather than less conversation, several models support a 1 million token window: Fable models, Sonnet 5 and later, Haiku 5.5, Opus 4.6 and later, and Sonnet 4.6. Most are selected through a [1m] model variant; Sonnet 5.5 and Sonnet 5 always run with the 1M window and have no separate variant. Compaction works the same at the larger size. Model configuration covers availability by plan, the default auto-compact threshold for each model, and how to correct the assumed window when you route through an LLM gateway via ANTHROPIC_BASE_URL or use a custom model ID.
Checking your own session
/contextgives a live breakdown by category with suggestions, including whichCLAUDE.mdand auto memory files loaded./context allalso shows the token cost of each loaded MCP tool./memoryopens the instruction and memory files so you can trim them.
I run /context whenever a session starts to feel sluggish or forgetful. More often than not the fix is a /clear and a sharper prompt.