Skip to content

Context window

What fills Claude Code's context window during a real session, what you see versus what Claude sees, what survives compaction, and how to keep it lean.

The context window is Claude's working memory for a session. Everything it knows about your task at a given moment, from the system prompt to the last test run, has to fit in there. When people tell me Claude Code "got worse" halfway through a long session, nine times out of ten the context window is full of old file reads and the important instructions have been summarised away.

This page follows a realistic session from start to finish so you can see what goes in, when, and what you can do about it. The token numbers are illustrative; your own will depend on your CLAUDE.md, MCP servers and file sizes. Run /context at any point to see the real breakdown for your session.

Three levels of visibility

A useful mental model: every item in context is one of three kinds.

VisibilityWhat it meansExamples
HiddenClaude sees it, you never doSystem prompt, CLAUDE.md, auto memory, hook output
One-linerYou see a short notice, Claude sees the full content"Read auth.ts", "Loaded .claude/rules/api.md", "Running tests..."
FullBoth of you see itYour prompts, Claude's replies, diffs

The gap between a one-liner and its real size is where context disappears. A file read that shows as five words in your terminal may be several thousand tokens for Claude.

Before you type anything

A surprising amount loads before your first prompt:

ItemRough sizeNotes
System promptSeveral thousand tokensCore behaviour, tool use and formatting rules. Always first, never shown
Auto memory (MEMORY.md)Hundreds of tokensClaude's notes from earlier sessions. The first 200 lines or 25 KB, whichever comes first
Environment detailsA few hundred tokensWorking directory, platform, shell, OS version, whether it is a git repo. Branch, status and recent commits arrive as a separate block
MCP tool namesSmallJust the names. Full schemas are deferred and fetched through tool search when needed
Skill descriptionsHundreds of tokensOne line per skill Claude may invoke. Skills with disable-model-invocation: true are left out entirely
~/.claude/CLAUDE.mdVariesYour personal preferences for every project
Project CLAUDE.mdOften the largest item you controlConventions, commands, architecture

Depending on your setup there may be more: an output style, text from --append-system-prompt, or an AGENTS.md read in place of CLAUDE.md.

The lesson is that your first prompt is tiny next to what is already loaded. That is fine as long as what is loaded is useful. Keep CLAUDE.md under about 200 lines and move reference material into skills or path-scoped rules.

Note: You can change how MCP schemas load with the ENABLE_TOOL_SEARCH environment variable. auto loads them up front if they fit within 10% of the window; false loads every schema immediately. The default defers them all.

While Claude works

Suppose you ask: "users on the mobile app get logged out after exactly an hour; find out why". A typical sequence:

  1. Claude reads the session service. You see a one-line "Read" notice. Claude receives the whole file, perhaps a couple of thousand tokens.
  2. It follows imports into the token helper and the middleware. More one-liners, more tokens.
  3. A path-scoped rule loads. Because .claude/rules/backend.md has paths: matching src/server/**, it attaches itself the moment Claude reads a file there. You see "Loaded .claude/rules/backend.md", not its contents.
  4. It reads the existing tests, which trigger a second rule scoped to *.test.ts.
  5. It greps for refreshToken. You see that the search ran; Claude sees every matching line.
  6. It explains the bug (refresh tokens were issued with the access token's expiry). This text is visible to both of you.
  7. It edits the file. You see the diff.
  8. A PostToolUse hook runs the formatter. Silent for you. Only output returned through hookSpecificOutput.additionalContext reaches Claude; plain stdout from a hook that exits 0 goes to the debug log, not into context.
  9. It adds a regression test, and the hook fires again.
  10. It runs the test suite. You see a pass count; Claude sees all the output.
  11. It summarises what changed.

File reads dominate. The easiest saving is a specific prompt: "the bug is in src/server/session.ts" means fewer files opened while hunting.

Tip: For a PostToolUse hook, exiting with code 2 shows stderr to Claude as an error but cannot block anything, because the tool has already run. Hook output over 10,000 characters is saved to a file and Claude gets a preview and the path.

Delegating to a subagent

Now you follow up: "use a subagent to check how idle timeouts are configured across environments, then fix any mismatch".

Claude writes a task and hands it to a subagent, which starts with its own fresh context window:

  • Its own system prompt, shorter than the main one. The main session's auto memory is not included; a custom agent with memory: in its frontmatter loads its own MEMORY.md instead.
  • Its own copy of CLAUDE.md, which counts against its window, not yours. The built-in Explore and Plan agents skip it.
  • The same MCP servers and skills, and most of the parent's tools, minus a few that make no sense when nested (plan mode controls, background task tools and, by default, the Agent tool itself to stop recursion).
  • The task Claude wrote for it, in place of a user prompt.

The subagent then reads config files, environment overrides and the timeout module. Thousands of tokens, none of them in your window. When it finishes, you receive only its final answer plus a small metadata trailer with token counts and duration. Reading six thousand tokens of files to return a four-hundred-token summary is exactly the trade you want.

Compaction

Eventually, either because you run /compact or because the window is nearly full, Claude Code replaces the conversation with a structured summary. The summary keeps your requests and intent, key technical concepts, files examined or changed with important snippets, errors and their fixes, pending tasks and the current piece of work. Verbatim tool output and intermediate reasoning are gone.

From v2.1.198, the summarisation request uses the same extended thinking setting as your session. Your settings are not changed afterwards.

What comes back afterwards

ItemAfter compaction
System prompt and output styleStill in force
Project-root CLAUDE.md and rules without paths:Re-read from disk
Auto memoryRe-read from disk
Git statusA fresh snapshot is taken
A plan written in plan modeRe-read from disk
Rules with paths:Reload only when a matching file is touched again
Nested CLAUDE.md files in subfoldersReload only when Claude works in that folder again
Files Claude read or editedUp to five are re-read, most recently modified first
Bodies of skills you invokedRe-injected, up to 5,000 tokens per skill and 25,000 in total; the oldest are dropped first
Skill description listingNot re-injected; only skills you actually used come back
Background shell commands and background subagentsKeep running; Claude is reminded which are active so it does not start duplicates
Context earlier hooks addedSummarised along with everything else
SessionStart hooks matching the compact sourceRun again, and their output is added

A re-read file larger than 5,000 tokens comes back as a path only, labelled Referenced file rather than Read.

Two practical consequences:

  • Path-scoped rules and nested CLAUDE.md files are summarised away. They live in message history from the moment they load. If a rule must survive compaction, remove its paths: frontmatter or move it into the root CLAUDE.md.
  • Long skills get truncated from the bottom. Put the instructions that matter most near the top of SKILL.md.

In the terminal, compaction shows as a "Conversation compacted" message followed by one-liners for each re-read file and the restored skills.

Staying ahead of a full window

A full context window does not end a session, because auto-compaction kicks in. You will usually get better results by acting first:

  • Compact with a focus. Before a big new chunk of work, run something like /compact keep the timeout investigation and the list of affected environments. You choose what survives rather than leaving it to the automatic pass.
  • Summarise part of the conversation. Run /rewind, pick a message and choose Summarize from here or Summarize up to here. See checkpointing.
  • Compact earlier. /autocompact with a token count, for example /autocompact 400k, sets how full the window gets before the automatic pass. Model configuration lists accepted values.
  • Clear between unrelated tasks. /clear costs nothing and stops yesterday's exploration being re-sent with every message.
  • Delegate heavy reading to subagents, as above.

Bigger windows

If you need more room rather than less conversation, several models support a 1 million token window: Fable models, Sonnet 5 and later, Haiku 5.5, Opus 4.6 and later, and Sonnet 4.6. Most are selected through a [1m] model variant; Sonnet 5.5 and Sonnet 5 always run with the 1M window and have no separate variant. Compaction works the same at the larger size. Model configuration covers availability by plan, the default auto-compact threshold for each model, and how to correct the assumed window when you route through an LLM gateway via ANTHROPIC_BASE_URL or use a custom model ID.

Checking your own session

  • /context gives a live breakdown by category with suggestions, including which CLAUDE.md and auto memory files loaded.
  • /context all also shows the token cost of each loaded MCP tool.
  • /memory opens the instruction and memory files so you can trim them.

I run /context whenever a session starts to feel sluggish or forgetful. More often than not the fix is a /clear and a sharper prompt.