context-mode: keep tool output out of your context
How the context-mode plugin runs commands in a sandbox so only the answer reaches Claude, indexes large content for search, and restores relevant memory after compaction.
Transcript
The biggest token drain is tool output. One test run, one log file or one web page can dump tens of thousands of tokens into the context window.
Context-mode runs the command in a sandbox instead. The raw output stays outside, and only the answer you asked for comes back.
Large content is indexed into a local database, with ranked full-text search.
Later, the agent searches, and gets back only the matching snippets instead of the whole document.
It also records session events: edits, tasks and decisions.
After the context is compacted, it retrieves just the relevant memories, so the agent keeps its thread.
Sandbox the output. Search, don't read. Remember what matters.
More in this series
2:29How to cut Claude Code token usage
A workflow that saves tokens: index the codebase once with graphify and share it with the team, plan from the graph and requirements, confirm the plan, delegate to subagents, then validate.
1:16Headroom: compress what the model reads
How Headroom, an open-source context-compression layer, shrinks tool output, logs and code before they reach the model, with content-aware compressors and reversible caching.