Headroom: compress what the model reads
How Headroom, an open-source context-compression layer, shrinks tool output, logs and code before they reach the model, with content-aware compressors and reversible caching.
Transcript
Even when output must reach the model, most of it is repetition: repeated JSON fields, log noise, boilerplate.
Headroom is an open-source layer between your tools and the model. It compresses that content before the model reads it.
It is content-aware. JSON, code and logs each get their own compressor, so structure and the key lines survive.
And it is reversible. The original is cached locally, so the model can ask for the full detail when it needs it.
Run it as a local proxy, a library, or an M C P server. Point Claude Code at it, and your workflow stays the same.
Sandbox what you can. Compress what you must.
More in this series
2:29How to cut Claude Code token usage
A workflow that saves tokens: index the codebase once with graphify and share it with the team, plan from the graph and requirements, confirm the plan, delegate to subagents, then validate.
1:26context-mode: keep tool output out of your context
How the context-mode plugin runs commands in a sandbox so only the answer reaches Claude, indexes large content for search, and restores relevant memory after compaction.