Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.
Market Signal
Why It Has Market Pull
Headroom is one of the fastest-growing open-source AI-infrastructure projects of 2026, built by a Netflix senior engineer and reaching roughly 41.7K GitHub stars within about five months of launch. It sits between agents and LLM providers to compress tool outputs, logs, files, and RAG chunks by a claimed 60-95% with reversible retrieval, and has drawn mainstream tech-press coverage plus third-party integrations.
- 41,700+ GitHub stars and 2,860+ forks; repo created Jan 2026, so growth is roughly five months.
- Actively maintained: release v0.26.0 on 2026-06-16 with the latest push on 2026-06-20 and 150+ releases total.
- Mainstream press including The Register (May 2026), with an estimated $700K in user token-cost savings reported across ~200B tokens.
- Third-party adoption: Tailscale published an Aperture LLM-gateway hook that runs Headroom's compress() function.
- Deployable as a library, proxy, or MCP server with integrations for LangChain, LiteLLM, Agno and Strands.
feedbacks
What People Are Saying
"Thanks for building headroom, it's quite interesting, useful, and you've made it a joy to use!"HN comment
"I've just made a v0.0.1 adapter to run headroom as a hook with Tailscale's Aperture Gateway"HN comment
"The idea of keeping originals cached and injecting a retrieval tool is the right architecture."Independent review
"Aggregate token savings does not equal bottom-line cost savings if rework cancels out the gains."Independent review
"Can you really cut your AI API bill? I deployed Headroom and measured it"Independent benchmark
"The github seems to be a better source, since it supports both JavaScript and Python"HN comment
"runs Headroom's compress() function (not affiliated, just an adapter)"GitHub repo

















