Netflix Open-Sources Headroom: Cutting 90% of Redundant Tokens and Saving $700,000
Headroom (v0.22), open-sourced by a Netflix senior engineer, runs as a local proxy that intercepts LLM inputs and applies lossless, reversible compression to logs, JSON, and code. It has already saved users roughly $700,000 and 200 billion tokens.


Netflix Open-Sources Headroom: Cutting 90% of Redundant Tokens and Saving $700,000
Headroom (v0.22), open-sourced by a Netflix senior engineer, runs as a local proxy that intercepts LLM inputs and applies lossless, reversible compression to logs, JSON, and code. It has already saved users roughly $700,000 and 200 billion tokens.
AI bills are becoming a chronic headache for many teams. Uber and Microsoft COOs have both learned the hard way: encourage engineers to use AI, and the bill can grow large enough to wipe out the savings from layoffs. Netflix senior engineer Tejas Chopra offers an open-source answer — Headroom (v0.22), which acts as a gate between you and the large model, cutting up to 90% of redundant tokens out of instructions before sending them. Since being open-sourced this January, the project has saved users roughly $700,000 and 200 billion tokens, and collected 2000 GitHub stars.
This piece explains what problem Headroom solves, how it works, and how to install and use it.
What problem it solves
A $287 bill from Claude Sonnet got Chopra digging into token costs. He found that the instructions he wrote by hand were not the main culprit — the real offender was the attached machine metadata: deeply nested JSON, API response templates, repeated database fields, server logs.
"This is not prose writing, and it is not creative writing — it is compressible data disguised as text."
A 2025 study found that reading user input accounts for roughly 76% of all token consumption. Model vendors do offer optimizations like prefix caching, but the default TTL is only 5 minutes, and write costs double in exchange for 90% read savings — the break-even point is something you have to work out yourself.
How Headroom works
Headroom is built on Python and Node and runs as a local proxy (on port 8787) on the engineer's own device, compressing instructions before they reach the large model. The pipeline has several stages:

- CacheAligner: looks for changes only in what has already been sent, forwarding only the new parts to the large model so the entire KV cache is not invalidated. If a system prompt contains a date or UUID that changes every turn — which would otherwise cause a cache miss every time — this step plugs the leak directly.
- Route by data type: traffic is automatically dispatched to the matching compressor.
- Type-specific compressors:
- AST compressor: compresses program code (abstract syntax tree)
- JSON compressor: strips redundant JSON
- DOM compressor: trims web template code
- Reduction processor: filters for essential content using statistical analysis and iterates via a feedback loop — it judges whether compression went too far based on how often the large model reaches back for the original uncompressed prompt.
- CCR (Compressed Cache and Retrieval): lets the large model pull back the original data when needed. Compressed sections are flagged, and the model can fetch the corresponding original content from local Redis or SQLite through Headroom MCP.
Observed compression ratios by type:
| Data type | Discardable share |
|---|---|
| Server logs | ~90% |
| MCP tool output (redundant JSON) | ~70% |
| Database output / file trees | heavy repeated metadata |
The key difference: reversible compression
Headroom's biggest difference from other token-reduction tools on the market is that it is reversible. Commercial offerings such as Y Combinator-backed Token Company sell "compression as a service"; on the open-source side there is RTK (Rust Token Killer) for pruning verbose output, and LeanCTX (an RTK variant). These tools are useful, but most of them prune lossily.
Headroom keeps the original prompt intact locally (Redis / SQLite), so the large model can pull back the full context at any time. That means compression can be aggressive without losing information, and whether it went too far is decided by the model's own "look-back frequency."
💡 Tip: Trimming tokens does not just save money — it can improve output quality. Stanford research found that large models pay more attention to the beginning and end of the context and tend to ignore the middle; the Chroma team observed across 18 large models that "the longer the input, the less stable the output," a phenomenon they named "Context Rot."
How to use it
Headroom wraps your large model calls on the command line:
headroom wrap codexThe tool automatically parses the input and applies compression. It works best on server logs, MCP tool output, database output, and file trees.
GitHub repository: https://github.com/chopratejas/headroom
Who it is for
- Developers burned by token bills: especially agent workflows that lean heavily on MCP tools, RAG retrieval, and long log analysis.
- Voice-driven apps: for latency-sensitive scenarios (users expect responses within 200ms), fewer tokens mean shorter latency. One user has already adapted Headroom for voice applications.
- Energy-conscious teams: fewer tokens = smaller context windows = less energy consumed.
Known limitations
Chopra admits the tool stack is still maturing, particularly around testing accuracy. Compression for audio, images, and video is not done yet (though a user has already adapted it for video parsing); a related project, Headlight, will be open-sourced soon and will trace where every token came from, which should help accuracy in multi-model collaboration.
References
- GitHub: https://github.com/chopratejas/headroom
- Original story (The Register): https://www.theregister.com/ai-ml/2026/05/31/netflix-wiz-creates-app-to-slash-ai-bills-then-open-sources-it/