Headroom: A Token-Slimming Tool for AI Bills, Open-Sourced by a Netflix Engineer, Cuts 90% of Redundant Tokens
Netflix senior engineer Tejas Chopra open-sources Headroom, which losslessly compresses context before it reaches a large model; it has already saved users about $700,000 and 200 billion tokens.


Headroom: A Token-Slimming Tool for AI Bills, Open-Sourced by a Netflix Engineer, Cuts 90% of Redundant Tokens
Netflix senior engineer Tejas Chopra open-sources Headroom, which losslessly compresses context before it reaches a large model; it has already saved users about $700,000 and 200 billion tokens.
You encourage engineers to use AI aggressively, and then the bill explodes — Uber and Microsoft's COO have both recently learned that lesson. Tejas Chopra, a senior engineer at Netflix, has open-sourced a tool called Headroom that losslessly trims agent input, token by token, before instructions reach a large language model. Chopra estimates that up to 90% of tokens are redundant for the chosen model. The open-source tool, only released in January 2026 and currently at v0.22, has gathered about 2,000 stars and 120+ forks on GitHub, and has saved users roughly $700,000 and 200 billion tokens.
What Is Headroom
Headroom is built on Python and Node and runs as a local proxy (port 8787) on the engineer's machine. Wrap an LLM from the command line and it parses input automatically:
headroom wrap codexIts core selling point is something no competing tool offers: reversible compression. Compressed data is tagged, and when the model needs the original context, it can pull the corresponding content from the user's local device (Redis or SQLite) via Headroom MCP. You save tokens without losing information.
What It Compresses Best
Headroom does compress some code and human instructions, but what it trims best is machine-generated boilerplate:
- Server logs: about 90% can be discarded
- MCP tool output: about 70% is redundant JSON
- Database output and file trees: masses of repeated metadata
A $287 bill from Claude Sonnet first brought the problem to Chopra's attention: the unit price looks cheap ($3 per million input tokens), but the data shipped to the model is full of deeply nested JSON, boilerplate API response code, and repeated database fields. As he put it — "this isn't prose writing, or creative writing; it's compressible data disguised as text." Research shows that reading user input accounts for about 76% of all token consumption.
The Four-Step Compression Pipeline
- CacheAligner: looks for only the changed information within already-sent content and sends just the delta, avoiding rewriting unchanged full text into the KV cache. If your system prompt carries a date field or a per-turn UUID, every call is a cache miss and costs spike — this step exists to fix exactly that.
- Routing identifies the data type and dispatches it to the matching compressor.
- Specialized compressors: an abstract syntax tree (AST) compressor handles code; JSON and document object model (DOM) compressors strip redundant JSON and web template markup respectively; a slimming processor filters for salient content using statistical analysis, and iteratively tunes the compression level through a feedback loop (judging whether compression went too far by how often the model fetches back the original prompt).
- CCR (compressed cache and retrieval): lets the model fetch back the original uncompressed data, stored uniformly in Redis or SQLite.
When to Use It
Related research shows that managing tokens well both saves money and improves output. When agents push more context than the model actually needs, the extra overhead comes with degraded output — Stanford researchers found that LLMs attend most to the beginning and end of the context window and tend to ignore the middle, and the Chroma team dubbed the phenomenon of "longer input, less stable output" Context Rot. On top of that, leaner prompts cut response latency; one fork of the tool already serves voice-interaction applications (which are extremely sensitive to a 200-millisecond response window).
Typical scenarios where Headroom fits:
- Teams whose agents fire constantly and whose bills are out of control
- Development flows heavy on MCP tool calls, database queries, and file-tree traversal that generate massive redundant JSON
- Latency-sensitive real-time applications (voice, interactive demos)
How It Compares With Other Tools
Model vendors offer their own token-cost optimization tools, but they're fairly opaque to end users. Claude's prefix caching, for example, defaults to a 5-minute TTL, and the one-hour TTL stated in the docs comes with a trap — "your write costs double to earn a 90% saving on reads," and you have to find the balance yourself.
On the commercial side there's Y Combinator-backed Token Company (token compression as a service); in open source there's RTK (Rust Token Killer, which prunes verbose command output) and its variant LeanCTX. Chopra concedes all of these are useful, but Headroom's differentiation is embedding the operation into the dev workflow and offering reversible compression.
Tip: Headroom's toolchain still has gaps, particularly around testing accuracy. CCR stores the original prompts, so dedicated compressors could later be built for special data types like financial data; compression for audio, images, and video is also in progress — a fork already uses it for video parsing, and a related project, Headlight, will soon be open-sourced to track the provenance of every token, which should be especially useful for multi-model accuracy.
Fewer tokens mean a smaller context window and lower energy consumption — provided the Jevons paradox hasn't kicked in first.
Related articles

Tabbit: A Permanently Free AI Browser with 10+ Top Models at Your Fingertips
Meituan launches AI browser Tabbit V1.0 with core features permanently free, 10+ top domestic large models and agent capabilities built in, plus one-click access to 300+ ready-made trick skills.

Finding Overseas Influencers with AI Employees: A Hands-On AhaCreator Guide
A step-by-step walkthrough of the full overseas influencer marketing workflow with AhaCreator — from creator sourcing and content review to cross-border payouts. Ideal for indie developers and teams going global.

AI Long-Video Generation: A Head-to-Head Review of Two Open-Source Frameworks
VideoClaw and JoyAI-Echo, two open-source frameworks released the same day, tackle AI long-video consistency through multi-agent collaboration and cross-modal memory banks respectively — this article compares their technical approaches.

ChatGPT's Memory System Gets a Full Overhaul: Dreaming V3 Is Live
OpenAI ships the new Dreaming V3 memory architecture — ChatGPT now "dreams" in the background to organize what it knows about you. Free access for 1 billion users for the first time, with doubled memory capacity for Plus/Pro.

Cloudflare Integrates Claude Managed Agents: A Developer's Practical Guide
Cloudflare adds support for Claude Managed Agents, letting developers run Claude agents on Cloudflare's platform, connect to private systems, and deploy AI agents securely.

Claude Code vs Codex: A Panoramic Timeline of 24 Features
A timeline breakdown of the 24 features both AI coding agents share — Claude Code shipped 18 of them first, but the gap is now closing in days, not months.