Netflix Open-Sources Headroom: Cutting 90% of Wasted Tokens and Saving $700,000
Headroom, an open-source tool built by a senior Netflix engineer, trims redundant tokens before prompts ever reach the LLM — reversible compression, lossless context — and has already saved users $700,000.


Netflix Open-Sources Headroom: Cutting 90% of Wasted Tokens and Saving $700,000
Headroom, an open-source tool built by a senior Netflix engineer, trims redundant tokens before prompts ever reach the LLM — reversible compression, lossless context — and has already saved users $700,000.
Encouraging engineers to use AI freely can produce staggering bills — Uber and Microsoft's COO have both learned this firsthand. Headroom, an open-source tool built by senior Netflix engineer Tejas Chopra, slims prompts down token by token before they reach the model, cutting up to 90% of redundant tokens by its estimates. The little tool, released in January 2025 and still at v0.22, has already saved users about $700,000 — roughly 200 billion tokens. If your team is getting burned by AI bills, this one is worth a look.

What Headroom Is
Headroom is a local proxy tool built on Python and Node that runs on the engineer's own device (listening on port 8787 by default). It works by wrapping the LLM client, automatically parsing and slimming the input.
The simplest usage is wrapping your LLM client on the command line:
headroom wrap codexThe tool intercepts every request headed to the model and slims it down before it arrives. GitHub: github.com/chopratejas/headroom, currently at 2000+ stars and 120+ forks.
What It Slims Down
A $287 bill from Claude Sonnet first made Chopra notice the token cost problem. On close inspection he found the real problem was not hand-written instructions but the accompanying boilerplate code and machine metadata:
- Redundant, verbose JSON structures
- Layer upon layer of templated content inside API responses
- Masses of repeated database fields
"This is not prose, and it is not creative writing — it is compressible data masquerading as text," Chopra wrote. In 2025 a group of researchers found that reading user input accounts for about 76% of all token consumption.
What Headroom slim best:
- Server logs: about 90% can be discarded
- MCP tool outputs: about 70% is redundant JSON
- Database outputs and file trees: huge amounts of repeated metadata
The Four-Step Pipeline
Step 1: CacheAligner
It looks only for what has changed among content already submitted and sends the model only the new parts, skipping the replacement of the vast majority of unchanged full text inside the KV cache.
💡 Tip: If your system prompt carries a date field or a UUID that changes automatically every turn, every call produces a cache miss and costs spike dramatically. Headroom detects and handles this class of problem automatically.
Step 2: Routing by Data Type
Routing identifies the data type and forwards it to the matching compressor:
- AST compressor: compresses program code (abstract syntax trees)
- JSON compressor: strips redundant JSON
- DOM compressor: strips web template code
A set of slimming processors also filter for meaningful content using statistical analysis, iterating on compression levels through a feedback loop — judging whether compression is too aggressive or too lax by how often the model pulls back the original uncompressed prompt.
Step 3: CCR (Compressed Cache and Retrieve)
This is Headroom's biggest difference from other token-slimming tools: reversible compression.
Data is tagged at the point of compression; if the model needs the original context, it can pull the corresponding content back through the Headroom MCP from the user's local device, where the raw data is stored uniformly in a Redis or SQLite database.
Step 4: The Model Retrieves On Demand
When needed, the model actively requests the original uncompressed data through the Headroom MCP instead of passively accepting everything. This mechanism lets compression be more aggressive while preserving accuracy.
The Core Advantage: Reversible Compression
Token-slimming tools are hardly scarce, but Headroom is unique in being losslessly reversible:
- Commercial: Token Company (backed by Y Combinator), offering token compression as a service.
- Open source: RTK (Rust Token Killer) prunes the output of verbose commands; LeanCTX is an RTK variant.
These tools are all useful, but Headroom embeds the operation inside the developer's workflow and offers something none of the others can: reversible compression. The caching offered by model vendors has limits too: Claude's prefix cache defaults to just 5 minutes, and after idling the entire context window has to be refreshed; a one-hour TTL configuration does exist, but "write costs double to buy a 90% saving on reads."
Why Slim Tokens at All
Beyond saving money, token slimming brings two more important benefits.
First, it improves model output quality. When agents push more context than the model actually needs, it not only adds overhead but degrades generation quality. Stanford research found that large models attend most to the beginning and end of the context window and tend to ignore the middle. The research team at data integrator Chroma calls this phenomenon "Context Rot" — across 18 large models, the longer the input text, the less stable the output.
Second, it lowers response latency. Chopra mentions a user who replicated Headroom for a voice application — with voice, even silence produces tokens, users expect responses within 200 milliseconds to sound natural, and Headroom shrinks the latency window as much as possible.
Use Cases
Headroom suits these kinds of users:
- Teams getting burned by AI bills: especially heavy agent-workflow usage where context routinely carries large tool outputs and logs.
- Voice applications: latency-sensitive scenarios that need input compressed to the maximum.
- Energy-conscious teams: fewer tokens mean smaller context windows and less energy consumption.
⚠️ Note: Chopra admits the tool stack still needs polish, especially around testing accuracy. But because CCR stores the original prompts, dedicated compressors for special data types like financial data can be built later. Validate on non-critical workflows first, then expand to production.
A related project, Headlight, is about to be open-sourced; it will track where every token came from, which could prove especially useful for ensuring accuracy in multi-model work.
Related articles

Doubao Seed-Audio 1.0 Hands-On: Character Dialogue, Sound Effects, and BGM in One Pass
Volcano Engine's Seed-Audio 1.0 upgrades to film-grade all-element direct output — a single prompt generates multi-character dialogue, sound effects, and background music, approaching finished-production sound.

WeChat's AI Assistant "Xiaowei" Hands-On: 12 Scenarios to Map Its Powers and Limits
WeChat's native AI assistant Xiaowei has opened gray-scale testing. Built on Tencent's in-house WeLM model, it can send messages, check bills, and analyze Moments — but no scheduled sending or bulk operations yet.

Alibaba HappyHorse 1.1 Hands-On: The Greasy Look Is Gone, and 1080P Just Got 25% Cheaper
Alibaba releases the HappyHorse 1.1 video generation model with upgrades across five dimensions; 1080P drops from ¥1.2 to ¥0.9 per second, with hands-on comparisons and access links.

Baidu Open-Sources Unlimited OCR: a 500M-Active Small Model Reads 40 Pages in One Pass Without Forgetting
Baidu open-sources Unlimited OCR, an end-to-end OCR model with 3B total / 500M active parameters that sets a new OmniDocBench SOTA and transcribes dozens of pages per inference without forgetting.

Sakana Fugu: the Orchestrator That Doesn't Answer Itself, Just Directs Other Models to Do the Work
Sakana AI releases the Fugu family of orchestrator models, smartly dispatching GPT, Claude, and Gemini to finish tasks, with performance approaching Fable 5 and Mythos Preview.

Seko Infinite Canvas Hands-On: Drop in One Idea, and the Agent Finishes a Whole AI Video for You
A hands-on guide to the Seko infinite canvas + Seedance 2.0: 720P cost down 50% and 1080P down 80%, letting even beginners produce multi-episode AI video epics in 10 minutes.