GitHub Cuts Agent Token Costs 62% with Daily Audits
GitHub's gh-aw CLI tool builds an audit-optimize loop through daily token audits and MCP pruning; the Auto-Triage task has sustained a 62% reduction in equivalent token cost.


GitHub Cuts Agent Token Costs 62% with Daily Audits
GitHub's gh-aw CLI tool builds an audit-optimize loop through daily token audits and MCP pruning; the Auto-Triage task has sustained a 62% reduction in equivalent token cost.
If your team runs LLM Agents in CI (scheduled automation tasks), token costs can quietly pile up to staggering numbers. GitHub's engineering team shared a closed-loop "audit-optimize" method that cut the equivalent token cost of Agent workflows by as much as 62%. The tooling is integrated into the gh-aw CLI, and you can borrow this playbook directly.
Equivalent Tokens (ET): A Metric for Cross-Model Cost Comparison
Model prices vary widely, so comparing raw token counts means nothing. GitHub designed an "Equivalent Token (ET)" metric:
- Output tokens are weighted at 4x (output is far more expensive than input)
- Cache-read tokens are weighted at 0.1x
- Coefficients are applied by model type: Haiku 0.25x, Sonnet 1.0x, Opus 5.0x
This way, no matter which model you use, a 10% drop in ET corresponds to roughly a 10% drop in cost.
The Audit-Optimize Loop: Two Agents Running Automatically
Step 1: The Daily Token Usage Auditor
This Agent runs automatically every day and does three things:
- Aggregates resource consumption by workflow: all Agent calls are funneled through an API proxy, and each run generates a
token-usage.jsonlfile - Flags anomalous runs: spots tasks whose consumption suddenly spikes
- Finds the costliest tasks: ranked by priority, marking the workflows that need optimizing
Step 2: The Daily Token Optimizer
When the auditor finds a workflow worth attention, the optimizer kicks in automatically:
- Reads the relevant source code and recent logs
- Analyzes the sources of inefficiency
- Automatically creates a GitHub Issue with concrete optimization suggestions
The two Agents' own token consumption is also counted in the same daily report.
The Biggest Source of Savings: Pruning MCP Tools
The most common source of inefficiency the optimizer finds is unused MCP tools.
The reason: the LLM API is fundamentally stateless, so every request has to carry the tool schemas. A GitHub MCP Server containing 40 tools can add an extra 10KB-15KB of schema content to every turn of interaction.
Concrete optimization moves:
- Delete unused tool definitions: in the smoke-test workflow, this alone reduced context per call by about 8KB-12KB
- Replace MCP calls with the
ghCLI: PR diffs and file contents are instead fetched directly from the command line, with the data pre-downloaded into the working directory before the Agent starts - Avoid exposing auth tokens: data is fetched at runtime through a transparent HTTP proxy
Measured Results
Optimization results across a dozen-plus production workflows:
| Workflow | ET reduction |
|---|---|
| Auto-Triage Issues | 62% (validated over 109 runs) |
| Security Guard | 43% |
| Smoke Claude | 59% |
| Daily Community Attribution | 37% |
Caveats
MCP pruning strategies have their limits. GitHub's own "Daily Community Attribution" workflow removed 8 unused tools, yet ET didn't drop meaningfully. The reason: in that workflow, the tool inventory made up only a small share of the overall context. Profile first to confirm the bottleneck, then make your move.
Tools You Can Use Right Away
Auditor and Optimiser ship as part of the gh-aw CLI. You can:
- Deploy a similar
token-usage.jsonllogging mechanism in your own CI - Take advantage of Anthropic's and OpenAI's prompt caching features
- Use LangChain's Callback mechanism to track Agent token usage
The core idea: the cheapest LLM call is the one that never happens.
Source reference: GitHub Slashes Agent Workflow Token Spend up to 62% with Daily Audits and MCP Pruning - InfoQ
Toolin Editorial Team
Categories
Related articles

Xiaomi Miloco 2.0: Smart Homes Finally Get a True AI Steward
Xiaomi open sources its whole-home AI solution Xiaomi Miloco 2.0 — multimodal perception, proactive intelligence, and household memory bring the Agent into the smart home ecosystem.

Agnes AI Makes Its Omnimodal API Free Indefinitely, with 1M Context and 4K Image Generation Upgrades This Week
Agnes AI has opened its text, image, and video omnimodal model APIs for free indefinitely, with 1M ultra-long context and 4K ultra-HD text-to-image upgrades landing this week.

Hands-On with the AI Version of Alipay: Order McDonald's and Collect Energy with a Single Sentence
The AI version of Alipay has entered beta testing with an AI assistant named A Bao that operates mini programs on command; this post covers how to get an invitation code plus the hands-on experience.

Major Claude Design Update: One-Click Design System Import and Two-Way Code Sync
Anthropic has shipped a major Claude Design update with design system import, two-way /design-sync and /design code sync, and one-click export to 9 platforms.

GLM-5.2 Open Source, Put to the Test: On the Same Level as Opus 4.8 and GPT 5.5
GLM-5.2 is the first Chinese open-source model to enter the new "Big Three", with 1M long context and coding ability close to closed-source flagship level, MIT-licensed and ready to use.

Kimi Work Adds Goal Mode and a Plugin Center, with 50% Off Quota in June
Kimi Work launches Goal Mode with 24-hour continuous autonomous operation and a new Plugin Center supporting Baidu Netdisk, DingTalk, Feishu, and other apps, with all task quota consumption halved in June.