The caveman Plugin: Making Claude Code Cut the Fluff and Save Tokens
The caveman plugin that blew up on GitHub with 20K stars saves tokens by streamlining AI output style. It supports Claude Code and Codex, with three compression levels to switch between as needed.


The caveman Plugin: Making Claude Code Cut the Fluff and Save Tokens
The caveman plugin that blew up on GitHub with 20K stars saves tokens by streamlining AI output style. It supports Claude Code and Codex, with three compression levels to switch between as needed.
If you develop with Claude Code or Codex, you have almost certainly hit this pain point: you only want two lines of code, and the AI hands you five paragraphs of flowery prose. The caveman plugin exists to solve exactly this -- it makes the AI talk like a caveman, stripping out every pleasantry that carries no technical meaning, and claims to save about 75% of output tokens.
What Is caveman
caveman is a Skill/plugin that works with both Claude Code and Codex, and its GitHub stars have already topped 20,000. The core idea is simple: have the AI Agent output technical content in the leanest possible way, cutting only the fluff without sacrificing accuracy.
- Project: https://github.com/JuliusBrussee/caveman
- Author: Julius Brussee

caveman's GitHub star growth curve; after blowing up on Hacker News, it soared from a few dozen to 20,000.
How to Install
caveman runs on Claude Code's Skill mechanism. Installation is trivially simple -- one command does it:
npx skills add JuliusBrussee/cavemanAfter installation, trigger any of the following keywords in Claude Code to activate it:
caveman modetalk like cavemanuse cavemanless tokensbe brief/caveman(slash command)
Three Compression Levels
caveman supports three compression levels -- choose as needed:
lite: Strip filler words, keep full sentences
For scenarios that call for professional, concise output.
"A connection pool reuses database connections that are already open instead of creating a new one for every request, avoiding repeated handshake overhead."
full (default): Fragmented sentences + shorter words
Classic caveman style; the function words get cut.
"Connection pool reuses open DB connections. Not a new one per request. Saves handshake overhead."
ultra: Heavy abbreviation + arrow notation
Maximum compression; common terms replaced with abbreviations.
"Conn pool=reuse DB conns. Skip handshake->faster at high concurrency."
Real-World Comparison
Here is an output comparison for a bug-fix scenario:
Normal mode:
Of course! I'd be happy to help. The problem you are running into is most likely caused by the token-expiration check logic in the authentication middleware. Specifically, the condition uses a less-than sign instead of a less-than-or-equal-to sign...
caveman mode:
Bug in auth middleware. Token expiration check uses
<, not<=. Fix here:
Code blocks and error messages are left untouched; only the natural-language description gets compressed.
Core Compression Rules
caveman's SKILL.md defines the specific compression strategy:
- Delete: articles, filler words, pleasantries, hedging language
- Allow: short sentences, sentence fragments
- Prefer: shorter synonyms ("big" over "extensive", "fix" over "implement a solution")
- Keep unchanged: technical terms, code blocks, error messages
- Recommended pattern: [Problem][Action][Reason]. [Next step].
Things to Know
Before deciding to adopt it, there are a few important limitations to understand:
1. It compresses visible output only, not the thinking process
The author himself has clarified on Hacker News: caveman does not touch hidden reasoning tokens or thinking tokens. The model's background reasoning does not get any shorter; only what it finally says out loud is compressed.
2. The Skill itself consumes context
Loading the caveman Skill takes a bite out of your context budget (Anthropic's official docs say so). So the true end-to-end cost savings may not match the "75%" in the README.
3. Where the 75% number comes from
The author provides a benchmark script and token comparisons for a handful of tasks (ranging from 22% to 87%, averaging 65%), but he also notes these are preliminary tests, not rigorous benchmarks.
Token savings comparison data; the savings ratio varies widely across tasks.
Is There Academic Backing?
Two related papers provide background for the idea that concise output does not necessarily hurt performance:
- The 2024 paper "The Benefits of a Concise Chain of Thought" found that when a model is asked to use a concise chain of thought, response length drops by 48.70% with almost no noticeable loss in problem-solving ability
- The 2026 paper "Brevity Constraints Reverse Performance Hierarchies" reports that adding brevity constraints to large models can raise accuracy by 26 percentage points
But both papers study generic conciseness prompting strategies; neither is a dedicated evaluation of caveman.
When It Fits
caveman suits these scenarios:
- Everyday coding tasks: bug fixes, new features, utility scripts -- no need for the AI to write prose
- Tight token budgets: projects billed by the token, where every cent counts
- Bulk code changes: editing many files at once, with less redundant output
- Agents in CI/CD: automated pipelines that do not need human-friendly explanations
Scenarios where it does not fit:
- Learning a new concept, when you need the AI to explain in detail
- Code review, when you need the AI to analyze design rationale
- Collaborating with non-technical team members
The fact that caveman went viral is itself a signal: developers have had enough of verbose AI output. When users would rather make the AI talk like a "caveman" than keep paying for redundant tokens, it is a sign that "restraint" should be a baseline capability of AI tools.
Toolin Editorial Team
Categories
Related articles

NVIDIA RTX Spark: NVIDIA Redefines the AI PC, 128G Unified Memory Runs a 120B Model Locally
NVIDIA has released the RTX Spark consumer AI chip: 128GB of unified memory and 1 PFLOP of compute, able to run a 120B large model locally on a 14mm laptop — the Windows ecosystem enters the AI PC era

Step 3.7 Flash Hands-On: 400 TPS Inference, Agent Tasks at 1/9 the Cost of Claude
StepFun has released Step 3.7 Flash: 400 tokens/second inference, 11B activated parameters delivering 97% of Claude Opus 4.6's performance, open source and locally deployable

ClawGym: An Open-Source Framework Unifying Agent Training and Evaluation
RUC has open-sourced a full-pipeline Claw Agent framework spanning data, training, and evaluation, with 13.5K executable tasks and support for sandbox-parallel reinforcement learning

Codex Computer Use Lands on Windows: A Hands-On Guide
OpenAI Codex now officially supports operating Windows PCs, with complete setup steps, limitation notes, and how to control it remotely from your phone

Gamma-World: An Open-Source Multi-Agent World Model
NVIDIA and Tsinghua have open-sourced a multi-agent world model: trained on two players, it generalizes directly to four, supporting zero-shot real-time rollouts of multiplayer scenarios

Step 3.7 Flash in Claude Code: A Hands-On Guide
A hands-on test of StepFun's open-source Flash model integrated into Claude Code, using complex Agent workflows to see whether a Chinese model can stand in for closed-source foundations