The caveman Plugin: Making Claude Code Cut the Fluff and Save Tokens

·Toolin Editorial Team

The caveman plugin that blew up on GitHub with 20K stars saves tokens by streamlining AI output style. It supports Claude Code and Codex, with three compression levels to switch between as needed.

The caveman Plugin: Making Claude Code Cut the Fluff and Save Tokens

If you develop with Claude Code or Codex, you have almost certainly hit this pain point: you only want two lines of code, and the AI hands you five paragraphs of flowery prose. The caveman plugin exists to solve exactly this -- it makes the AI talk like a caveman, stripping out every pleasantry that carries no technical meaning, and claims to save about 75% of output tokens.

What Is caveman

caveman is a Skill/plugin that works with both Claude Code and Codex, and its GitHub stars have already topped 20,000. The core idea is simple: have the AI Agent output technical content in the leanest possible way, cutting only the fluff without sacrificing accuracy.

caveman star growth curve

caveman's GitHub star growth curve; after blowing up on Hacker News, it soared from a few dozen to 20,000.

How to Install

caveman runs on Claude Code's Skill mechanism. Installation is trivially simple -- one command does it:

npx skills add JuliusBrussee/caveman

After installation, trigger any of the following keywords in Claude Code to activate it:

  • caveman mode
  • talk like caveman
  • use caveman
  • less tokens
  • be brief
  • /caveman (slash command)

Three Compression Levels

caveman supports three compression levels -- choose as needed:

lite: Strip filler words, keep full sentences

For scenarios that call for professional, concise output.

"A connection pool reuses database connections that are already open instead of creating a new one for every request, avoiding repeated handshake overhead."

full (default): Fragmented sentences + shorter words

Classic caveman style; the function words get cut.

"Connection pool reuses open DB connections. Not a new one per request. Saves handshake overhead."

ultra: Heavy abbreviation + arrow notation

Maximum compression; common terms replaced with abbreviations.

"Conn pool=reuse DB conns. Skip handshake->faster at high concurrency."

Real-World Comparison

Here is an output comparison for a bug-fix scenario:

Normal mode:

Of course! I'd be happy to help. The problem you are running into is most likely caused by the token-expiration check logic in the authentication middleware. Specifically, the condition uses a less-than sign instead of a less-than-or-equal-to sign...

caveman mode:

Bug in auth middleware. Token expiration check uses <, not <=. Fix here:

Code blocks and error messages are left untouched; only the natural-language description gets compressed.

Core Compression Rules

caveman's SKILL.md defines the specific compression strategy:

  • Delete: articles, filler words, pleasantries, hedging language
  • Allow: short sentences, sentence fragments
  • Prefer: shorter synonyms ("big" over "extensive", "fix" over "implement a solution")
  • Keep unchanged: technical terms, code blocks, error messages
  • Recommended pattern: [Problem][Action][Reason]. [Next step].

Things to Know

Before deciding to adopt it, there are a few important limitations to understand:

1. It compresses visible output only, not the thinking process

The author himself has clarified on Hacker News: caveman does not touch hidden reasoning tokens or thinking tokens. The model's background reasoning does not get any shorter; only what it finally says out loud is compressed.

2. The Skill itself consumes context

Loading the caveman Skill takes a bite out of your context budget (Anthropic's official docs say so). So the true end-to-end cost savings may not match the "75%" in the README.

3. Where the 75% number comes from

The author provides a benchmark script and token comparisons for a handful of tasks (ranging from 22% to 87%, averaging 65%), but he also notes these are preliminary tests, not rigorous benchmarks.

Token savings comparison data; the savings ratio varies widely across tasks.

Is There Academic Backing?

Two related papers provide background for the idea that concise output does not necessarily hurt performance:

  • The 2024 paper "The Benefits of a Concise Chain of Thought" found that when a model is asked to use a concise chain of thought, response length drops by 48.70% with almost no noticeable loss in problem-solving ability
  • The 2026 paper "Brevity Constraints Reverse Performance Hierarchies" reports that adding brevity constraints to large models can raise accuracy by 26 percentage points

But both papers study generic conciseness prompting strategies; neither is a dedicated evaluation of caveman.

When It Fits

caveman suits these scenarios:

  • Everyday coding tasks: bug fixes, new features, utility scripts -- no need for the AI to write prose
  • Tight token budgets: projects billed by the token, where every cent counts
  • Bulk code changes: editing many files at once, with less redundant output
  • Agents in CI/CD: automated pipelines that do not need human-friendly explanations

Scenarios where it does not fit:

  • Learning a new concept, when you need the AI to explain in detail
  • Code review, when you need the AI to analyze design rationale
  • Collaborating with non-technical team members

The fact that caveman went viral is itself a signal: developers have had enough of verbose AI output. When users would rather make the AI talk like a "caveman" than keep paying for redundant tokens, it is a sign that "restraint" should be a baseline capability of AI tools.

Related articles

NVIDIA RTX Spark: NVIDIA Redefines the AI PC, 128G Unified Memory Runs a 120B Model Locally
AI Products

NVIDIA RTX Spark: NVIDIA Redefines the AI PC, 128G Unified Memory Runs a 120B Model Locally

NVIDIA has released the RTX Spark consumer AI chip: 128GB of unified memory and 1 PFLOP of compute, able to run a 120B large model locally on a 14mm laptop — the Windows ecosystem enters the AI PC era

Toolin Editorial Team
Step 3.7 Flash Hands-On: 400 TPS Inference, Agent Tasks at 1/9 the Cost of Claude
AI Products

Step 3.7 Flash Hands-On: 400 TPS Inference, Agent Tasks at 1/9 the Cost of Claude

StepFun has released Step 3.7 Flash: 400 tokens/second inference, 11B activated parameters delivering 97% of Claude Opus 4.6's performance, open source and locally deployable

Toolin Editorial Team
ClawGym: An Open-Source Framework Unifying Agent Training and Evaluation
AI Products

ClawGym: An Open-Source Framework Unifying Agent Training and Evaluation

RUC has open-sourced a full-pipeline Claw Agent framework spanning data, training, and evaluation, with 13.5K executable tasks and support for sandbox-parallel reinforcement learning

Toolin Editorial Team
Codex Computer Use Lands on Windows: A Hands-On Guide
AI Tutorials

Codex Computer Use Lands on Windows: A Hands-On Guide

OpenAI Codex now officially supports operating Windows PCs, with complete setup steps, limitation notes, and how to control it remotely from your phone

Toolin Editorial Team
Gamma-World: An Open-Source Multi-Agent World Model
AI Products

Gamma-World: An Open-Source Multi-Agent World Model

NVIDIA and Tsinghua have open-sourced a multi-agent world model: trained on two players, it generalizes directly to four, supporting zero-shot real-time rollouts of multiplayer scenarios

Toolin Editorial Team
Step 3.7 Flash in Claude Code: A Hands-On Guide
AI Tutorials

Step 3.7 Flash in Claude Code: A Hands-On Guide

A hands-on test of StepFun's open-source Flash model integrated into Claude Code, using complex Agent workflows to see whether a Chinese model can stand in for closed-source foundations

Toolin Editorial Team