A Practical Guide to Saving Tokens in Claude Code

·Toolin Editorial Team

Master the prompt caching mechanism, session management strategies, and six core rules to make your Claude Code quota last longer and spend better.

A Practical Guide to Saving Tokens in Claude Code

Burning through your Claude Code quota too fast? Some Max users blow through a week's allowance in two days, and a single session can cost over $134 in real terms. The problem usually isn't heavy usage — it's that you don't understand the caching mechanism underneath. This article helps you figure out where tokens actually go and how to save them.

Prompt Caching: The Core of Saving Tokens

Every time a large language model receives your message, it has to "read" the entire input from scratch. In Claude Code, that input usually includes the system instructions, tool definitions, CLAUDE.md project rules, conversation history, and your new message. The first two parts barely change within a session, yet the model re-"reads" them every time.

Prompt caching does something simple: after computing the first time, it stores the intermediate results; the next time it encounters the same input prefix, it reuses the stored results and skips the recomputation. Reading from cache costs one tenth of recomputing.

But caching comes with two preconditions:

  • The cache only works on "prefixes": the match must start from the very beginning and be character-for-character exact. Change a single character on page one and the entire cache is invalidated.
  • The cache has a time-to-live: the main agent's cache window is 1 hour, sub-agents' is 5 minutes. Every cache hit refreshes the timer.

How prompt caching works

Where your cache hits land, and how often, directly determines your token spend

Three Counterintuitive Money-Saving Strategies

Once you understand caching, some "common sense" needs flipping.

While the Cache Is Warm, Continuing Beats a New Session

Every new Claude Code session has to reload the system prompt, tool definitions, CLAUDE.md, and project configuration. That "infrastructure" is roughly 50,000 tokens. Frequent /clear means repeatedly paying the full-price write fee for content that never changes.

In an active session, this content stays cached, and you only pay one tenth of the price each time.

Getting a Complex Task Right Once Beats Three Rounds of Back-and-Forth

Turning off extended thinking does save tokens on a single request. But for a complex refactoring task, doing it in one shot with extended thinking on versus turning it off and iterating for three rounds — the latter is very likely more expensive, because every extra round resends the entire context.

Simple tasks are the opposite: lowering /effort or turning off thinking mode in /config has an immediate effect.

Give Paths for Long Content; Don't Paste It into the Conversation

Don't copy-paste 10,000 lines of logs into the conversation and ask Claude to find the error — just send it the log file's path. Claude Code will use tools like grep to retrieve what it needs, pulling only the relevant parts into context.

Remember: the cheapest tokens are always the ones that never entered the context at all.

Input quality determines token consumption

Controlling input quality is more effective than controlling output length

Keep Chatting or Start a New Session: A Decision Table

This may be the single most important judgment call for saving tokens in Claude Code. Many people's default habit is "clear when done," but the cheapest default is actually the reverse: keep going as long as you can, and treat starting a new session as a conditionally triggered operation.

Continue the current session when:

  • The task hasn't changed — still debugging the same bug or writing the same module
  • Less than 1 hour since the last message, so the cache is still alive
  • The content in context is still useful for the current work

Start a new session when:

  • The task has changed and the two contexts are completely different
  • Idle for more than 1 hour, so the cache has most likely expired
  • The context is stuffed with irrelevant content and too noisy

In one sentence: cache warm and task unchanged, keep chatting. Cache expired, task switched, or context too noisy — restart without hesitation.

A one-task-per-session working style almost never triggers quota problems

The 1M Context Window: Use with Caution

As of March 2026, Max, Team, and Enterprise plans default to Opus 4.6's 1M context window. Anthropic removed the 2x price premium for long context, but the 1M window is becoming the number-one reason many people's quotas hit bottom.

The problem lies in the cost of cache invalidation. You build up a very long session on the 1M window, step away from your computer for over an hour, come back and continue — the entire 1M-token cache has expired, and a single message triggers a full rebuild.

Most everyday sessions trigger compaction at 80-120K of context — they never even reach 200K, let alone 1M.

If you want to disable the 1M context window, add this to ~/.claude/settings.json:

{
  "env": {
    "CLAUDE_CODE_DISABLE_1M_CONTEXT": "1"
  }
}

If you want to set the auto-compaction threshold for context:

{
  "env": {
    "CLAUDE_CODE_AUTO_COMPACT_WINDOW": "200000"
  }
}

Context gets automatically compacted and summarized as it approaches 200,000 tokens, preserving continuity while keeping costs from spiraling.

Six Core Operating Rules

1. Use Sonnet for Everyday Work

Opus's input cost is about 1.7x Sonnet's, and it burns tokens roughly twice as fast. Sonnet is enough for most coding tasks; save Opus for complex architecture decisions and multi-step reasoning. Type /model in Claude Code to switch.

2. Don't Switch Models Mid-Session

Prompt caches are isolated per model. If you've built up 100,000 tokens of cache on Opus and switch to Sonnet to ask a simple question, Sonnet has to build its own cache from zero. When you need a lighter model, use a sub-agent instead of switching the main model.

3. Keep CLAUDE.md Lean and Control the Number of Skills

CLAUDE.md's content gets injected into every request. The official advice is to keep it within 200 lines and retain only rules that genuinely hold long term. More skills isn't better either — remember to turn off MCP services you aren't using.

One small trick: write maintainer notes in CLAUDE.md as HTML comments — Claude strips comments before injecting the context, so they cost no tokens.

4. CLI First, MCP Second

GitHub's gh CLI consumes far fewer tokens than the GitHub MCP server. If the command line can do it, don't install an MCP.

5. Spend a Few Tokens Planning First

For complex tasks, enter plan mode first — let Claude explore the code and propose an approach before implementation. What's truly expensive is rescanning the code, rewriting the implementation, and rerunning tests after the direction proves wrong.

6. Use permissions.deny to Limit What Gets Read

In .claude/settings.json, use permissions.deny to strictly limit what the model can read:

{
  "permissions": {
    "deny": [
      "Read(./.env)",
      "Read(./.env.)",
      "Read(./secrets/)",
      "Read(./node_modules/)",
      "Read(./build)"
    ]
  }
}

Cut unnecessary file reads at the source to avoid wasted tokens

Delegate Tasks to Reduce Main-Session Consumption

Sub-agents: Claude Code's sub-agents have their own context and return only a brief summary to the main session when done. The detailed output of code reviews, test runs, and doc lookups never sits in the main session, so you don't pay for that content on every subsequent message.

Codex plugin: If you also have an OpenAI subscription, people in the community use the following command to offload some tasks:

claude mcp add codex -- npx -y @openai/codex-plugin-cc

Good candidates to delegate: structured bug fixes, code reviews, writing tests. Keep for Claude Code: architecture design, cross-file refactoring, and complex work that requires understanding the whole codebase.

Wrap-Up

The core idea of saving tokens comes down to one sentence: maximize cache hits and keep irrelevant content out of your context.

Starting a new session is a means, not the default. Once you understand prompt caching, you'll see that "continuing to work in an active session" is the default strategy, and "starting a new session" is an optimization triggered by specific conditions.

Related articles

OpenWorker: Andrew Ng's Open-Source Desktop AI Coworker That Ships Finished Work
AI Products

OpenWorker: Andrew Ng's Open-Source Desktop AI Coworker That Ships Finished Work

Andrew Ng open-sources OpenWorker (MIT), a local-first, model-agnostic desktop agent connecting to 25+ tools, turning "prepare a customer brief" into a finished document you can actually open.

Toolin Editorial Team
AutoMIA: Two Images In, a 3D-Printable Mirror Illusion Out — Runs on a Single RTX 3090
AI Products

AutoMIA: Two Images In, a 3D-Printable Mirror Illusion Out — Runs on a Single RTX 3090

AutoMIA, a Tsinghua CVPR'26 Highlight open-source project, turns two input images into a 3D-printable voxel mirror-illusion art model (STL-ready) — single RTX 3090, 76 seconds per design.

Toolin Editorial Team
Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL
AI Products

Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL

Mind Lab open-sources Macaron-V1 (Venti / Coding-Venti / Tall tiers), a Mixture-of-LoRA architecture built on GLM-5.2 744B plus four 1B LoRA experts, with a free hosted API and a UI4A plugin.

Toolin Editorial Team
UniWorld-View: Single Image/Video to Any Camera Trajectory, Tops the WorldScore Leaderboard, Fully Open-Sourced
AI Products

UniWorld-View: Single Image/Video to Any Camera Trajectory, Tops the WorldScore Leaderboard, Fully Open-Sourced

UniWorld-View, open-sourced by Peking University + Rabbitpre + Pengcheng Laboratory, generates novel-view videos along any camera trajectory from a single image/video — it tops Fei-Fei Li's team's WorldScore leaderboard, with code and weights fully open under Apache-2.0 and one-click downloads.

Toolin Editorial Team
Claude Managed Agents Ships Six Updates: Skill Cap Raised to 500, With Ready-to-Run Payloads
AI Products

Claude Managed Agents Ships Six Updates: Skill Cap Raised to 500, With Ready-to-Run Payloads

Anthropic's managed agent platform CMA ships six updates at once: per-session skills up from 20 to 500, a five-level effort setting writable into per-agent config, and seeded sessions created in one step with initial_events.

Toolin Editorial Team
DojoAgents: Build a Financial Research Agent Locally in 10 Minutes
AI Tutorials

DojoAgents: Build a Financial Research Agent Locally in 10 Minutes

An open-source agent framework from Shenchong Intelligence covering A-shares, US stocks, and Hong Kong stocks — deploy a financial research agent that autonomously analyzes market themes, locally, in 10 minutes.

Toolin Editorial Team