Toolin.ai

The Claude Code System Prompt Cut List: 6 Categories of Old Rules to Delete, 2 New Sections to Add

Published · toolin小编

Anthropic cut the Claude Code system prompt by 80% and added 70% back—deleting 6 categories of rules that patched older models, and adding two new sections, Delivering work and Corrections, to manage Opus 5's temper. An actionable checklist you can apply directly to your own agent prompts.

The Claude Code System Prompt Cut List: 6 Categories of Old Rules to Delete, 2 New Sections to Add

Claude Code team member Thariq Shihipar revealed that they cut Claude's system prompt by more than 80% with no noticeable drop on internal coding evals. Two days later, Qoder engineer Chen Cheng captured actual requests and found a more nuanced truth—Opus 4.7 was about 15,225 characters, 4.8 dropped to 4,467, and Opus 5 climbed back to 7,694 (72% longer than 4.8).

The 80% cut is real, and so is the 70% added back. What was cut: rules that patched holes in older models. What was added: guardrails managing Opus 5's new "temper."

This is an actionable checklist for anyone designing their own agent system prompts: 6 categories to delete + 2 sections to add, plus one core principle. Apply it directly against your own CLAUDE.md and system prompts.

The Deleted 80%: 6 Categories of "Patching Old Models"

Digging through internal usage records, Anthropic found that the same request often contained conflicting requirements—the system prompt said "don't add comments," Skills said "add documentation where appropriate," and the user would then ask to "explain the complex logic clearly." The model had to process these contradictory instructions before doing any work, wasting tokens and time.

These rules weren't useless—they were compensating for older models' capabilities. Older models lacked judgment; without hard-coded rules they really would delete files recklessly. Newer models can judge from context on their own, and the old rules have become shackles.

1. Delete "hard-coded expression constraints"; switch to "let the model judge"

Old rules to delete:

  • No comments by default
  • No multi-paragraph docstrings
  • Don't create planning or analysis documents unless asked

Replace with one sentence:

Match the comment density, naming, and conventions of the surrounding code.

The instruction looks smaller, but Claude has to make more judgments—it decides how to write based on the task instead of mechanically executing a dead rule.

2. Delete "tool-call examples"; switch to "design interfaces"

Stuffing tool-call examples into the prompt used to be iron law in agent development. But new models easily treat the approach in the example as the whole answer.

What to do: fewer examples, better interface design. For instance, define the Todo tool's states as pending / in_progress / completed and restrict it to one in-progress task at a time, and Claude can infer the usage on its own.

3. Delete "always-on review and tool instructions"; switch to "load on demand"

Old rules to delete: large amounts of review, verification, and tool instructions permanently resident in the system prompt. Even when the user only changes one line of copy, they first sit through Claude running a long context.

Replace with: split those workflows into standalone Skills, and load tool definitions on demand via ToolSearch. Claude only needs to know where things live and fetches them when truly needed.

4. Delete "repeated reminders"; let it be "said once"

Old models forget instructions easily, so the same rule often had to appear in the system prompt, tool descriptions, and examples. New models understand long context much better—repeated reminders just create conflicts.

What to do: for any given thing, give exactly one place ownership.

5. Delete "preferences piled into CLAUDE.md"; hand off to "automatic memory"

Old rules to delete: piling project notes, user preferences, and work experience all into CLAUDE.md.

Replace with: long-term state goes to automatic memory; CLAUDE.md keeps only a codebase introduction and truly counterintuitive gotchas.

The test is clear:

  • "Write high-quality code"—no need to remind
  • "All types must live in the same file"—this is worth writing

6. Delete "human-translated simplified specs"; hand over "real reference material"

New models no longer need humans to translate every requirement into an "AI-readable" simplified spec. HTML prototypes, test cases, existing code, and scoring rubrics can all serve directly as references.

Core insight: one real prototype mockup tells Claude far more than "the page should feel premium, clean, and airy."

💡 To sum up: what Anthropic deleted wasn't context itself, but judgments humans had pre-made on behalf of older models.

The Added-Back 70%: 2 Sections Managing Opus 5's "Temper"

Most of the 72% by which Opus 5 exceeds 4.8 sits in two new sections: Delivering work and Corrections. They target two new problems that surfaced after the Opus 5 upgrade.

Added section 1: Delivering work (preventing doing too much)

The problem: Opus 5 is more autonomous. It inserts steps the user never asked for and modifies the problem the user wanted solved according to its own judgment. Ask it to fix one error, and it may casually refactor the surrounding code, add tests, update docs, and check similar modules.

Constraints to write into the prompt:

  • Deliver work within the scope the user originally asked for
  • Don't quietly shrink, expand, or change the task because you found a "better approach"
  • Don't complete only the easy parts and present the result as done
  • When a part genuinely can't be pushed forward, finish the remaining parts first, then clearly state what's missing and why
  • Ordinary ambiguity is the model's call; only stop and ask the user when a different interpretation would materially change the outcome

Added section 2: Corrections (preventing explaining too much)

The problem: Opus 5 is far more eager than before to walk users through its own corrections—pointing out which earlier statement was wrong, analyzing the cause in detail, and explaining how it now plans to adjust. In long tasks, many of these corrections never change the final result; they just add output length.

Constraints to write into the prompt:

  • Only explain errors that genuinely affect code, conclusions, or decisions; for small mistakes irrelevant to the outcome, fix them and move on
  • No apologies or preamble
  • No repeated self-criticism
  • No itemized tallies of previous mistakes
  • A user follow-up question ≠ Claude was necessarily wrong earlier
  • When another agent flags an issue, first judge whether it's actually right—don't automatically overturn your own conclusion

💡 The new logic in one line: one section stops Claude from doing too much, the other stops Claude from saying too much.

The Core Principle: Claude Code's Version of the "Bitter Lesson"

Richard Sutton's 2019 essay "The Bitter Lesson" looked back at seventy years of AI history: humans can't resist baking their own experience directly into machines—for chess, tell the machine human grandmaster openings; for speech, hard-code phoneme structures. It pays off immediately, but once compute scales up, these carefully designed rules quickly fail, and general methods that leverage compute end up far ahead.

The lesson is bitter because what gets discarded isn't the dumb design—it's precisely the parts researchers were proudest of and invested in the most.

Claude Code's cuts are a miniature "Bitter Lesson":

While model capability isn't growing, rules fill holes; once model capability grows, rules you refuse to delete become the holes themselves.

Applied to your own agent prompts

Don't write every step of how you work into the prompt, forcing the model to forever walk in the footprints of old experience. What survives model generations is an environment that can keep scaling—which exact path to take is ultimately the model's call.

Before adding a rule to your agent next time, ask yourself one question:

Is this rule patching an old model's hole, or capping a new model's ability?