The Claude Advisor Strategy in Practice: Opus as the Strategist, Sonnet Does the Heavy Lifting

·Toolin Editorial Team

Anthropic has launched the Claude Advisor Strategy and the Monitor tool: one line of code lets Opus direct Sonnet/Haiku from behind the scenes, cutting costs by 85%

The Claude Advisor Strategy in Practice: Opus as the Strategist, Sonnet Does the Heavy Lifting

Developers building AI Agents face a dilemma: Opus is smart but expensive, Sonnet is cheap but fumbles at critical moments. Anthropic's newly released 'Advisor Strategy' and Monitor tool turn this hard problem into a matter of one API parameter.

This article walks you through both features hands-on, from API calls to cost calculation, without skipping a step.

What Is the Advisor Strategy

The traditional approach is to make the strongest model the 'orchestrator', breaking big tasks into subtasks for smaller models. The problem: whether the task is simple or complex, every request starts by burning the most expensive tokens.

The Advisor Strategy flips it around: Sonnet/Haiku act as executors on the front line, and only when a hard problem shows up do they consult Opus. Opus never talks to humans directly and never calls tools; it only reads the context and offers advice.

Measured results:

CombinationSWE-bench score changeCost change
Sonnet 4.6 + Opus advisor+2.7%-11.9%
Haiku 4.5 + Opus advisorPerformance doubledCost only $1.07 (vs Sonnet $7)
Haiku 4.5 + Opus (BrowseComp)19.7% -> 41.2%Cost down 85%

Enable the Advisor Strategy with One Line of Code

Just declare advisor_20260301 in your Messages API request. The model handoff happens silently inside a single /v1/messages request -- no extra callbacks or context management needed.

response = client.messages.create(
    model="claude-sonnet-4-6",  # executor
    tools=[
        {
            "type": "advisor_20260301",
            "name": "advisor",
            "model": "claude-opus-4-6",  # advisor
            "max_uses": 3,  # at most 3 consultations with Opus
        },
        # ... your other tools
    ],
    messages=[...]
)
# The tokens the advisor consumes are listed separately in usage

Key parameters:

  • model: the executor model (Sonnet or Haiku)
  • advisor_20260301: the advisor tool type, a fixed value
  • advisor.model: the advisor model (usually Opus)
  • max_uses: caps the number of advisor calls per request, preventing token overspend

How Token Costs Are Calculated

The tokens the advisor consumes are billed at Opus pricing; the tokens the executor consumes are billed at Sonnet or Haiku pricing.

The key point: each time the advisor weighs in, it generates only a short plan, usually 400 to 700 tokens. The bulk of the actual output is handled entirely by the executor at a much lower rate.

Which combination fits which scenario:

  • Haiku + Opus advisor: high-concurrency, budget-sensitive scenarios, the lowest cost
  • Sonnet + Opus advisor: the everyday choice balancing performance and cost
  • Opus + Opus advisor: highest quality, for mission-critical tasks

Tip: Use max_uses to control how often the advisor is called. The advisor's token consumption is listed separately in the usage information, making it easy to track each model tier's spend.

The Monitor Tool: From Polling to Event-Driven

On the same day, Anthropic also launched the Monitor tool. It solves a different money-burning problem: Agents idling.

In the past, when you had an Agent watch a task (say, waiting for CI to finish), it had to loop and check constantly, burning a round of tokens with every check. Monitor lets Claude write its own background monitoring code, shifting from 'active polling' to 'event-driven'.

Typical uses:

  • Watch system logs for errors continuously, and call the Agent in only when something is wrong
  • Automatically track the status of a GitHub PR; a script polls in the background while the Agent burns no tokens
  • Wake the Agent after an external event arrives (an email lands, a deploy finishes)

Monitor and the Advisor Strategy share the same logic: find the step that does not need to cost money and strip it out. One saves on model calls, the other saves on idle loops.

How to Wire It Up

Step 1: Confirm API Access

Make sure your Anthropic API account is active and has enough token quota. The Advisor Strategy is currently live on the Claude platform in beta.

Step 2: Modify Your Existing API Calls

In your existing Agent code, find the client.messages.create call and add the advisor tool configuration:

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=4096,
    tools=[
        {
            "type": "advisor_20260301",
            "name": "advisor",
            "model": "claude-opus-4-6",
            "max_uses": 3,
        },
        # your existing tool definitions...
    ],
    messages=[
        {"role": "user", "content": "Your task description"}
    ]
)

# Check the advisor's usage
print(response.usage)

Step 3: Monitor Costs

In the returned usage object, the advisor's and the executor's token consumption are tracked separately. You can use that to tune the max_uses parameter and find the balance between quality and cost.

Step 4 (Optional): Configure Monitor

For scenarios that need long-running monitoring, explicitly ask Claude in the prompt to create a background monitoring script instead of a polling loop.

FAQ

  • Q: How is the Advisor Strategy different from ordinary model routing? A: Traditional routing makes you write the logic for when to switch models. The Advisor Strategy lets the executor decide on its own when to 'call for backup', and everything happens inside a single API request.

  • Q: What is a good value for max_uses? A: 1-2 for simple tasks, 3-5 for complex programming tasks. Start with 3 and tune it based on the usage data.

  • Q: Is the Advisor Strategy compatible with my existing tool stack? A: Fully compatible. The advisor is just an ordinary tool entry in the Messages API request; your Agent can search the web and execute code while consulting Opus.

References

Related articles

Xiaomi Miloco 2.0: Smart Homes Finally Get a True AI Steward
AI Products

Xiaomi Miloco 2.0: Smart Homes Finally Get a True AI Steward

Xiaomi open sources its whole-home AI solution Xiaomi Miloco 2.0 — multimodal perception, proactive intelligence, and household memory bring the Agent into the smart home ecosystem.

Toolin Editorial Team
Agnes AI Makes Its Omnimodal API Free Indefinitely, with 1M Context and 4K Image Generation Upgrades This Week
AI Products

Agnes AI Makes Its Omnimodal API Free Indefinitely, with 1M Context and 4K Image Generation Upgrades This Week

Agnes AI has opened its text, image, and video omnimodal model APIs for free indefinitely, with 1M ultra-long context and 4K ultra-HD text-to-image upgrades landing this week.

Toolin Editorial Team
Hands-On with the AI Version of Alipay: Order McDonald's and Collect Energy with a Single Sentence
AI Products

Hands-On with the AI Version of Alipay: Order McDonald's and Collect Energy with a Single Sentence

The AI version of Alipay has entered beta testing with an AI assistant named A Bao that operates mini programs on command; this post covers how to get an invitation code plus the hands-on experience.

Toolin Editorial Team
Major Claude Design Update: One-Click Design System Import and Two-Way Code Sync
AI Products

Major Claude Design Update: One-Click Design System Import and Two-Way Code Sync

Anthropic has shipped a major Claude Design update with design system import, two-way /design-sync and /design code sync, and one-click export to 9 platforms.

Toolin Editorial Team
GLM-5.2 Open Source, Put to the Test: On the Same Level as Opus 4.8 and GPT 5.5
AI Products

GLM-5.2 Open Source, Put to the Test: On the Same Level as Opus 4.8 and GPT 5.5

GLM-5.2 is the first Chinese open-source model to enter the new "Big Three", with 1M long context and coding ability close to closed-source flagship level, MIT-licensed and ready to use.

Toolin Editorial Team
Kimi Work Adds Goal Mode and a Plugin Center, with 50% Off Quota in June
AI Products

Kimi Work Adds Goal Mode and a Plugin Center, with 50% Off Quota in June

Kimi Work launches Goal Mode with 24-hour continuous autonomous operation and a new Plugin Center supporting Baidu Netdisk, DingTalk, Feishu, and other apps, with all task quota consumption halved in June.

Toolin Editorial Team