Claude Sonnet 5 Arrives: Near-Opus 4.8 Performance at 60% of the Price

·Toolin Editorial Team

Anthropic releases Claude Sonnet 5 with adaptive thinking on by default and a new tokenizer. Pricing stays at $3/$15, with a limited-time $2/$10 rate through August 31.

Claude Sonnet 5 Arrives: Near-Opus 4.8 Performance at 60% of the Price

On June 30, Anthropic released Claude Sonnet 5, positioned as a drop-in upgrade to Sonnet 4.6. It delivers clear gains on coding and agent tasks, with performance approaching the pricier Opus 4.8 while costing just 60% of the latter per million tokens. If you're on Sonnet 4.6, migration cost is minimal; if you were weighing a move up to Opus, this Sonnet deserves a fresh look.

Key Updates

Adaptive Thinking On by Default

This is the most significant behavior change. On Sonnet 4.6, requests without a thinking field don't think; on Sonnet 5, the same request automatically enables adaptive thinking. If you don't want it, you must explicitly pass thinking: {type: "disabled"}.

# Sonnet 5 uses adaptive thinking by default; no extra config needed
# To turn it off:
thinking = {"type": "disabled"}

Meanwhile, the old manual extended thinking (thinking: {type: "enabled", budget_tokens: N}) has been fully removed, and calls using it return a 400 error.

Manual Sampling Parameters No Longer Supported

Setting temperature, top_p, or top_k to non-default values returns a 400 error outright. During migration you must remove these parameters and steer model behavior via the system prompt instead. This restriction was previously introduced on Opus 4.7.

New Tokenizer: Roughly 30% More Tokens for the Same Text

Sonnet 5 ships a new tokenizer: the same input text produces about 30% more tokens than on Sonnet 4.6. This is not an API shape change (request/response/streaming event structures are unchanged), but it materially affects your budget:

  • Token counts: the usage field will read higher for the same text; don't reuse counts from the old model
  • Context capacity: the 1M window is unchanged, but less text actually fits inside it
  • max_tokens ceilings: output limits tuned on the old model may now truncate; re-evaluate them
  • Per-request cost: unit price is unchanged, but the same request carries more tokens, so total cost may rise

The First Sonnet with Real-Time Cyber Safety

Sonnet 5 is the first Sonnet-class model with built-in real-time cyber safety. Requests touching banned or high-risk cybersecurity topics may be refused; refusals come back as HTTP 200 + stop_reason: "refusal", not as errors.

Pricing

ItemPrice (per million tokens)
Standard input$3
Standard output$15
Limited-time input (through 2026-08-31)$2
Limited-time output (through 2026-08-31)$10

For comparison, Claude Opus 4.8 is $5/$25. Note that because of the new tokenizer, the actual spend on equivalent requests may be higher than the unit-price comparison suggests.

Availability

  • Claude API: available to all customers
  • AWS: via Claude on Amazon Bedrock and Claude Platform on AWS; legacy Bedrock (InvokeModel/Converse API) is not supported
  • Google Cloud: via Claude on Google Cloud
  • Microsoft Foundry: via Claude in Microsoft Foundry
  • Supports ZDR (zero data retention) agreements

Migration Checklist

Sonnet 5 is a drop-in replacement for Sonnet 4.6 — changing the model ID is enough:

model = "claude-sonnet-4-6"  # before migration
model = "claude-sonnet-5"    # after migration

Then check three things, in order:

  1. Token budget and counting: re-count your prompts with the token counting tool, and re-evaluate max_tokens values that sit near output-length ceilings
  2. Extended thinking: if you're still using budget_tokens, migrate to adaptive thinking
  3. Sampling parameters: remove non-default temperature/top_p/top_k values

Other API constraints (such as no assistant message prefilling) are the same as on Sonnet 4.6 — no further changes needed.

Where It Fits

  • Coding agents: the biggest gains over Sonnet 4.6 are in coding and agentic tasks
  • Need Opus-level capability on a budget: Sonnet 5 narrows the gap with Opus 4.8 on multiple benchmarks
  • Production workloads already on Sonnet 4.6: a drop-in upgrade with low migration risk

If you rely heavily on fine-grained sampling parameters to tune temperature, or reuse historical token counts extensively for cost forecasting, this upgrade takes some adaptation work. Otherwise, it's the most worthwhile switch in the current Sonnet lineup.

Primary sources: