Claude Opus 5 launch: benchmarks and pricing
TECH

Claude Opus 5 launch: benchmarks and pricing

27+
Signals

Strategic Overview

  • 01.
    Anthropic released Claude Opus 5 on July 24, 2026 across Claude Max, Pro, and the API, with a 1M-token context window, priced identically to predecessor Opus 4.8 at $5/$25 per million input/output tokens (fast mode at double that rate).
  • 02.
    The model leads or nearly matches rival Fable 5 across major benchmarks - more than doubling Opus 4.8's score on Frontier-Bench, landing within 0.5% of Fable 5 on CursorBench at half the cost per task, and scoring roughly three times the next-best model on ARC-AGI 3.
  • 03.
    Anthropic expects Opus 5's safety classifiers to trigger about 85% less often than Fable 5's, cutting spurious refusals during coding sessions, and also ships an opt-in 'Automatic Fallbacks' feature that reroutes blocked requests to a weaker model.
  • 04.
    Early enterprise deployments at Zapier, Box, and Cognition's Devin report concrete workflow gains, and Opus 5 marks Anthropic's fourth Claude 5-generation release in under two months.

Deep Analysis

Frontier intelligence at Opus prices: what the benchmarks actually show

Anthropic's benchmark claims are unusually specific about the cost angle, not just raw capability. On Frontier-Bench v0.1, Opus 5 more than doubles predecessor Opus 4.8's score while costing less per task [1]. Independent benchmarking site MarkTechPost logged the exact numbers behind that claim: 43.3% for Opus 5 versus 18.7% for Opus 4.8 and 33.7% for Fable 5 [2]. On CursorBench 3.2 at maximum effort, Opus 5 lands within half a percentage point of Fable 5's peak score at half the cost per task [1], and on OSWorld 2.0, a computer-use benchmark, it beats Fable 5's best result at roughly a third of the cost [1], hitting 70.57% per MarkTechPost's breakdown [2]. Artificial Analysis's own knowledge-work evaluation, AA-Briefcase, put Opus 5 at a 1,720 Elo rating - 146 points ahead of Fable 5's 1,574 - with an Analytical Quality Elo nearly 300 points higher [3]. The pattern across every named benchmark is the same: Opus 5 doesn't need to beat Fable 5 outright to be the more attractive buy, it just needs to get close while costing a fraction as much, at pricing identical to the 4.8 generation it replaces.

Fewer refusals, tighter guardrails - and a case where they still misfired

Anthropic says it expects Opus 5's safety classifiers to trigger around 85% less often than Fable 5's [1], and independent analyst Zvi Mowshowitz's read of the system card confirms the scale of that shift: FrontierBench sessions saw classifier interventions fall from 42% of trials to 5% [4]. The same system-card analysis reports prompt-injection attack success dropping from 5.5% to 2.0%, and computer-use attack success falling from 7.14% to 0.54% with extended thinking enabled [4]- Anthropic is framing Opus 5 as its most aligned model yet, not just its least annoying one. Anthropic also shipped an opt-in 'Automatic Fallbacks' feature that reroutes a blocked request to a weaker model rather than refusing outright [5], an implicit admission that even a tighter classifier still misfires sometimes. Community reaction to the guardrail changes has been split: many describe the model as looser and less prone to false refusals than before, but a vocal subset - particularly people doing legitimate security and network-testing work - reported the classifier still blocked benign requests, before the issue got resolved through Anthropic's own verification channels. An 85% average reduction is still an average, not a guarantee for any individual prompt that happens to resemble something adversarial.

Is this actually a big deal, or incremental? The contrarian read

Not every reviewer bought the benchmark story wholesale. Code-review vendor CodeRabbit ran its own independent test and found a genuine improvement in precision on actionable review comments (39.3% versus 35.2% for the prior model), but also found Opus 5 missed some previously-caught known issues and generated roughly four times as many low-value 'nitpick' comments [6]- a reminder that benchmark gains and day-to-day usefulness don't always move together. Every's hands-on review after a week of real use described the model as 'brilliant in flashes, frustrating in practice' [7], citing friction with existing skills and plugins built for earlier Claude versions. Broader community sentiment echoed that split: alongside real enthusiasm for the cost-adjusted benchmark gains, there was also a Reddit thread praising Opus 5 Low for beating Sonnet 5 High and its looser guardrails, alongside grumbling that usage limits weren't reset at launch - both signs that 'wins the leaderboard' and 'feels like a clear upgrade to daily users' are not the same claim.

Real-world validation and the four-releases-in-two-months cadence

Away from benchmarks, early enterprise integrations offer a second, more grounded data point. Cognition's CEO Scott Wu said Opus 5 approaches Fable-level performance at half the cost inside its Devin coding agent, particularly on debugging and root-cause analysis [8]. Zapier's CEO Wade Foster reported the model topped its internal AutomationBench leaderboard without using more tokens than earlier, identically-priced Claude releases [9]. Box integrated Opus 5 and measured an 8% overall improvement over Opus 4.8, with gains concentrated in data analysis (11%) and due-diligence workflows (17%) [10]. Those numbers matter more than any single benchmark because they come from production workloads rather than curated test suites. They also land inside a broader pattern: Opus 5 is Anthropic's fourth Claude 5-generation release in under two months [11], a cadence that reads less like a singular blockbuster launch and more like continuous iteration on cost, speed, and reliability - exactly the axis enterprise buyers say they now care about most.

Historical Context

2026-05-28
Released Opus 4.8, the predecessor model whose pricing ($5/$25 per million tokens) Opus 5 matches exactly.
2026-07-24
Launched Claude Opus 5 as its fourth Claude 5-generation model release in under two months, positioned as near-Fable-5 intelligence at roughly half the cost per task.

Power Map

Key Players
Subject

Claude Opus 5 launch: benchmarks and pricing

AN

Anthropic

Model developer and publisher; released Opus 5 simultaneously on Claude Max, Pro, and API, positioning it as a cost-efficient near-frontier alternative to its own top-tier Fable 5 model.

CO

Cognition (Devin)

Coding-agent vendor; CEO Scott Wu said Opus 5 approaches Fable-level performance at half the cost inside Devin, with particular strength on debugging and root-cause analysis.

ZA

Zapier

Automation platform; CEO Wade Foster reported Opus 5 topped its AutomationBench leaderboard without spending more tokens than prior Claude releases.

BO

Box

Enterprise content-management customer; integrated Opus 5 and reported it outperforms Opus 4.8 by 8% overall, with larger gains in data analysis and due-diligence workflows.

CO

CodeRabbit

AI code-review vendor; independently benchmarked Opus 5, finding higher precision on actionable review comments but more missed known issues and far more low-value nitpicks than the prior model.

EV

Every

Independent reviewer; found the model 'brilliant in flashes, frustrating in practice' during its first week of real use, citing friction with existing skills and plugins.

Fact Check

11 cited
  1. [1] Claude Opus 5
  2. [2] Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing
  3. [3] Anthropic Launches Claude Opus 5, Tops AI Benchmark Index at Half the Cost of Fable 5
  4. [4] Claude Opus 5: The System Card
  5. [5] Anthropic launches Opus 5
  6. [6] Opus 5 Model Review
  7. [7] Vibe Check: Opus 5
  8. [8] Claude Opus 5
  9. [9] Claude Opus 5 Release Details
  10. [10] Anthropic Launches Claude Opus 5 and Puts a Dial on Your AI Bill: Here's What It Means for B2B Teams
  11. [11] Anthropic releases new model Opus 5

Source Articles

Top 3

THE SIGNAL.

Analysts

"Said Opus 5 approaches Fable-level performance at half the cost inside Devin, especially strong at debugging and root-cause analysis."

Scott Wu
CEO, Cognition

"Reported Opus 5 topped Zapier's internal AutomationBench leaderboard without using more tokens than earlier, identically-priced Claude releases."

Wade Foster
CEO, Zapier

"Analyzed the Opus 5 system card in depth, highlighting the drop in safety-classifier triggers on FrontierBench."

Zvi Mowshowitz
Independent AI safety analyst / newsletter author
The Crowd

"Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price."

@@claudeai60799

"Just over 6 months later, Opus 5 now produces near-superhuman level spreadsheets and slide decks that match what a consultant would make. Things are changing fast."

@@alexalbert__3610

"claude opus 5 vs fable 5 vs gpt 5.6 sol vs kimi k3 @AnthropicAI released claude opus 5 today. per the announcement, it's "a thoughtful and proactive model that comes close to the frontier intelligence of claude fable 5 at half the price." key facts from the release: • priced [thread continues, truncated by X's "Show more"]"

@@thehypedotnews62

"Introducing Claude Opus 5"

@u/ClaudeOfficial2900
Broadcast
Claude Opus 5 Is INSANE – Is This the BEST Model Yet?

Claude Opus 5 Is INSANE – Is This the BEST Model Yet?

Opus 5 Is Here: It Beats Fable

Opus 5 Is Here: It Beats Fable

We Tested Claude Opus 5. It's Frustrating with Flashes of Brilliance.

We Tested Claude Opus 5. It's Frustrating with Flashes of Brilliance.