7/26/2026, 1:01:57 PM · foundation-models

Claude Opus 5 Launches; Anthropic Retakes Agentic Coding Lead with 43.3% FrontierBench Score

Anthropic's new Claude Opus 5, released July 24, tops the FrontierBench v0.1 agentic coding leaderboard at 43.3%—surpassing both its own Fable 5 and OpenAI's GPT-5.6 Sol—while holding pricing flat at $5/$25 per million tokens.

Overview

<cite index="20-5">Anthropic released Claude Opus 5 on July 24, 2026, positioning it as a daily-use model that approaches the capability of its flagship Claude Fable 5 at roughly half the price.</cite> <cite index="5-1">The model is the fourth in Anthropic's Claude 5 family, which also includes Mythos 5 and Fable 5.</cite> The launch arrives fifteen days after <cite index="28-1">GPT-5.6 (Generative Pre-trained Transformer 5.6), a large language model (LLM) developed by OpenAI, reached general availability on July 9, 2026.</cite>

Benchmark Performance

<cite index="21-1,21-2,21-3">On FrontierBench v0.1, a 74-task successor to Terminal-Bench 2.1, Opus 5 scored 43.3% at max effort. Opus 4.8 had scored 18.7%. Fable 5 reached 33.7% and GPT-5.6 Sol reached 37.5%.</cite> <cite index="27-1,27-2">According to the Anthropic Claude Opus 5 system card, Table 8.1.A reports the 43.3% figure from an internal FrontierBench v0.1 run using mini-SWE-agent, a GKE backend, and five attempts per task.</cite>

<cite index="21-16">At xhigh effort, Opus 5 reaches 44.4% mean reward, its best result.</cite> Beyond coding, <cite index="21-22,21-23,21-24">the ARC Prize Foundation reports a verified 30.16% score for Opus 5 on ARC-AGI-3 at high effort—roughly four times the best previously reported leaderboard score; GPT-5.6 Sol reached 7.78% and Opus 4.8 reached 1.52%.</cite> <cite index="21-26">On Humanity's Last Exam, Opus 5 scored 56.3% without tools and 64.7% with them.</cite>

<cite index="2-3">Opus 5 is now Anthropic's most capable generally available model for scientific research, scoring 10.2 percentage points higher than Opus 4.8 on the company's internal chemistry benchmark.</cite>

Pricing and Availability

<cite index="2-6,2-7">Anthropic says the model delivers nearly all the intelligence of its top-of-the-line Claude Fable 5 at half the cost. The model is priced at $5 per million input tokens and $25 per million output tokens, unchanged from its predecessor, Opus 4.8.</cite> <cite index="3-11">In contrast, Fable 5 costs $10 per million input tokens and $50 per million output tokens.</cite>

<cite index="20-6,20-7">Anthropic announced the model on its official blog, stating it is available immediately through the Claude application programming interface (API) and consumer tiers. Opus 5 is now the default model on Claude Max and the strongest model available on Claude Pro.</cite>

<cite index="1-10">Key technical specifications include a 1 million token context window, an "xhigh" reasoning effort mode, per-turn controls, and a safety fallback to Claude Opus 4.8.</cite> <cite index="21-9,21-10">Context is 1M tokens as both default and maximum, with no smaller variant; maximum output is 128k tokens on the synchronous Messages API.</cite>

Competitive Context and Enterprise Signals

<cite index="28-9">OpenAI's flagship GPT-5.6 Sol had been described by OpenAI as its "workhorse" and "best coding model yet" suited for complex reasoning, coding, and agentic workflow.</cite> Opus 5's FrontierBench result now places Anthropic's generally available offering ahead of Sol on that specific evaluation.

<cite index="19-7">Early-access enterprise partners reported immediate cost benefits: AI legal provider Harvey reported that Opus 5 matched maximum-reasoning outputs while generating 26% fewer tokens on average, while Zapier Chief Executive Wade Foster noted that the model completed a complex churn-prevention sequence end-to-end on AutomationBench—a test previous models failed entirely.</cite>

<cite index="6-1">Opus 5 is Anthropic's fourth Claude 5 model release in less than two months, underscoring how AI deployment has shifted from blockbuster launches to rapid improvements on capability, cost, and speed.</cite>

<cite index="5-4">Mythos 5 remains Anthropic's current cutting-edge model, but the AI lab has limited its availability to a select group of companies and organizations due to its impressive hacking capabilities.</cite> <cite index="3-5">Opus 5 features less restrictive cybersecurity safeguards than Fable 5 and is not subject to the 30-day data retention policy.</cite>

Financial Backdrop

<cite index="2-5">Reuters reported in February that Anthropic was valued at roughly $380 billion in its latest funding round, following a period in which its annualized revenue climbed from about $1 billion at the end of 2024 to a projected $9 billion by the end of 2025, with internal targets reportedly reaching $20 to $26 billion for 2026.</cite>

Cross-references

Sources

  1. [1]
    What Is Claude Opus 5? Anthropic's Honeycomb Flagship
  2. [2]
    Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows | VentureBeat
  3. [3]
    Anthropic Launches Claude Opus 5, Matching Near-Flagship Performance for Half the Cost — BigGo Finance
  4. [4]
    Introducing Claude Sonnet 5 \ Anthropic
  5. [5]
    Anthropic debuts Opus 5 model as company preps for IPO later this year
  6. [6]
    Anthropic releases new model, Opus 5
  7. [7]
    Anthropic preparing for potential Claude Opus 5 rollout
  8. [8]
    Anthropic Prepares for Claude Opus 5 Rollout | ClaudeAINews
  9. [9]
    Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
  10. [10]
    Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge
  11. [11]
    Automated Benchmark Auditing for AI Agents and Large Language Models
  12. [12]
    Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
  13. [13]
    Introducing the AI Frontier Model Tracker | DemandSphere
  14. [14]
    Quantifying Frontier LLM Capabilities for Container Sandbox Escape
  15. [15]
    SciDesignBench: Benchmarking and Improving Language Models for Scientific Inverse Design
  16. [16]
    FrontierMath Leaderboard
  17. [17]
    SWE-bench Verified - AI Benchmark Explained | DemandSphere
  18. [18]
    AI benchmark scores — every frontier model on every major benchmark
  19. [19]
    Anthropic Launches Claude Opus 5: Smarter, Faster and Half the Price
  20. [20]
    Anthropic Releases Claude Opus 5, Claims State-of-the-Art Coding at Half the Cost of Fable 5
  21. [21]
    Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing - MarkTechPost
  22. [22]
    Claude Opus 5 Benchmarks Explained: Frontier-Bench, ARC-AGI-3 and SWE-bench (2026)
  23. [23]
    Claude Opus 5 Benchmarks Explained - Vellum
  24. [24]
    Anthropic Launches Claude Opus 5, Bringing Near-Fable AI Performance at Half the Price
  25. [25]
    Claude Opus 5: Benchmarks, System Card & Review (July 2026) - AIToolsReview
  26. [26]
    Claude Opus 5: Everything You Need to Know
  27. [27]
    Claude Opus 5 Benchmarks, Pricing & Speed (July 2026) | BenchLM.ai
  28. [28]
    GPT-5.6 - Wikipedia
  29. [29]
    ChatGPT 5.6: Release Date, Models, Pricing & Access
  30. [30]
    GPT-5.6 Sol: Benchmarks, Pricing & API Access Guide 2026
  31. [31]
    GPT-5.6 Sol: Benchmarks, API Pricing & Review | Coursiv Blog
  32. [32]
    GPT-5.6 Release Date: When You Can Actually Use It
  33. [33]
    GPT-5.5
  34. [34]
    GPT-5.4
  35. [35]
    GPT-5.6 Sol, Terra, Luna rolling out in ChatGPT, Codex, ...
  36. [36]
    GPT-5.6: Frontier intelligence that scales with your ambition | OpenAI
Claude Opus 5 Launches; Anthropic Retakes Agentic Coding Lead with 43.3% FrontierBench Score · AIDB