Claude Opus 5 launched on July 24, 2026 as Anthropic’s new flagship in the Opus tier, succeeding Opus 4.8 at the exact same price point of 5 USD per million input tokens and 25 USD per million output tokens. Anthropic describes it as a “thoughtful and proactive” model that comes close to the frontier intelligence of Claude Fable 5, but at roughly half the cost, explicitly pitching it as an everyday workhorse for coding, complex office work and long-running agents. Opus 5 is available across all Anthropic surfaces, has become the default model on Claude Max, and is the strongest model accessible on Claude Pro, which means frontier-adjacent performance is now in front of ordinary subscribers rather than just high-end enterprise contracts.
Where Opus 5 Fits in the Claude Lineup
As of mid‑2026, Anthropic’s lineup is relatively straightforward: Fable 5 for the most demanding long‑running agents, Sonnet 5 for the best cost–speed–intelligence balance, Haiku 4.5 for high‑volume, low‑latency classification, and the Opus tier for deep reasoning and agentic workflows, now led by Opus 5.
Fable 5 is priced at around 10 USD per million input tokens and 50 USD per million output tokens, while Sonnet 5 sits lower at 3 USD / 15 USD (with an introductory 2 USD / 10 USD offer until late August 2026) and Haiku 4.5 starts at 1 USD / 5 USD. In that context, Opus 5 occupies the “sweet spot” of near‑Fable capabilities at Opus pricing, making it a compelling choice for agencies like CreativDigital and software teams that want powerful agents without paying Fable rates for every request.
Benchmarks That Actually Matter
Opus 5’s launch is heavily backed by benchmark numbers, and they are unusually strong across the board.
On Frontier‑Bench v0.1, a hard set of real computer tasks including coding, system administration and data work that an AI agent has to complete autonomously, Opus 5 reaches 43.3% solved, more than double Opus 4.8’s 21.1% and nearly 10 points clear of Fable 5 for the best published score at launch.
On ARC‑AGI‑3, a novel problem‑solving test where the model is dropped into interactive environments it has never seen, Opus 5 scores 30.2%, around three times the next‑best model (GPT‑5.6 Sol at 7.8%) and roughly twenty times Opus 4.8’s 1.5%, which Anthropic points to as evidence of genuine reasoning gains rather than overfitting benchmarks.
Knowledge Work and Business Automation
For economically valuable knowledge work and business workflows, Opus 5 currently tops several independent evaluations. On GDPval‑AA v2, which measures economically valuable knowledge work in a rebased Elo format, Opus 5 comes in at 1861, ahead of Fable 5 at 1747 and GPT‑5.6 Sol at 1736, claiming the best published score at the time of launch.
On Zapier’s AutomationBench, a test of end‑to‑end business workflows, Opus 5 completes 26% of tasks, about 1.5x the next‑best model for the same cost per task, and even at its lowest effort setting it passes more tasks than any competitor.
For a digital agency or product shop, that translates into agents that can:
- Set up and optimize marketing campaigns
- Generate detailed performance audit reports
- Create technical documents and customized email workflows
- Coordinate complex end-to-end multi-app processes without manual steps
Because AutomationBench mimics real multi‑step office processes, Opus 5’s lead there is a strong signal that it is well‑suited to automating marketing operations, CRM workflows and internal reporting pipelines.
Agentic Software Engineering and Coding
Opus 5 is explicitly tuned for agentic software engineering: planning tasks, editing files, running tests and iterating on its own rather than just spitting out static code snippets.
On FrontierCode v1.1, a benchmark of very hard, frontier‑difficulty coding tasks that an agent has to carry out end‑to‑end, Opus 5 scores 53.4%, and on DeepSWE 1.1 it reaches 68.8%, placing it near the top of tracked models (fourth out of thirteen, with GPT‑5.6 Sol leading at 73%). Combined with a 1‑million‑token context window, Opus 5 can analyze large codebases, understand complex architectures and propose refactors or new features without forcing you to break prompts into tiny slices.
If you’re a developer building modern web applications, Opus 5 is the kind of model that can generate not just isolated snippets but entire modules, CI scripts, data pipelines or agents that run tests and self‑repair failing cases. Its tunable effort ladder – from low through medium, high, xhigh up to max – lets you decide when to pay for deep reasoning and multi‑step planning versus when a quick, cheap answer is all you need, which is crucial when running agents in production.
Computer Use, Agents and Web Search
One of the big promises of the Claude 5 generation is robust computer use: models that can operate real desktops, browsers and apps, not just chat about them.
On OSWorld 2.0, a refreshed and harder computer‑use benchmark, Opus 5 achieves 70.6%, beating Fable 5 (66.1%), GPT‑5.6 Sol (62.6%) and Opus 4.8 (55.7%) while reportedly delivering Fable‑level performance at around a third of Fable’s cost per task. On BrowseComp, a test of agentic web search – can the model browse the web and track down hard‑to‑find answers? – Opus 5 hits 90.8%, just behind Kimi K3 at 91.2% and slightly ahead of GPT‑5.6 Sol at 90.4%.
For a team like CreativDigital, that opens the door to agents that actually click through dashboards, fill in forms, scrape competitor data and assemble research reports by operating standard web tools.
Combined with the large context window, you can run agents that open dozens of tabs, work inside CRM systems, spreadsheets and project‑management tools, with a single Opus 5‑powered brain coordinating the whole workflow.
Pricing, Access and the Fast Mode
From a business perspective, the most striking aspect of Opus 5 is that it keeps Opus 4.8’s prices unchanged: 5 USD per million input tokens (around 3.95 GBP) and 25 USD per million output tokens (around 19.75 GBP). At that price point, Anthropic emphasizes token efficiency – Opus 5 tends to reach answers with fewer wasted steps and fewer tool calls than its predecessors, which lowers effective cost per task even before you factor in the better benchmark scores.
A dedicated Fast mode runs Opus 5 at roughly 2.5x the default speed at twice the base price, for situations where latency matters more than strict token cost, and the launch also introduces mid‑conversation tool changes and automatic API fallbacks so that safety‑flagged requests can be routed to other models instead of being hard‑blocked.
Safety, Alignment and Known Limitations
Anthropic describes Opus 5 as its most aligned model to date, with an automated behavioral audit score of 2.3 versus 2.85 for Opus 4.8, and reports that it has the lowest cooperation rate with misuse among the models it tested. External cyber‑range evaluations by the UK AI Security Institute on early Opus 5 checkpoints indicated that Opus 5 is less capable at exploiting vulnerabilities than Mythos 5, which allows Anthropic to give it somewhat looser safeguards for general access while still keeping offensive capabilities constrained relative to their specialist cyber model.
At the same time, the system card notes that Opus 5 “hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall,” and that it deliberately lags Mythos 5 on advanced biology and offensive cybersecurity work, underscoring that it is meant as a safe general‑purpose flagship rather than a cutting‑edge research tool.
Positioning Versus Fable, Sonnet and Haiku
- Claude Fable 5 remains Anthropic’s frontier‑class offering for long‑running agents, with adaptive thinking permanently enabled and a design optimized for multi‑day autonomous projects (~$10 / ~$50).
- Claude Sonnet 5 is pitched as the best balance between speed and intelligence, and several ecosystem guides recommend keeping most production traffic on Sonnet 5 for cost‑efficiency (~$3 / ~$15).
- Claude Haiku 4.5 is tuned for high‑volume workloads like classification and routing, with a 200K‑token context window, minimal latency and the lowest price point in the lineup (~$1 / ~$5).
- Claude Opus 5 serves as the bridge between Sonnet and Fable: near‑frontier scores on coding, knowledge work and computer use, at mid‑tier prices (~$5 / ~$25), making it the natural choice for powerful production agents in medium and large organizations.
Practical Scenarios for Agencies and Dev Teams
For a digital agency or product studio, Opus 5 unlocks a range of practical scenarios:
- Team coding assistant: Agents that read entire repos, propose refactors, generate tests and even apply patches, leveraging Opus 5’s strong FrontierCode and DeepSWE scores and large context.
- Campaign and operations automation: Agents that set up ads, validate CRM data, coordinate email flows and produce weekly reports, built on the strengths demonstrated in GDPval and AutomationBench.
- Semi‑autonomous competitive research: Agents that browse the web, compile pricing and feature comparisons and synthesize findings into strategic briefs using BrowseComp‑level browsing performance.
- Internal tool orchestration: With mid‑conversation tool changes and automatic fallbacks, you can design flows where Opus 5 chooses which tools (scrapers, databases, email senders) to call and gracefully reroutes when a request trips safety classifiers.
If you’re looking to ship maximum value by deploying autonomous AI agents, Opus 5 is arguably the best current balance between agent capabilities and monthly API cost. Contact the CreativDigital team to integrate Claude Opus 5 into your production stack.



