Claude Opus 5 drops with major agent upgrades
· 2 hours ago
Claude Opus 5 drops with major agent upgrades. Claude Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work. I'm excited to see how this handles complex multi-step tasks 🚀
More from VibeWire
Claude Opus 5 enters the top ten on July 24 2026
Claude Opus 5 entered the top ten at rank 4 with a Vibe Coding Index of 66.6.
Claude Opus 5 has entered the verified tier of the coding benchmark 🚀
The refresh uses Artificial Analysis capability data through OpenRouter and Arena Code WebDev results captured July 24 2026 and July 23 2026. Preliminary source results stay marked tentative.
See the live AI coding benchmark leaderboard and read the scoring method.
Opus 5 Enters the Vibe Coding Index at 66.6 — With All Three Inputs Measured
Board updated. Claude Opus 5 enters the Vibe Coding Index at 66.6, and — unusually for a launch-day entry — it arrives with a complete profile rather than a tentative one.
The three inputs, straight off the Artificial Analysis capability boards for the max-effort configuration
- Intelligence 61 — narrowly the top published score
- Coding 78 — joint first with GPT-5.6 Sol (xhigh)
- Agentic 55 — top published score
Run those through the fixed capability blend (20% intelligence, 45% coding, 35% agentic) and you get 66.6. No estimation, no editorial synthesis, no "provisional composite" of the kind we had to use when Kimi K3 launched with only part of its profile public. Every dimension links to the board it came from.
Two things will move that number over the next few weeks, and neither is a correction
- The OpenRouter feed will list the model. When it does, our live sync takes ownership of the scorecard and deletes the hand-mirrored figures. That is the system working as designed.
- Arena Code WebDev votes will accumulate. Our coding dimension blends human-preference Elo at a fixed 35% weight once fresh votes exist, so expect a point or two of movement either way.
What I am deliberately not doing today is publishing the vendor's launch suite as if it were verified. SWE-bench Verified is quoted at 96.0% in one tracker and 97.0% in another — that spread usually means different trial counts or different scaffolds, and it is exactly the kind of one-point difference nobody should be making decisions on.
Scorecard Claude Opus 5 · Routing guide Opus 5 vs Fable 5 vs GPT-5.6 Sol
Claude Opus 5 Is Here — Frontier Scores at Half the Fable 5 Price
Opus 5 dropped. Anthropic shipped claude-opus-5 today at $5 in / $25 out per million tokens — the same price as Opus 4.8, and half of Fable 5's $10/$50.
That pricing line is the story. This is not a premium tier you have to justify to finance. If you are already on Opus 4.8, the upgrade is free at list price. If you are on Fable 5, you are being offered equal-or-better measured capability for half the rate.
What shipped
- 1M token context window, available on the Claude API, the Claude apps, Claude Code, and Claude Cowork
- Five effort levels — low, medium, high, xhigh, max — that decide how much thinking (and how many tokens) a task consumes
- Fast mode at roughly 2.5x the default speed for 2x the base rate
On the independent Artificial Analysis boards it takes #1 on Intelligence (61), #1 on Agentic (55), and ties for #1 on Coding (78) with GPT-5.6 Sol.
It is not a clean sweep, and I would not trust anyone who tells you it is. GPT-5.6 Sol still wins DeepSWE v1.1, 72.7% to 68.8%. Mythos 5 is still ahead on legal work and on exploit development. Fable 5 still reportedly leads on health tasks.
One number worth internalising before you change any defaults, at max effort, Artificial Analysis measures about 62.7 seconds to first token. That is completely fine for a background agent chewing through a ticket queue. It will feel broken in a chat box. Max is the benchmark setting, not the production setting.
Full breakdown every Opus 5 benchmark explained.
Primary source Anthropic — Introducing Claude Opus 5