Claude Opus 5 launched on July 24, 2026 at $5 in / $25 out per million tokens and immediately took the top published position on two of the three index boards we score. This page is the hub, every Opus 5 matchup in one place, with a link to the live head-to-head for each one.
A note on method before the tables. The live comparison pages recompute from current data on every request, so they will always be more accurate than any prose written on launch day. Use this page to decide which matchup you care about, then use the live page for the numbers you act on. Where a competitor's index has not been published, we say so instead of estimating.
Deeper reading every Opus 5 benchmark explained, the three-way routing decision, and the pricing and effort-level guide.
The one-screen summary
Published Intelligence Index results as of July 24, 2026. Opus 5 leads by a single point — a narrow margin worth reading honestly.
Claude Opus 5's own numbers, for reference in every matchup below
| Metric | Claude Opus 5 |
|---|---|
| Vibe Coding Index | 66.6 |
| Intelligence Index | 61 — top published |
| Coding Index | 78 — joint top published |
| Agentic Index | 55 — top published |
| Price | $5 in / $25 out per 1M |
| Context window | 1M tokens |
| Model ID | claude-opus-5 |
Every matchup at a glance
| Matchup | Short answer | Live head-to-head |
|---|---|---|
| vs GPT-5.6 Sol | Tied on coding, Opus 5 leads agentic, Sol wins DeepSWE | Compare → |
| vs Claude Fable 5 | Opus 5 edges intelligence at half the price | Compare → |
| vs Claude Opus 4.8 | Same price, large capability jump | Compare → |
| vs Claude Sonnet 5 | Different lanes — cost and latency vs peak capability | Compare → |
| vs Kimi K3 | Opus 5 leads the index, Kimi K3 is far cheaper | Compare → |
| vs Grok 4.5 | See the live board — Grok's published coverage is thinner | Compare → |
| vs MiniMax-M3 | Frontier capability against open-weight economics | Compare → |
Claude Opus 5 vs GPT-5.6 Sol
This is the matchup that decides most defaults, and it is genuinely close. Both models post a Coding Index of 78 — Opus 5 at max effort, GPT-5.6 Sol at xhigh. Opus 5 leads the Agentic Index 55 to 54 and the Intelligence Index 61 to 59.
| Benchmark | Opus 5 | GPT-5.6 Sol | Winner |
|---|---|---|---|
| Coding Index | 78 | 78 (xhigh) | Tie |
| Agentic Index | 55 | 54 | Opus 5 |
| Intelligence Index | 61 | 59 | Opus 5 |
| SWE-bench Pro | 79.2% | 64.6% | Opus 5 |
| DeepSWE v1.1 | 68.8% | 72.7% | GPT-5.6 Sol |
| OSWorld 2.0 | 70.6% | 62.6% | Opus 5 |
| ARC-AGI-3 | 30.2% | 7.8% | Opus 5 |
| GDPval-AA v2 | 1,861 Elo | 1,736 Elo | Opus 5 |
The DeepSWE result is the one to weigh if your work is large multi-file changes in an established repository — that is the benchmark aimed squarely at repo-scale engineering, and GPT-5.6 Sol wins it. Everywhere else, especially on unfamiliar problems and long-horizon agent runs, Opus 5 leads. Keep both routes alive see the live comparison.
Claude Opus 5 vs Claude Fable 5
The internal matchup, and the one with the clearest financial answer. Opus 5 posts an Intelligence Index of 61 against Fable 5's 60, leads Frontier-Bench v0.1 43.3% to 33.7%, leads GDPval-AA v2 1,861 to 1,747, and leads AA-Briefcase by 146 Elo — at half the token price, $5/$25 against $10/$50.
| Metric | Opus 5 | Fable 5 |
|---|---|---|
| Intelligence Index | 61 | 60 |
| Frontier-Bench v0.1 | 43.3% | 33.7% |
| GDPval-AA v2 | 1,861 Elo | 1,747 Elo |
| Price per 1M | $5 / $25 | $10 / $50 |
| Health tasks | Behind | Leads |
Be careful with one thing. Fable 5's Coding and Agentic index values are not separately published, so a coding-specific Opus-5-beats-Fable-5 claim rests on vendor comparisons rather than a third-party board. The intelligence and price facts are solid. Live comparison →
Claude Opus 5 vs Claude Opus 4.8
The easiest decision on this page, because the price is identical at $5/$25. Opus 5 more than doubles Opus 4.8 on Frontier-Bench v0.1 — 43.3% against 18.7% — and the ARC-AGI-3 gap is 30.2% to 1.5%. Anthropic also reports domain gains of 10.2 percentage points on organic chemistry evaluations and 7.7 points on protein tasks.
If you are running Opus 4.8 today, there is no budget conversation to have. Run your own twenty tickets, confirm the improvement holds on your codebase, pin the revision and move. Live comparison →
Claude Opus 5 vs Claude Sonnet 5
This is not a "which is better" question — it is a lane question, and getting it wrong is the most common way teams overspend on AI coding.
| Consideration | Opus 5 | Sonnet 5 |
|---|---|---|
| Price per 1M | $5 / $25 | $2 / $10 introductory, then $3 / $15 |
| Latency at high effort | Slow — reasoning first | Much faster |
| Best lane | Hard debugging, long-horizon agents | Routine features, mechanical edits, interactive surfaces |
Sonnet 5's introductory rate runs through August 31, 2026. One caveat that catches people. Sonnet 5 uses a newer tokenizer that counts roughly 30% more tokens for the same text, so per-token parity is not per-request parity. Measure on your own prompts before you model the saving. The right architecture sends each task to the cheapest model that reliably completes it and escalates to Opus 5 on failure. Live comparison →
Claude Opus 5 vs Kimi K3
Kimi K3 held the top of our board before today at a Vibe Coding Index of 64.1, built from a tentative multi-source profile anchored by independent Artificial Analysis indices. Opus 5 enters at 66.6 with all three dimensions independently published.
The interesting axis here is not capability, it is economics. Kimi K3 lists at $3 in / $15 out per million tokens against Opus 5's $5/$25, with the same 1M context window. If your workload is high-volume implementation where the capability gap costs you little, the cheaper model can win on cost per accepted change even while losing the leaderboard. That is exactly the calculation the index cannot make for you. Live comparison →
Claude Opus 5 vs Grok 4.5 and MiniMax-M3
These two matchups are where we are most careful, because published index coverage for both is thinner than for the Anthropic and OpenAI flagships. Rather than extrapolate, the live comparison pages show exactly which dimensions have measured values and which render an em-dash.
That em-dash is deliberate design on our part. A missing score is information — it means nobody has published a comparable measurement — and filling it with an estimate would quietly turn an unknown into a ranking. See Opus 5 vs Grok 4.5 and Opus 5 vs MiniMax-M3 for current coverage.
How to read any of these comparisons
Every head-to-head page shows the same six rows. Vibe Coding Index, the three source dimensions, blended price, and context window. Here is what each is good for and where it misleads.
| Row | Good for | Misleads when |
|---|---|---|
| Vibe Coding Index | One-number ranking across capability | Your workload is not weighted like the blend |
| Intelligence | General reasoning headroom | Read as a coding score |
| Coding | Code generation and software tasks | Compared across different effort tiers |
| Agentic | Multi-step tool use | Your agent is single-shot |
| Blended price | Rough cost comparison | Models differ in verbosity or tokenizer |
| Context window | Ceiling on what fits | Treated as a target rather than a limit |
The row that trips up the most people is Coding. Index values are published per effort configuration, and comparing a max-effort number against another vendor's default is not a comparison, it is a category error. When the boards publish multiple tiers, we score the tier the board ranks at the top and say which one it is.
Which comparison should you actually read?
- Currently on Opus 4.8 read the Opus 4.8 matchup. Same price, large jump, easy decision.
- Currently on Fable 5 read the Fable 5 matchup. You are likely overpaying by 2x unless you are in health-adjacent work.
- Currently on GPT-5.6 Sol read the Sol matchup carefully, and weight DeepSWE by how repo-scale your work really is.
- Cost-constrained read the Sonnet 5 and Kimi K3 matchups. The frontier model is often not the right default for volume work.
- Evaluating from scratch start with the full leaderboard rather than any single pair.
The comparison nobody runs, and should
Every matchup on this page is model against model. The comparison that changes budgets more often is model against your current setup — same model, better harness.
A mid-tier model with cached repository context, summarised tool output, a failing test as the goal, and a three-attempt retry cap routinely outperforms a frontier model driven by a vague prompt with the whole repo pasted in. The second setup costs several times more and produces worse patches. We have watched teams switch models three times looking for a quality improvement that was sitting in their context strategy the whole time.
| Change | Typical quality effect | Typical cost effect |
|---|---|---|
| Upgrade to a frontier model | Moderate improvement | Higher per token |
| Add a verifiable definition of done | Large improvement | Lower — fewer retries |
| Cache stable context | Neutral | Large decrease |
| Summarise tool output | Often improves | Decrease |
| Cap retries and escalate to a human | Improves predictability | Decrease |
Run the harness comparison before the model comparison. If Opus 5 still wins afterwards — and on hard work it usually will — you will at least know what you actually bought.
The evidence standard behind these pages
Every number on a comparison page traces to a source you can open. The three index dimensions come from published Artificial Analysis capability boards. Where we mirror a board by hand on launch day — as we did for Opus 5 this morning — that mirror is recorded as its own evidence type and is deleted automatically when the live OpenRouter feed lists the model and takes ownership of the scorecard.
Vendor-reported launch numbers are stored separately from measured ones and never feed the index. That is why an Opus 5 comparison page shows a Vibe Coding Index of 66.6 rather than a number assembled from the launch slide deck. It is also why some cells are empty, we would rather show you a gap than a guess.
What changes when the live feed catches up
Comparison pages are only as complete as the roster behind them, and the roster is still settling after a launch. A few of these matchups are queued rather than live, the pair is configured, and the page activates automatically on the sync that introduces the opponent to the public catalog. Until then the URL returns a 404 rather than a half-empty page, and it stays out of the sitemap.
That is a deliberate choice. A comparison page with one populated column is worse than no page — it looks authoritative while showing you nothing, and it teaches readers to distrust the rest of the board. We would rather ship the pair configuration now and let the page appear the moment it has something to say.
| State | What you see | In the sitemap? |
|---|---|---|
| Both models scored | Full head-to-head with all six rows | Yes |
| One model unscored | Page renders, unmeasured cells show an em-dash | Yes |
| Opponent not in the roster | 404 — the pair is queued | No |
The same logic governs Opus 5's own scorecard. It entered with a hand-mirrored set of published index values recorded as their own evidence type. When the OpenRouter benchmark feed lists the model, the live sync takes ownership, deletes the mirror, and every comparison page that references Opus 5 starts reading from the feed instead. No page needs editing for that to happen.
Common questions
Is Claude Opus 5 the best coding model right now?
It ties GPT-5.6 Sol at 78 on the published Coding Index and leads the Agentic and Intelligence boards. "Best" depends on your workload — GPT-5.6 Sol still wins DeepSWE v1.1, which is the most repo-scale of the coding benchmarks.
Is Opus 5 worth it over Opus 4.8?
At identical pricing and roughly double the Frontier-Bench result, the evaluation is nearly free to run. Confirm on your own tickets, then switch.
Should I move off Fable 5?
Probably, unless you do health-adjacent work. Same or better published intelligence at half the token price is a hard combination to argue against.
What about cheaper models like Kimi K3 or Sonnet 5?
For high-volume implementation work they frequently win on cost per accepted change. Route by lane rather than picking one model for everything.
Do these comparison pages update automatically?
Yes. They recompute from current benchmark data on every request, so a live page always beats a number quoted in an article — including this one.
Sources and further reading
- Anthropic. Introducing Claude Opus 5
- Artificial Analysis Coding Index
- Artificial Analysis Agentic Index
- OpenRouter. Claude Opus 5 pricing and providers
- Claude Opus 5 scorecard
- How the Vibe Coding Index is calculated
A final word on comparison content generally. Most "X vs Y" pages on the internet are written once, never updated, and quietly become wrong within a month as prices change and new revisions ship. That is why the numbered claims here live on generated pages that recompute, and why this hub exists mainly to route you to them. If you find a figure on this page that disagrees with the live comparison, trust the live comparison — it read the database, and this paragraph did not.