Skip to content
← All models
Inception

Mercury 2 benchmarks & scores

Inception Proprietary Released Mar 2026
20.4
Vibe Coding Index · #66 of 103

Independent Mercury 2 AI coding benchmark results Intelligence 21, Coding 28, Agentic 10, with a Vibe Coding Index of 20.4. Compare its overall rank, token price, context window, and practical agentic workflow strengths below.

Blended price #34/108
$0.38
$/M tokens · 3 to 1 in, out
Context
128K
Window size

Evidence confidence

High confidence

100% core coverage

Complete core scores supported by at least two current verified sources.

3/3
Core scores
2
Evidence sources
Jul 21
Last measured

Catalog capabilities

What the model supports

128K context
Include reasoning Max tokens Reasoning Reasoning effort Response format Stop Structured outputs Temperature Tool choice Tools
Input
Text
Output
Text
Knowledge cutoff
Not listed
Cache read price
$0.03 per million tokens

Suite radar

Intelligence Coding Agentic

Suite scores

Intelligence #65/103
General reasoning & knowledge (Intelligence Index)
21.4
Coding #71/120
Code generation & software tasks (Coding Index)
28.4
Agentic #63/104
Multi-step tool use & agentic workflows (Agentic Index)
9.6

Task placements

Where Mercury 2 places outside the VCI blend

Head-to-head preference boards that never rewrite the suite scores above — Arena Code WebDev and Design Arena categories only.

Arena Code WebDev

Human preference for web development outputs. Separate from Artificial Analysis coding until blended into the published Coding suite score.

Coding suite score above (28.4, multi-source index) includes this signal when current.

Rank
#45 /45
ELO
1,164 ±23
Source #
96
Votes
947

Design Arena categories

Task leaderboards for Mercury 2

Website, UI, game, data viz, and other tournament boards — not used in the Vibe Coding Index formula.

Category Rank ELO Win rate Avg time Tournaments
ASCII art
#65 / 66 1,025 27.0% 10.4s 4,036 Source ↗
SVG
#84 / 90 1,026 25.4% 2.4s 4,036 Source ↗
3D
#110 / 121 1,049 24.6% 3.4s 4,036 Full board →
Game development
#111 / 129 1,036 21.3% 3.8s 4,036 Full board →
UI components
#111 / 126 1,016 20.2% 3.9s 4,036 Full board →
Data visualization
#114 / 128 1,009 20.9% 3.1s 4,036 Full board →
Code categories
#118 / 132 1,028 21.3% 4.3s 4,036 Full board →
Web development
#126 / 141 1,016 20.1% 5.1s 4,036 Full board →

Source record

Source. Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings).

TheVibeFather rankings

Mercury 2 practical agentic coding scorecard

This is our algorithmic assessment of where Mercury 2 ranks for common vibe coding and agentic engineering workflows. Every stable 0–10 score is calculated consistently from the sourced Intelligence, Coding, and Agentic signals above—never from an unsourced opinion score.

Overall VibeFather rating
2.0 /10
#66 of 103 models
Practical area

Small, well-defined code changes

Precision on scoped edits, fixes, and implementation tasks.

Developing 2.6/10
Field rank #67 / 103

UI and CSS iteration

Front-end implementation with iterative tool-driven refinement.

Developing 2.5/10
Field rank #67 / 104

Routine debugging

Diagnosing failures and turning reasoning into correct code changes.

Developing 2.5/10
Field rank #67 / 103

Repository-wide refactors

Coordinating larger edits across files while preserving intent.

Developing 2.1/10
Field rank #66 / 103

Architecture decisions

Reasoning through tradeoffs, constraints, and system-level choices.

Developing 1.7/10
Field rank #65 / 103

Tool use and workflow execution

Planning and completing multi-step work with tools and feedback loops.

Developing 1.4/10
Field rank #65 / 104

Long autonomous coding tasks

Sustaining coherent progress across longer agentic engineering runs.

Developing 1.2/10
Field rank #63 / 103

Overall VibeFather rating

Our complete Vibe Coding Index, expressed on the same 0–10 practical scale.

Developing 2.0/10
Field rank #66 / 103
1 · Sourced signals

We start with Intelligence, Coding, Agentic benchmark evidence, and preserve missing values as unverified.

2 · Practical lenses

Our code-controlled algorithm blends the native 0–100 signals by workflow and expresses the result on a stable 0–10 scale. Every required input must be present.

3 · Field ranking

Each workflow score is ranked against every model with comparable evidence, so positions update when the benchmark field changes.

The Vibe Coding Index is TheVibeFather’s code-controlled 0–100 composite for practical coding and agentic capability. Price and adoption are reported independently and never alter quality. Exact weighting remains proprietary, inputs, missing-data behavior, and per-category ranks are disclosed here. Read the benchmark methodology →

Recorded score history

How Mercury 2 has moved

All benchmark history →
24.0 21.0 18.0
Jul 21, 2026 4 AM CT
20.4 rank 66

No material score or rank change.

Jul 21, 2026 3 AM CT
20.4 rank 66

No material score or rank change.

Jul 21, 2026 2 AM CT
20.4 rank 66

No material score or rank change.

Where Mercury 2 sits

Its logo stays full-color and highlighted while the rest of the field recedes for quick comparison.

Efficient frontier Alibaba Amazon Anthropic Cohere DeepSeek Google Inception Inclusionai Kwaipilot Meta Meta Llama MiniMax Mistral Moonshot NVIDIA Nex Agi OpenAI Stepfun Tencent Thinkingmachines Upstage Xiaomi Zhipu xAI Hover, tap or focus a model for axis guides · resting dashed lines = medians 27 model(s) hidden — data not yet verified

Model benchmark FAQ

Mercury 2 for vibe coding and agentic engineering

Answers update from this model’s current sourced scores, practical rankings, and verified pricing.

What is Mercury 2's Vibe Coding Index?

Mercury 2 has a Vibe Coding Index of 20.4/100 and is #66 of 103 in the current field. The index is TheVibeFather's stable composite for practical AI coding, combining independently sourced Intelligence, Coding, and Agentic benchmark signals. Price and adoption are compared separately and never inflate the quality score.

How does Mercury 2 rank for vibe coding?

For vibe coding, Mercury 2's strongest practical area is Small, well-defined code changes at 2.6/10, ranking #67 of 103 models in TheVibeFather rankings.

Is Mercury 2 good for agentic coding and autonomous tasks?

Mercury 2 scores 1.4/10 for tool use and workflow execution and 1.2/10 for long autonomous coding tasks. Those categories emphasize the Agentic benchmark signal used for multi-step planning, tool calls, edits, tests, and feedback loops.

What coding benchmark scores does Mercury 2 have?

The current sourced benchmark profile for Mercury 2 is Intelligence 21/100, Coding 28/100, Agentic 10/100. TheVibeFather converts those 0–100 signals into consistent practical 0–10 workflow scores so models can be compared for real agentic engineering work.

How much does Mercury 2 cost for agentic engineering?

The current benchmark uses a 3 to 1 input-to-output blend of $0.38 per million blended tokens. That places it #34 of 108 by price, where a lower rank number means less expensive.

Are vibe coding, agentic coding, and agentic engineering the same?

They describe overlapping AI-assisted software workflows. Vibe coding emphasizes directing software through intent and iteration, agentic coding emphasizes a model planning and executing multi-step work with tools, agentic engineering is the broader discipline of designing, supervising, and validating those workflows. TheVibeFather benchmarks the shared capabilities behind all three terms on every model page, including Mercury 2.