The Grok 4.6 benchmark is no longer a pending page. After the August 12 launch, the August 13 Vibe Bench refresh pulled a complete Artificial Analysis profile. Grok 4.6 scores 67.3 on the Vibe Coding Index and sits sixth, just behind GPT-5.6 Sol at 67.5.
That is a measured placement, not a launch-day estimate. All three required inputs are present. Confidence is medium because the live feed currently cites one independent source family rather than a second confirming board.
The numbers on the live card
| Signal | Grok 4.6 | Grok 4.5 |
|---|---|---|
| Vibe Coding Index | 67.3 | 61.9 |
| Intelligence | 60.9 | 55.8 |
| Coding | 76.8 | 74.7 |
| Agentic | 58.7 | 48.9 |
| Board rank | 6 | 9 |
Artificial Analysis published its own write-up the same day. It scores Grok 4.6 (high) at 61 on the Intelligence Index, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62). Our live Intelligence cell is 60.9, which matches that board after ordinary rounding. AA also reported a GDPval-AA v2 Elo of 1753 and 88.4% on Terminal-Bench v2.1. Those are useful agentic details. They are not extra Vibe Coding Index inputs.
How to read the gap to the top
Claude Opus 5 still leads the live leaderboard at 71.7. Fable 5, Kimi K3, and Qwen3.8 Max sit in the high 68s. Grok 4.6 is in the next cluster with GPT-5.6 Sol. The interesting part is the agentic cell. Grok 4.6 gained about ten points on Grok 4.5 there, which is the largest of the three moves, and it did so at the same $2 / $6 headline price.
Vendor launch charts still exist and still need labels. SpaceXAI’s own table mixes first-party runs with third-party figures. We keep those on the launch briefing and let the independent 0–100 boards drive rank. See how we rank models if you want the method in one place.
What a sixth-place score should change
It should put Grok 4.6 on a short trial list for agentic coding and knowledge-work loops, especially if you already pay SpaceXAI prices. It should not automatically displace Opus 5, Fable 5, or Kimi K3. Those models still win the composite. Grok 4.6 wins the “same intelligence class as Sol, far cheaper output tokens” argument, which matters when a long agent loop is mostly output.
Bottom line
The Grok 4.6 benchmark now has a real score. Use 67.3 and rank 6 as the current Vibe Bench reading, check the model page after the next refresh, and keep vendor charts in a separate column. A measured sixth is more useful than an unconfirmed first.