Grok 4.6 vs Grok 4.5 is the cleanest same-lab upgrade on this week’s board. Headline price does not move. Context stays at 500,000 tokens. The live Vibe Coding Index rises from 61.9 to 67.3, and almost all of that extra quality sits in the agentic column.
SpaceXAI says the new run used a longer supplemental training pass, regenerated supervised traces with Grok 4.5, and then trained on agentic reinforcement-learning tasks across knowledge work, general coding, kernels, web development, and computer-aided design. That is a training story. The scorecard is how we decide whether to switch a seat.
Side by side on August 13
| Field | Grok 4.5 | Grok 4.6 |
|---|---|---|
| Vibe Coding Index | 61.9 | 67.3 |
| Intelligence | 55.8 | 60.9 |
| Coding | 74.7 | 76.8 |
| Agentic | 48.9 | 58.7 |
| Input / output | $2 / $6 | $2 / $6 |
| Cached input under 200k | $0.30 | $0.50 |
| Board rank | 9 | 6 |
Coding barely moved. Intelligence moved a useful five points. Agentic moved about ten. If your workload is single-shot code generation, the upgrade is modest. If your workload is a long tool loop that has to stay on a task, the upgrade is the whole point of the release.
When to switch, and when to wait
Switch a trial seat if you already use Grok 4.5 inside Grok Build or Cursor and the jobs fail by stalling, skipping verification, or losing the thread. Keep 4.5 as the default if your traffic is short completions, if cached-input volume is huge, or if you have not yet run your own repository suite. The cache line got more expensive. High-volume prompt reuse can erase the quality gain.
Neither model is the board leader. Opus 5, Fable 5, and Kimi K3 still sit above both. The comparison here is family-internal. Use the Grok 4.6 scorecard and the Grok 4.5 scorecard when you need the live cells, not a screenshot from launch day.
A fair trial is a dozen repository tasks you already scored on 4.5, plus one long agent loop that has to use tools. If 4.6 only wins the long loop, route that class of job and leave short completions on 4.5 until the cache bill is clear. If it wins both, switch the default and keep 4.5 as rollback for a week.
Bottom line
Treat Grok 4.6 vs Grok 4.5 as a same-price agentic step, not a new cost class and not a reason to abandon every other frontier model. Trial it on the jobs 4.5 already almost finished. Keep 4.5 until those jobs actually close.