Skip to content

Models

DeepSeek V4 Pro Agent Scores Are Vendor Numbers First

DeepSeek published a long agent table with the GA. DeepSeek V4 Pro agent scores from the lab are not the same thing as the live Vibe Coding Index.

The Vibe Father 6 min read
DeepSeek whale symbol and wordmark on black
DeepSeek product wordmark. Editorial reference TheVibeFather media library Editorial reference
Share Post to X LinkedIn

DeepSeek V4 Pro agent scores in the August 13 changelog are the lab’s own table. They are useful. They are not the Vibe Coding Index. The live August 13 refresh puts DeepSeek V4 Pro 0813 at 58.9, rank 16, with Intelligence 53, Coding 68.8, and Agentic 49.6.

That gap is normal on launch week. A vendor can report Terminal-Bench 2.1 at 87.9 in its harness while an independent 0–100 index lands much lower. Both can be honest. They are not interchangeable.

What DeepSeek reported for GA

Vendor figureReported scoreHow we treat it
HLE without / with tools42.7 / 60.0Launch claim, not a VCI cell
Terminal Bench 2.187.9Launch claim
NL2Repo61.5Launch claim
Cybergym83.3Launch claim
DeepSWE62.7Launch claim
Toolathlon-Verified74.1Launch claim
Agents' Last Exam25.7Launch claim
DSBench-FullStack / Hard71.1 / 67.2Internal DeepSeek sets

DeepSeek’s July 31 Flash notes already warned that some of these runs use a DeepSeek harness in minimal mode at max effort. Read every vendor row with that in mind. A 87.9 Terminal-Bench number in-house is not a promise that your agent loop will hit 87.9.

What the live index says

On Vibe Bench, V4 Pro 0813 is a strong mid-board open-weight model, not a new number one. It sits near DeepSeek V4 Flash 0731 (58.4) and well below Grok 4.6 (67.3), GPT-5.6 Sol (67.5), and Claude Opus 5 (71.7). The agentic cell at 49.6 is respectable and cheaper than the frontier closed models. It is not the 80-plus story the raw vendor table can suggest if you skim.

Use the live DeepSeek V4 Pro card for staffing. Keep the vendor table as a “what the lab wants you to try” list. The method is in how we rank models.

If you already run V4 Pro from the April preview, do not assume the GA scores transfer. Re-run the same agent suite at the same thinking effort. If Terminal-Bench in your harness does not move, the changelog table is not your production result. If agent tasks do move, keep the win and still quote 58.9 when someone asks where the model ranks.

Bottom line

DeepSeek V4 Pro agent scores from the lab are a reason to re-test agents, not a reason to reprint the leaderboard. Quote the changelog as DeepSeek’s numbers. Quote 58.9 and rank 16 as the current independent placement.

Sources

Reader check

Was this article helpful?

One click helps us decide what to research next.

The app behind this research

TheVibeFather is the multi-CLI AI coding harness

You just read field notes from the same team that ships TheVibeFather — the multi-CLI AI coding harness that runs Claude Code, Codex, OpenCode and more with shared memory and a verify gate. Bring your own keys.

Keep reading