Skip to content

DeepSeek-V4-Pro Official Version Agent Ability Surges, DeepSWE Spikes Nearly 50 Points in One Go

Aug 13, 11:03

According to Sentinel Beating monitoring, the Agent score of DeepSeek-V4-Pro-0813 has seen a significant leap. A self-test table released by DeepSeek's official team shows that DeepSWE soared from 12.8 in the Preview version to 62.7, marking a 49.9-point increase. CyberGym also rose from 52.7 to 83.3, while AutomationBench climbed from 12.8 to 31.8.

The new version has outperformed Claude Opus 4.8 in multiple benchmarks. Terminal Bench 2.1 achieved 87.9 compared to 85.0, CyberGym scored 83.3 versus 78.3, and DeepSWE reached 62.7 against 58.0. AutomationBench even surpassed Fable 5 with a score of 31.8 over 29.1.

Remarkably, despite the upgrade from Preview to 0813, there has been no price increase. The V4-Pro API still maintains a cost of 3 USD per million tokens for input and 6 USD for output.

However, these results are currently based on DeepSeek's self-testing, and third parties have not yet replicated them entirely. Especially considering DeepSWE's nearly 50-point surge in one go, and the fact that Agent evaluation heavily relies on Harness, the actual improvement will require external testing.

Source