Skip to content

The old Pro scored only 8 points, while the DeepSeek V4 Flash surged to 54.4 points in DeepSWE, approaching the Opus 4.8.

Jul 31, 14:37

According to Dynamic Insight Beating monitoring, on the DeepSWE long-range software engineering benchmark, DeepSeek-V4-Flash Official Version scored 54.4 points, approaching Claude Opus 4.8 with 59 points, far exceeding the previous 8 points of V4-Pro-Preview.

This score is lower than Kimi K3's 69 points but higher than GLM-5.2's 44 points, placing it in the top tier of domestic models.

However, this is not a raw model test under exactly the same conditions. The official V4-Flash version uses DeepSeek's as-yet-unreleased Harness minimalist mode, where the Agent's performance is affected by both the model and execution framework.

Source