DeepSeek-V4-Pro Keeps Getting Stronger: Three-Wheel Recovers 87.5% of Bugs, Surpassing Claude Opus 5 and Qwen3.8 Max
According to CyberBeat monitoring, security firm Aikido tested 8 cutting-edge models against 32 real-world vulnerabilities. The DeepSeek-V4-Pro-0813 solo run only managed to uncover 58.3% on average, but after 3 consecutive runs followed by result merging, it directly identified 28 vulnerabilities, achieving an 87.5% coverage rate, pushing it to the top. Both Claude Opus 5 and Qwen3.8 Max scored only 81.3%.
Even more remarkable is the cost. DeepSeek spends an average of $7.34 per vulnerability found, while Opus 5 costs $66.97, making the former roughly 1/9 the cost of the latter.
DeepSeek's performance is also quite unusual: unstable in solo runs, but multiple consecutive runs yield the best results. In each round, it uncovers different vulnerabilities, and after three rounds, its coverage surpasses all competitors. However, this comes at the expense of a high false positive rate; only 65.6% of the issues it reports are ultimately confirmed as true vulnerabilities, the lowest among the 8 models.