Bug Bounty Round 3 Matches Opus5 in Recall Rate with Half the Cost
According to the DeepInspect Beating Monitor, the security firm Aikido conducted code audits using the Agent testing tool on 7 models, including Qwen3.8-Max, Claude Opus 5, Kimi K3 Max, DeepSeek V4 Flash, as well as GPT-5.6 Sol, Luna, and Terra. The test covered 32 recent publicly known vulnerabilities, with each model being run 3 times.
Qwen3.8-Max in total found 26 vulnerabilities over the three runs, achieving a recall rate of 81.3%, tied for first place with Opus 5. Its F1 composite score (balancing both missed detections and false alarms) was 83.2%, slightly below Opus 5, Kimi K3 Max, and GPT-5.6 Sol.
However, Qwen's performance was inconsistent in individual runs. Out of the 26 vulnerabilities, only 10 were consistently identified in all three runs, while Opus 5 and Sol had 19. Nevertheless, Qwen's total cost was approximately $821, only half of Opus 5 and Sol, but still five times that of DeepSeek V4 Flash.