DeepSeek V4 Flash to Replace 8 Harness Sets: Pi Agent Offers Highest Success Rate and Cost Savings, While Claude Code is the Fastest but Most Expensive
According to the Dynamic Watch Beating Monitor, AI Agent infrastructure company Composio integrated DeepSeek V4 Flash into 8 sets of Harness, testing with the same batch of 30 Agent tasks. The Harness is the execution framework outside the model, responsible for context, tool invocation, and task execution. The tasks involved real applications such as Gmail, GitHub, Slack, Calendar, and Notion.
The pass rates of the 8 sets of Harness are as follows:
1. Pi Agent: 20/30
2. Oh My Pi: 17/30
3. Claude Code: 16/30
4. Codex: 16/30
5. Deep Agents: 16/30
6. Hermes Agent: 15/30
7. Prime Agent: 15/24
8. OpenCode: 14/30
Simply by changing the Harness, DeepSeek V4 Flash was able to increase its passed tasks from 14 to 20.
There is also a significant difference in cost and speed. Pi Agent costs approximately $0.028 per successful task, while Claude Code costs $0.195, nearly 7 times as much. Claude Code is the fastest, with a median runtime of 122.7 seconds, while Oh My Pi is the slowest at 272.4 seconds.
However, Pi's test configuration differs from the uniform conditions: it uses high instead of max inference intensity, and out of the 30 tasks, 24 use the DeepSeek official API instead of OpenRouter. Therefore, this group of variances cannot be fully attributed to the Harness.
Prior to this, Composio also conducted another round of Harness cross-evaluation using Kimi K3. At that time, Oh My Pi ranked first with 22/25, while the original Pi Agent scored 18/25. OMP itself is an enhanced version of Pi forked off for Coding, with added capabilities such as sub-Agents. After switching to DeepSeek V4 Flash, the original Pi surprisingly took the lead.