Claude Code runs Kimi K3, the most expensive and slowest option, costing nearly 4 times as much as Hermes Agent.
According to Insight Beating monitoring, AI Agent infrastructure company Composio expanded its testing by connecting the same Kimi K3 to 6 sets of Agent frameworks, completing 26 identical tasks. The comparison here is the performance of the frameworks, with all underlying models being Kimi K3.
Kimi Code completed 21 tasks ranking first. Hermes completed 20 tasks, Pi Agent and Claude Code both completed 19 tasks, OpenCode completed 18 tasks, and Codex lagged behind with 17 tasks.
Based on the official price of Kimi K3, Hermes and Pi Agent have the lowest average cost per task, at $0.39 and $0.40, respectively. Claude Code reached $1.47, approximately 3.8 times that of Hermes.
Pi Agent was the fastest, with a median task completion time of 161.7 seconds. Claude Code was the slowest, taking 347.6 seconds. Even with a difference of only 4 tasks in successful completion when using the same model with different frameworks, the average cost differed by nearly 4 times, and the time taken differed by over 2 times.