Tencent has established its own Agent benchmark, but its in-house CodeBuddy has not passed the Claude Code test.
According to the latest Beating Monitor report, Tencent conducted a set of Agent benchmark tests using the WorkBuddy Bench framework, pitting their in-house solutions CodeBuddy Code and Claude Code against each other.
The benchmark included 260 real-world work questions divided into four categories: code, web, office, and security. Seven models were each run through two sets of Harness for testing. The Harness serves as the outer execution layer of the model, handling context, tool invocation, and task execution.
Out of 28 model comparisons, Claude Code won in 17 instances, while CodeBuddy only triumphed in 11. When it came to code, it was a clean sweep of 7:0, with all seven models scoring higher when replaced with Claude Code.
However, CodeBuddy did not suffer a complete defeat. It slightly outperformed in web and office with a 4:3 score, but Claude Code took the lead in security with a 4:3 score.
Simply swapping out the Harness for the same model resulted in a difference of over a dozen points in performance. This indicates that evaluating the strength of an Agent cannot solely rely on the underlying model.
Nevertheless, Tencent's benchmark report also revealed some significant shortcomings. After splitting the 260 questions into four categories, each category only contained between 50 and 80 questions. Some members of the community directly questioned whether adding more difficult questions would change the rankings. Furthermore, individuals on GitHub expressed dissatisfaction with the limited testing of only two sets of Harness, without even including Codex.