Skip to content

Terminal-Bench 3.0 Leadership Change: Opus 5 Outperforms GPT-5.6 Sol by 43.5%

Aug 12, 16:50

According to Perceive Beating monitoring, Terminal-Bench 3.0 has a new leaderboard king. This evaluation requires the AI Agent to complete tasks in a real terminal environment and then directly check if the results are correct.

The current top three are:

1. Claude Opus 5 Max + mini-SWE-agent: 43.5%
2. GPT-5.6 Sol Max + Codex: 34.6%
3. Claude Fable 5 + Claude Code: 34.1%

Terminal-Bench 3.0 was launched on July 23rd with the aim of making the tasks more challenging. The scores of the old cutting-edge models have already converged, and the new version hopes to widen the gap again. The first edition includes 74 tasks covering 7 domains, and introduces more complex environments such as GPU and multi-container networks. The tasks may involve not only coding but also submitting model weights, formal proofs, spreadsheets, and CAD files.

However, this should not be directly interpreted as Opus 5 itself being nearly 9 percentage points stronger than Sol. Terminal-Bench allows models to run with different Agents, with the top three using mini-SWE-agent, Codex, and Claude Code, respectively. The Agent determines how the model interacts with tools, executes commands, and handles context, so the leaderboard more accurately measures the overall capability of the "model + Agent" system.

Source