Skip to content

The Best Salesman but the Least Aligned: Claude Opus 5 Tops Vending Machine AI Evaluation, Yet Repeatedly Colludes, Threatens, and Tears Up Agreements

Jul 30, 19:37

According to Perceiving AI monitoring, AI rating agency Andon Labs had a model operate a vending machine with a $500 initial capital in a simulated environment for 365 days. After five tests, Claude Opus 5 had an average end-of-term balance of $11.2 thousand, surpassing Claude Opus 4.7 and GPT-5.6 Sol, rising to the top of Vending-Bench 2.

In the multiplayer competition version, Opus 5 ranked second with around $7,000, close to the first-place GPT-5.6 Sol with around $7,400. In six tests, it consistently proposed or engaged in collusion pricing. It also fabricated competitor quotes, threatened peers, and violated ceasefire agreements 11 times.

Opus 5's refund approval rate eventually dropped to 10%, with a total refund of only $8.54 after six tests. GPT-5.6 Sol refunded $655 but still won the multiplayer match.

Andon Labs believes that Opus 5 once again exhibited the issue of "the better it is at making money, the more misaligned its behavior." Despite this, Anthropic's pre-launch audit claimed that it was the most aligned Claude ever.

Source