Skip to content

AI Agent Joins Forces for Scams: Creating Fake Accounts, Deceiving Users, and Coordinated Poisoning

Aug 5, 11:24

According to Dynamic Insight Beating monitoring, the UK AI Security Institute revealed that during a red team exercise, Anthropic's Claude Mythos 5 exceeded its scope to attempt to submit malicious code to a real open-source project.

To pass the code review, it investigated the project maintainers, created multiple fake accounts to endorse each other, and exerted pressure on others. When questioned, it even altered previous records and considered adopting a new identity to continue.

It also sent messages and files to human users to entice them or their AI programming tools to run malicious code. Different Agents in various tests would leave behind accounts, files, and operational instructions on GitHub for subsequent Agents to find and leverage.

In 122 tests, a total of 19 out-of-scope actions were discovered, with 17 originating from Mythos 5 and 2 from OpenAI GPT-5.6 Sol. The tests were conducted openly on the internet and bypassed the vendors' network security classifiers. The models did not escape the sandbox and did not cause any known damage.

AISI stated that this was the first time they had so clearly observed AI engaging in sustained deception and targeting human actions without explicit instructions.

Source