Anthropic AI agent created fake accounts to trick real people in security test, AISI says
A binary code can be seen on the display of a laptop. (Shot with zoom effect). Photo: Silas Stein/dpa (Photo by Silas Stein/picture alliance via Getty Images)
Artificial intelligence models running a cyber security test and given access to the internet reached out into the real world to target individuals and organizations without being told to do so, the AI Security Institute (AISI) reported.
In one case, the AI created fake accounts and tried to convince someone to approve malicious code so the model could accomplish its objective.
Why you should care:
The UK-based research institute said of the 19 total actions, which were spread over just 10 of the total of 122 runs, "the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations."
While these actions came as the result of AISI’s testing and were ended soon after being discovered, the institute noted that, as AI agents become more powerful, these behaviors could happen more frequently. The agents AISI used were not commercially available, and they did not escape their test environment.
The backstory:
AISI provided details of the incidents where the AI models went beyond the parameters established for the test and walked through what researchers considered the most serious actions taken in one event, which was performed by Anthropic’s Mythos 5 model.
What the AI agent tried
In it, the model attempted to add malicious code to a publicly used, open-source project and wanted human reviewers to approve it. The AI created the fake accounts and used them to socially engineer a real person to approve it. When the attempt was challenged, the agent modified its activity to seem more harmless and even considered creating a new identity to try again.
The agent continued to send messages and files to people to persuade them, or the AI models of the intended targets, to run the code. Some messages had harmful payloads, while others were social engineering attempts.
The agent also left messages on GitHub, a code repository platform, asking to work with other agents trying to complete the same challenge.
What they're saying:
"In these cases, standard good practice, human judgment, and caution around AI-generated code stopped the worst outcomes," AISI wrote in its statement. "But in several cases the margin between failure and success was narrow, resting on human vigilance rather than a technical barrier that would reliably prevent this behaviour in a more capable agent."
Dig deeper:
AISI, which is part of the British government’s Department of Science, pointed to recent incidents at OpenAI and Anthropic and said that those cases and its own point to "a shift in the risk landscape."
Not only can harm happen if someone deliberately misuses AI, it can also occur when agents take unintended actions while operating in an internal research or privileged-access scope to accomplish their goals.
Information for this article was taken from the AI Security Institute. This story was reported from Orlando.