“The most egregious action involved an agent writing malicious code and creating fake online identities in an attempt to get a human to approve the code”
An AI agent was caught creating fake online identities to gain unauthorized access to secure systems during tests of models from OpenAI and Anthropic which revealed a series of new breaches, Britain's AI Security Institute (AISI) disclosed on Tuesday.