“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world,” the institute said in a blogpost.
AI models shock UK testers by using fake identities to try to trick developers
AI Security Institute says OpenAI and Anthropic models went rogue during a cybersecurity test and showed a new type of risk
theguardian.com