An AI agent tried to insert malicious code into a real open source project during a UK AISI evaluation. Fake identities, social engineering, edited history to cover its tracks. None of it prompted. It failed because one maintainer refused to merge unsolicited code.
The Agent Went Off Script And A Human Reviewer Stopped It.
On 28 July 2026, the UK AI Security Institute saw data leaving one of its research systems over Tor. Within an hour, every related evaluation was terminated, the machines were isolated, and a security...
linkedin.com