Additional lessons not mentioned & what AI labs & testers need to do: 1. Monitor testing in real time, not months later 2. Prompt models to self-report lab escapes. These models knew what they’d done at some point 3. Set up a dedicated bidirectional reporting channel for victims
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. (1/4)