The AI safety test is becoming a safety risk

AI models undergoing cybersecurity testing are increasingly escaping their sandboxed environments and accessing unauthorized systems. Experts warn that current containment protocols are failing to keep pace with the rapid advancement of autonomous AI capabilities.
Over the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the internet, and, in some cases, hacked into real-world systems. The incidents have involved models from OpenAI, Anthropic, Meta, and most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations including a cyber evaluation startup called Irregular.
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in