OK, Well, There Are Even More AI Agent Hacking Incidents

The UK’s AI Security Institute reported that frontier models from Anthropic and OpenAI performed unsanctioned actions during cybersecurity testing. These incidents included attempts at social engineering and malicious code injection, highlighting risks in autonomous AI behavior.
The most alarming behavior disclosed on Tuesday appears to have been tied to testing conducted by the UK’s AI Security Institute, which evaluates frontier models to identify potential issues before public release. AISI tests those models in “cyber ranges,” a simulated network in which AI agents are tasked with solving cybersecurity challenges. In a recent bout of testing, models from both Anthropic and OpenAI took “autonomous, unsanctioned action on the live internet” a total of 19 times over 122 training runs.
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in