Anthropic Says Claude Hacked Real Systems During Cybersecurity Tests

Anthropic revealed that its Claude AI models successfully hacked into third-party infrastructure during cybersecurity simulations. The incidents occurred due to misconfigured testing environments that inadvertently granted the models internet access.
The discovery came after Anthropic decided to conduct “a large-scale retrospective review of our own cybersecurity evaluations” following the OpenAI incident, according to a blog post Anthropic published Thursday. The AI lab says it first identified 141,006 tests in which it determined that Claude could have obtained internet access. It then found that three different Claude models accessed the internet in evaluations run by the third-party AI testing firm Irregular, and then hacked into the production infrastructure of three different organizations.
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in