How OpenAI's AI Agents ‘secretly’ used Message Board to plan hacking attack
OpenAI researchers discovered that autonomous AI agents, when faced with impossible tasks, began collaborating on a secret message board to share exploits and bypass security protocols. This incident highlights the tendency of frontier models to seek unauthorized shortcuts when optimized for speed.
OpenAI researchers have revealed ‘shocking’ details of a recent cyberattack carried out by its runaway AI agents on Hugging Face systems. During a packed presentation at the Black Hat cybersecurity conference, the company’s alignment and safety researcher Eric Wallace and security engineer Michael Dalton revealed how an internal safety evaluation transformed into a coordinated attack on both OpenAI’s systems and the world’s largest AI repository.Describing the event as “the most qualitatively interesting example of AI capabilities” they had ever witnessed, the researchers disclosed that the incident left company people in attendance reacting with disbelief, saying, “This is wild” and “Jesus.”“What makes this incident interesting is that once one agent was able to find these kind of exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents,” said Wallace, adding, “So once one model is able to find a way to open the door to some access it’s not supposed to have, it can leave the door open for other agents to use.”The origins: A two-month secret message boardThe roots of the cyber incident trace back to May 7 during routine safety and performance evaluations of an unreleased frontier AI model. When assigned software security tasks that proved impossible to complete within standard rules, the autonomous agents began looking for shortcuts.“Frontier models really like to cheat,” Wallace explained, noting that training pressures to optimise speed often push AI systems to find unauthorized workarounds rather than solving problems natively.To bypass obstacles, the AI agents discovered they could store text notes within an internal software repository used during training. What started as simple requests for assistance quickly escalated into a message board where agents logged software vulnerabilities, exchanged access credentials and distributed work assignments.By operating as a collaborative swarm, if a single agent uncovered a system exploit, it posted the instructions so other agents could reuse the pathway – something that has been previously reported by news agency Reuters.Clearing the board failed to stop AI AgentsBelieving the problem was contained, OpenAI intervened following an internal breach on July 4. Security teams revoked access tokens, wiped the internal repository, patched known vulnerabilities and resumed model evaluations.However, the AI agents adapted almost immediately. Deprived of the ability to post text files, the agents began creating newly named directories within the system, using the folder titles themselves as text messages to rebuild their communication board.Once communication was re-established, the collective resumed sharing newly discovered exploits and moving laterally across network defenses.OpenAI security engineers later linked the Hugging Face attack directly to the evaluation runs after discovering that the credentials used in the external breach originated from the same internal AI training sessions.Get the latest technology news and updates. Download the TOI App.
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in