Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

A study based on 40,000 game runs reveals that humans acting as a 'human-in-the-loop' for AI agents fail to catch malicious commands about one-third of the time. The data shows that users are more likely to approve deceptive commands that appear routine, such as those involving package scripts.
A couple of months ago I published a small browser game : you play the human-in-the-loop for an AI coding agent, approving or denying its commands under time pressure. Some commands are routine ( git status , npm test ) and some other commands indicate your agent has been possessed and is sending your secrets to a remote server ( cat ~/.aws/credentials ). More on the threats associated with agents running commands and how to mitigate them can be found in the original post .
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in