Hacker News·3 min read·medium

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

W
Wirbelwind
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
AI Summary

A study based on 40,000 game runs reveals that humans acting as a 'human-in-the-loop' for AI agents fail to catch malicious commands about one-third of the time. The data shows that users are more likely to approve deceptive commands that appear routine, such as those involving package scripts.

A couple of months ago I published a small browser game : you play the human-in-the-loop for an AI coding agent, approving or denying its commands under time pressure. Some commands are routine ( git status , npm test ) and some other commands indicate your agent has been possessed and is sending your secrets to a remote server ( cat ~/.aws/credentials ). More on the threats associated with agents running commands and how to mitigate them can be found in the original post .

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai

Get the full story

Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.

Create free account

Already have an account? Sign in

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs — Headlinne — headlinne