Hacker News·4 min read·hard

What's the largest software project AI can complete on its own?

Y
yusufozkan
What's the largest software project AI can complete on its own?
AI Summary

Researchers have developed MirrorCode, a benchmark designed to test AI models on long-horizon software engineering tasks that require reimplementing entire programs. The study demonstrates that current models like Claude Opus can successfully complete complex, multi-week coding projects without human intervention.

AI has made rapid progress on software engineering benchmarks in the past few years. However, most such benchmarks tend to focus on shorter tasks like fixing bugs or implementing individual features. MirrorCode is our benchmark, co-developed with METR, to test AI models on long-horizon coding tasks. In a MirrorCode task, AI models are tasked with reimplementing an entire program end-to-end, without access to the original source code. AI-generated solutions must match the original program’s output exactly on end-to-end tests, including held-out tests. MirrorCode’s 25 target programs span different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience

Get the full story

Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.

Create free account

Already have an account? Sign in