Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
Needle2 is a new, highly compressed 45M-parameter AI model designed to run efficiently on low-power edge devices like wearables, robots, and budget smartphones. By focusing on tool calling and structured extraction, it provides functional AI capabilities without needing massive computational resources.
Today we release Needle 2: an open 45M-parameter model for tool calling, device use and structured extraction. The whole model is a single 14MB binary that runs a full session in 28MB of RAM. It is built on our Simple Attention Network findings, compressed to CQ2-bit with Cactus Quants , and baked into its own engine.
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in