Gemini Robotics 2 Brings Google's AI Into the Physical World

Google DeepMind has released Gemini Robotics 2, a system that integrates vision language models with physical action models to enable autonomous robot tasks. The technology allows robots to reason through complex environments, though experts warn of potential safety risks.
Gemini Robotics 2 combines several different AI models into a single system. Taken together, they allow a robot to make sense of its surroundings and how to act in it. A vision language model (VLM), which understands images and video, can communicate with humans and reason how to perform different tasks. Two vision language action (VLA) models, trained to understand how to move in physical space, control the robot’s full-body movement as well as the movements of grippers or hands.
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in