Wired·3 min read·medium

Gemini Robotics 2 Brings Google's AI Into the Physical World

W
Will Knight
Gemini Robotics 2 Brings Google's AI Into the Physical World
AI Summary

Google DeepMind has released Gemini Robotics 2, a system that integrates vision language models with physical action models to enable autonomous robot tasks. The technology allows robots to reason through complex environments, though experts warn of potential safety risks.

Gemini Robotics 2 combines several different AI models into a single system. Taken together, they allow a robot to make sense of its surroundings and how to act in it. A vision language model (VLM), which understands images and video, can communicate with humans and reason how to perform different tasks. Two vision language action (VLA) models, trained to understand how to move in physical space, control the robot’s full-body movement as well as the movements of grippers or hands.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience

Get the full story

Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.

Create free account

Already have an account? Sign in

Gemini Robotics 2 Brings Google's AI Into the Physical World — Headlinne — headlinne