Particle.news
Download on the App Store

Google DeepMind Launches Gemini Robotics 2 for Whole-Body Humanoid Control

A high-level embodied reasoning model coordinating vision-to-action with on-device action enables robots to plan multi-step tasks with low latency.

Overview

  • Gemini Robotics 2 was announced last week and DeepMind made the embodied reasoning component, ER 2, available to developers through the Gemini API and Google AI Studio preview while other action models remain in private or early access.
  • ER 2 serves as a high-level 'brain' that streams video, audio, and text to break instructions into steps, monitor continuous task progress, and call tools such as Google Search to guide decisions in real time.
  • The release adds vision-to-action models that translate ER 2 plans into whole-body motor commands so humanoids can walk, balance, crouch, and use five-fingered hands to perform tasks like tying knots or unscrewing light bulbs.
  • DeepMind published internal metrics showing ER 2 scores of 57.4% on continuous progress classification and 91.3% on moment finding with a 0.96s mean absolute distance and introduced the ASIMOV-Agentic benchmark to evaluate agentic safety behaviors.
  • Google and partners demonstrated multi-robot collaboration on hardware from Apptronik, Franka and Boston Dynamics, but independent validation, real-world intervention rates, deployment costs and broad production timelines remain unreported.