Google DeepMind has made a significant leap in robotics development by unveiling the updated Gemini Robotics 2 model. This solution, based on the VLA (Vision-Language-Action) architecture, is capable of controlling a robot's entire body, translating text and visual data into complex motor actions. Alongside it, the ER 2 tool was announced, designed to coordinate the actions of multiple devices.

From Manipulators to Humanoids: Full Body Control

The key difference of the new version from previous ones lies in the expanded control zone. While artificial intelligence previously focused primarily on controlling the upper body, Gemini Robotics 2 allows robots to master locomotion and complex physical interactions. Now, machines can walk, squat, bend, and reach for objects located in hard-to-reach places.

The model is universal and adapts to various form factors — from compact desktop manipulators to full-sized humanoids. In demonstration videos, the Apptronik Apollo 2 robot, controlled by the new system, successfully performed household tasks: picking up a watering can from the floor, taking a baseball glove off a shelf, and finding necessary items in a cluttered space.

Fine Motor Skills and Complex Manipulations

DeepMind also emphasized a significant improvement in motor skills. The new model supports working with complex robotic hands possessing 22 degrees of freedom. This opens up possibilities for performing tasks requiring high precision and dexterity. Such actions include tying knots, sealing bags, and unscrewing light bulbs.

ER 2: A "High-Level Brain" for Swarm Logic

Parallel to the body control model update, the company presented improvements for the ER 2 solution. DeepMind characterizes this tool as a "high-level brain" for robotics. The system is responsible for planning multi-step tasks, monitoring their execution, and making decisions to repeat actions in case of failure.

An important function of ER 2 is the distribution of work among multiple robots. In one of the demonstration scenarios, the Apollo 2 humanoid acted as a coordinator, assigning a two-armed Google robot to pack tools into a container during a garage cleanup.

Local Learning and Accessibility

To address tasks requiring autonomy, the company showcased the local Gemini Robotics On-Device 2 model. It operates without an internet connection, which is critical for safety and privacy in industrial environments. A feature of the model is its high learning speed: it adapts to new types of robots in less than 200 examples, taking only a few hours.

Currently, Gemini Robotics 2 is available in private preview mode. Access to ER 2 is open via the Gemini API and Google AI Studio, while the On-Device 2 local version is being tested by trusted partners. This event should be viewed in the context of the general trend of implementing AI in the physical world, which other market players, including the company Humanoid, have also noted, presenting an approach to training robots on real production tasks in July.