Will it run? Archive
Models

Google launches Gemini Robotics 2, the new intelligence layer of Google that teaches robots

By Rae Whitlock Clawpit staff
Google launches Gemini Robotics 2, the new intelligence layer of Google that teaches robots

to collaborate and adapt to new bodies Robots tend to be fairly stupid outside production lines. They are pre-programmed or remote-controlled to perform narrow, repetitive sequences, and struggle to learn autonomously or adapt to unpredictable environments. A harder problem is transferring skills from one robot body to another. To address this, Google is launching Gemini Robotics 2, an intelligence layer intended to let robots of any size or shape think, act and communicate on their own.

The system comprises three separate models. The first, Gemini Robotics 2, is a vision-language-action (VLA) model that converts visual and linguistic input into motor commands. The model can control full-body humanoid robots, from legs to fingertips, as well as dual-arm robots, offering enhanced manipulation for a variety of hands and grippers.

The second model, Gemini Robotics ER 2, handles embodied reasoning. It is a vision-language model (VLM) that serves as an agent, allowing robots to converse with humans, understand the physical world and plan multi-step tasks that last a few minutes. The model also introduces teamwork capability, so several robots can cooperate to finish a task more quickly. Google cites an example of a humanoid robot that can walk, crouch, stretch its body and clean a cluttered room while collaborating with another robot to complete the work.

The third model, Gemini Robotics On-Device 2, is a more efficient version of the VLA model, designed to run locally on the robot itself. The central claim is rapid adaptation to an entirely new robot body using a few hours of data. This represents a step change from the previous situation, where transferring skills between different platforms was especially complex. The announcement does not specify the exact amount of data required nor provide quantitative performance metrics for the tasks.

Google’s move envisions a comprehensive control layer that spans walking to delicate manual work under a single intelligence stack. The ability to plan long tasks and operate in teams hints at a shift from single-action scenarios to more complex use cases. However, the claim comes directly from the company and is not accompanied by independent benchmark data or external experiment documentation that would verify the stated adaptation times or the quality of inter-robot collaboration.

Clawpit — Back to top Clawpit