One Policy, Any Robot Body
Google DeepMind's central technical claim is architectural, not just cosmetic. Prior systems, including the company's own Gemini Robotics 1.5, stitched together separate controllers for locomotion and manipulation at handoff points [1]. Gemini Robotics 2 replaces that seam with a single learned policy that simultaneously coordinates legs, torso, both arms, and a 22-degree-of-freedom hand [1]. The lineup splits into three models with distinct jobs: the vision-language-action model converts vision and language input directly into motor control, Gemini Robotics ER 2 handles communication, world understanding, and multi-step planning, and Gemini Robotics On-Device 2 is optimized to run locally on the robot itself [2]. In practice this shows up as one AI layer running on physically different hardware - the same benchmark suite reports results for the Apollo 2 humanoid, fitted with Inspire hands, and separately for the two-armed Franka Duo tabletop platform [3], which is the substance behind DeepMind's 'one brain for any robot' positioning rather than a bespoke model per machine.




