7/31/2026, 1:03:11 PM · robotics-physical

Google DeepMind Releases Gemini Robotics 2, a Three-Model Suite for Whole-Body Humanoid Control

The suite introduces a vision-language-action model capable of coordinating full humanoid bodies from feet to fingertips, an embodied reasoning layer for multi-step planning, and an on-device model that adapts to new robot hardware in hours.

Google DeepMind on July 30, 2026, unveiled Gemini Robotics 2, <cite index="18-5,18-6">positioning the suite as "the intelligence layer powering the next generation of truly adaptable robots" and unlocking intelligent whole-body control, advanced dexterity, and multi-robot collaboration.</cite>

Three-Model Architecture

<cite index="12-1,12-2,12-3">The release comprises three distinct models. Gemini Robotics 2 is the vision-language-action (VLA) model that turns what a robot sees and hears into motor commands — it handles the physical execution.</cite> <cite index="12-7">Gemini Robotics ER 2 is the embodied reasoning (ER) layer, a high-level brain that plans multi-step jobs.</cite> <cite index="12-12">The third model, Gemini Robotics On-Device 2, runs locally with no internet connection.</cite>

<cite index="13-3">The flagship VLA is the first DeepMind model capable of controlling a full humanoid from feet to fingertips, running a single model checkpoint across three distinct hardware configurations: Apptronik's Apollo 2 robot with SharpaWave hands, Apollo 2 with Inspire hands, and the Franka Duo bi-arm platform with a Robotiq gripper.</cite>

Whole-Body Control and Dexterity

<cite index="9-10,9-11">Previous Gemini Robotics models controlled only the humanoid's upper body for tabletop tasks; Gemini Robotics 2 extends control to whole-body motion for the first time.</cite> <cite index="1-3">The new model can instruct an entire humanoid to walk to a table, crouch to a lower shelf, and place an object precisely in a bin — all under a single learned policy that simultaneously coordinates legs, torso, both arms, and a 22 degree-of-freedom hand.</cite> Demonstrated tasks shown in a company-released video include <cite index="3-4">cleaning up trash, picking up watering cans, inserting a tape into a boombox, screwing in a lightbulb, and tying a garbage bag.</cite>

Embodied Reasoning and Multi-Robot Collaboration

<cite index="13-4,13-5,13-6">Gemini Robotics ER 2 is the high-level cognitive layer that accepts natural language, streams continuous video, plans multi-step tasks lasting several minutes, monitors progress, and decides when to call for human help. It acts as the brain while the VLA acts as the hands, connecting through the Gemini Live application programming interface (API) via a bidirectional streaming endpoint — allowing it to reason about upcoming steps while simultaneously executing current actions.</cite> <cite index="12-9,12-10">ER 2 can call tools like Google Search and steer a Boston Dynamics Spot robot to fetch objects; it also enables different robots to work as a team, such as a wheeled machine and a humanoid splitting a task.</cite>

On-Device Adaptability

<cite index="13-7,13-8">Gemini Robotics On-Device 2 is an efficiency-optimized VLA designed to run entirely on local hardware without cloud connectivity. It inherits "motion transfer" techniques from the Gemini Robotics 1.5 generation and can adapt to a completely new robot embodiment — different shape, sensors, and degrees of freedom — in a few hours with fewer than 200 demonstration examples.</cite>

Availability and Hardware Partners

<cite index="11-7,11-8">Gemini Robotics ER 2 is available through Google AI Studio and in private preview on the Gemini Enterprise Agent Platform, while the VLA and on-device models are being released to early-access partners.</cite> <cite index="1-12">Named deployment partners include Apptronik (Apollo 2), Franka (Franka Duo), and Agile Robots.</cite>

Safety and Benchmarking

<cite index="12-14">Alongside the release, Google DeepMind published a new benchmark called ASIMOV-Agentic, which tests whether the reasoning model will refuse an unsafe command from the action model and whether it flags when a task is impossible or escalates to a human.</cite> <cite index="1-10">DeepMind also published a dedicated Gemini Robotics 2 Safety Technical Report alongside the model release.</cite> <cite index="3-9,3-10">Carolina Parada, head of robotics at Google DeepMind, noted that safety concerns intensify as robots encounter greater situational variety, adding that the technology remains at an early stage and Google is not planning to release consumer-facing robots in the near term.</cite>

Cross-references

Sources

  1. [1]
    Gemini Robotics 2 Controls Full Humanoids: Legs, Torso, Arms, and Fingers Under One Policy
  2. [2]
    Project Mariner
  3. [3]
    Google DeepMind unveils Gemini Robotics 2 for autonomous robots - TechBriefly
  4. [4]
    Gemini (language model)
  5. [5]
    Gemini Robotics
  6. [6]
    Google's new Gemini Robotics 2 platform allows for 'intelligent whole-body control' - Engadget
  7. [7]
    Gemini Robotics 2 brings whole body intelligence to robots — Google DeepMind
  8. [8]
    Gemini Robotics — Google DeepMind
  9. [9]
    Google DeepMind Ships Three Physical AI Models For Whole Body Control, Dexterity And Multi Robot Collaboration - MarkTechPost
  10. [10]
    Google DeepMind unveils Gemini Robotics 2 as Apptronik humanoid demonstrates whole-body AI
  11. [11]
    Google's Gemini Robotics 2 gives humanoid robots full-body control
  12. [12]
    Google DeepMind’s Gemini Robotics 2 controls whole humanoids
  13. [13]
    Vision Language Action Models in Robotic Manipulation: A Systematic Review
  14. [14]
    Google Introduces Gemini Robotics 2 with ‘Whole Body Intelligence’