Robots that fit into our world — we’ve been talking about it for decades. On July 30, Google DeepMind took a real step in that direction: Gemini Robotics 2, the intelligence layer for its next generation of robots. And this time it’s no longer just gripper arms on a tabletop.
What’s new
Until now, DeepMind’s models mostly controlled a robot’s upper body for table-height tasks. Gemini Robotics 2 can now control the entire body — walking, crouching, stretching, balancing. In the demo, an Apollo 2 humanoid tidies a cluttered room: it walks to a table, picks up a watering can, steps over to the shelves, and places it precisely where it belongs.
On top of that comes real dexterity. The model drives a five-fingered hand with 22 degrees of freedom, pulling off things like tying knots or sealing a ziplock bag. To be fair, success rates on the trickiest multi-finger tasks are still all over the place — screwing in a lightbulb sits at 36 percent. But the direction is right.
The three models
DeepMind isn’t shipping one model but three: Gemini Robotics 2, a vision-language-action model for the actual movement; Gemini Robotics ER 2, the reasoning brain that plans multi-minute tasks and coordinates several robots; and Gemini Robotics On-Device 2, which runs entirely locally — no network, no latency. That last one adapts to brand-new robot bodies in a few hours, often with fewer than 200 examples.
There’s also multi-robot collaboration: different robots talk to each other and solve tasks together that a single one couldn’t handle alone.
Why it matters
For those of us in the Claude corner, this is a Google story on the surface. But the trend behind it touches everyone: the big labs are pushing their language models out of the chat window and into the physical world. Anthropic is experimenting with Project Fetch on a robot dog, OpenAI is building a robotics team. The race for “AI that touches things” has started.
My take
What reassures me about this release: DeepMind makes safety a core theme. There’s a new benchmark called ASIMOV-Agentic that measures whether the robot refuses unsafe commands and, when uncertain, asks a human instead. That’s exactly what you want once these machines stand next to us. The videos still look a little slow and clumsy — but I remember how clumsy the first chatbots sounded. That moves fast.
Sources: