Gemini Robotics 2.0: DeepMind's New Robot Brain
One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids
▶ Listen to the full episode More from The Cognitive Revolution
The brief
Google DeepMind's Keerthana Gopalakrishnan says robotics is still in its GPT-2 era despite Gemini Robotics 2.0 driving humanoid hands and grippers alike. The new architecture splits reasoning (ER2) from action (the VLA), needs just 200 examples to adapt to a new robot body, and treats safety as a capability rather than a constraint.
This summer, a robotics event in China went viral for a simple reason: humanoid robots ran faster than Usain Bolt. The clips spread everywhere, carrying one message: the machines have arrived. But according to Keerthana Gopalakrishnan, staff research scientist at Google DeepMind and research lead for Gemini Robotics, that spectacle is close to beside the point. "I'm a very productive human. A lot of my friends are also very productively employed, but we don't run faster than Usain Bolt," she says (07:18). Her point is that foot speed almost never limits what a robot can do for a person. What limits it is manipulation, the far harder job of making hands and fingers grip, twist, and place real objects without crushing or dropping them.
The reason running went viral first is also the reason it was never going to be the real test. Locomotion is, in Gopalakrishnan's words, easy to train in simulation, because the ground is mostly flat and rigid, and contact with it is predictable (08:55). Manipulation is a different animal. The moment an object deforms, like cloth folding over itself, friction and contact physics get messy, and today's simulators start to break (09:23). Pick and place tasks with rigid objects, cubes and cans, still simulate fine. Cloth, and anything springy or crumpled, does not.
That gap between what simulates well and what actually matters in a kitchen or a warehouse is the quiet theme of the whole conversation. It is also why Gopalakrishnan, despite being visibly proud of her team's recent progress, still rates the field as being in its "GPT-2 era" (20:13). She does not mean the models are unsophisticated. She means two specific things are still missing: reliable few shot learning across a wide range of tasks, and what she calls cross-embodiment, the ability of one brain to control very different robot bodies without starting over (20:41). A laptop runs the same software whether it is a Mac or a PC. A robot foundation model, by contrast, can still be helpless the moment someone swaps in an unfamiliar robot. Until that stops being true, she argues, it is hard to call the underlying model a generic brain.
The Architecture Behind the Demo
This year DeepMind released Gemini Robotics 2.0 to answer exactly that question, and the shape of the system maps neatly onto the old split between thinking and doing. The suite has three parts. Gemini Robotics ER2, for embodied reasoning, is the slow, deliberate layer, built on the Gemini Flash line, and it handles reading instruments, pointing at objects, and parsing what a person actually wants (24:40). It is available through DeepMind's API, so outside developers can wire it up to their own robots. Below that sits Gemini Robotics 2, the vision-language-action model, which turns ER2's instructions into actual joint movements. It now controls the whole humanoid body from fingertips to feet in a single closed loop, rather than handing stabilization to a separate, simpler controller (50:05). A third, smaller version, Gemini Robotics On-Device, compresses the same capability so it can run on the robot's own local computer, with no cloud connection required (25:36). That matters anywhere network latency is a problem, since a round trip to a cloud model, which the host's own tests sometimes clocked at six seconds, is simply too slow for some tasks (26:41).
The ER2 model's working memory has a concrete limit worth knowing. Its context window, the amount of past information it can hold while reasoning, is 128,000 tokens, which Gopalakrishnan translates to roughly three minutes of densely packed robot memory (32:42, 33:05). Past that, the system has to compress. Her advice for anyone building on top of it is not to simply keep more video frames but to summarize in text instead, the way a person recalls flipping an egg without replaying every frame of the pan (34:22, 35:08). Text is more compressible than images, and it still carries the narrative of what happened.
"Working on generalization is like lifting the boat, lifting the wave for all the boats." (Keerthana Gopalakrishnan, 15:17)
The instruction channel between ER2 and the VLA is still fairly narrow, mostly language plus pointing at objects in an image (46:59). That is a deliberate research stage, not a permanent ceiling. DeepMind calls this problem steerability, and the team is actively trying to widen that bandwidth, because right now, if the connection between the reasoning model and the action model is only text, some nuance of a scene can get lost in translation (47:08).
Hands, Safety, and the Next Frontier
Perhaps the clearest sign of progress is hardware. A year ago the frontier was grippers, simple two-fingered clamps. Gemini Robotics 1, released in spring of last year, showed real dexterity, but with grippers. By this summer, Gemini Robotics 2 was tying trash bags and performing fine multi-fingered control (57:45). Gopalakrishnan now says hands, not grippers, are where the dexterity research happens, since hand-equipped robots can do everything grippers could plus tasks grippers cannot attempt at all (58:14). Capability still varies a lot by hardware: the Sharper hand can lift about 20 kilograms and has opened jars, while the Wuji hand is closer in strength to a ten year old's grip (59:08, 59:31).
Underneath this sits a less flashy but more consequential finding. With around 200 examples, a general foundation model can learn a new task on a new robot body, because it already carries broad physical common sense and multi-embodiment control data to build on (51:47). That is the clearest evidence of cross-embodiment working in practice, though Gopalakrishnan adds a caveat: taking a brand new robot body and getting high reliability on many tasks with zero extra data still has no strong precedent (51:18). One brain, instantly fluent in any body, is closer than it was, but not yet real.
Safety, in this telling, is not a brake on progress but a feature competing for the same engineering attention as everything else. "I think of safety as, like, a capability, right?" she says (63:32), arguing that people simply will not adopt robots that are unsafe, so the most useful systems will also turn out to be the safest ones. She separates this from the familiar AI safety conversation about goals and guardrails. Robotics has its own category, operational safety, making sure a robot does not hurt someone through clumsiness rather than malice (64:27). In one internal demo, someone dropped a basket over a working humanoid's head mid-task. The right response was not to blunder on blind but to notice the obstruction and ask for help removing it (65:08).
That example points to something Gopalakrishnan returns to more than once: humanoid robots carry a psychological burden gripper-arms never had. Because they look human, people expect them to act smart, and forgive them less when they fumble.
"A humanoid just, like, grappling around and creating a lot of failures would be judged much harshly." (Keerthana Gopalakrishnan, 66:55)
It is a strange asymmetry. A robot arm that drops a box gets a shrug. A humanoid that drops the same box, in the same way, reads as somehow disappointing.
Looking further out, the conversation turns to a genuinely open question: should robots keep reasoning the way large language models do, by looking at the present moment and reacting over and over, or do they need a forward model, a way of predicting what happens next so they can measure their own error against reality. Gopalakrishnan declines to pick a side early. "You're not emotionally attached to one method or the other method. You're emotionally attached to the problem itself," she says (73:53). The same caution shows up in how she talks about training data. Teleoperated data is precise but ages badly as hardware changes, sensor-glove data scales better but is costly to collect, and raw human video is cheap and abundant but noisy, since no two human hands move quite the same way (77:08, 77:37). Her guess is that the field keeps using all three in some shifting mixture, rather than betting everything on one.
What makes this conversation useful, more than flashy, is its refusal to round up. Robots can now run faster than Olympic sprinters, fold laundry on request, and improvise a wave mid-conversation. They still cannot reliably fry an egg in a stranger's kitchen, because an egg offers no do-over: drop it, and the room is a mess (54:00). Somewhere between those two facts sits the real timeline for a useful robot in anyone's home, and Gopalakrishnan, who builds the brains for a living, is in no hurry to guess the date.
Ask this episode anything
Pod's AI answers from the episode itself, with the minute mark so you can hear it yourself.
Or start with one of these
Ask this episode
Pod's AI listens to the whole episode to answer, and points you to the minute it comes from.
The rest of this answer, and any question after it
Answers come from the episode itself, never from a summary of it.
ContinueKey takeaways
- Robotics is still in its GPT-2 era, says Google DeepMind's Gopalakrishnan
- Running is easy to simulate because ground contact is predictable; manipulating cloth breaks current simulators
- Gemini Robotics 2.0 splits the brain into ER2 for reasoning and a VLA that controls the whole body
- About 200 examples let a foundation model learn a new task on an unfamiliar robot body
- Multi-fingered hands have overtaken grippers as the frontier of robot dexterity research, per Gopalakrishnan
The episode in cards
By the numbers
- 2027 predicted year humanoid robots become broadly useful, per the AI 2027 forecast
- 128,000 tokens context window size of Gemini Robotics ER2, DeepMind's embodied reasoning model
- 200 examples demonstrations needed for a foundation model to learn a new task on a new robot body
- 20 kg lifting capacity of the Sharper robotic hand, per Gopalakrishnan
In their words
“I'm a very productive human. A lot of my friends are also very productively employed, but we don't run faster than Usain Bolt.”
“Working on generalization is like lifting the boat, lifting the wave for all the boats.”
“I think of safety as, like, a capability, right?”
“A humanoid just, like, grappling around and creating a lot of failures would be judged much harshly.”
Protocols
-
Compress long robot memory with text summaries
Keerthana Gopalakrishnan recommends converting dense image history into text summaries once a task phase finishes, because text is more compressible than video frames and still preserves the narrative of what the robot did. The catch is that summarizing loses some precise spatial detail that dense image frames would have kept, so this trades precision for efficiency.
Applied continuously during long task episodes that exceed the model's context window
-
Keep low-level access alongside high-level commands
Gopalakrishnan advises that robot operators keep both high-level language control and low-level, fine-grained control available for the same robot, because safety and personalization depend on the ability to override or adjust a task in detail rather than issue only a general instruction. The catch is that supporting both channels adds interface complexity, since the robot must handle a simple language command and a detailed manual override at once.
Built into the interface design for every deployment
Questions this episode answers
What is Gemini Robotics 2.0?
Gemini Robotics 2.0 is a suite of three Google DeepMind models: ER2, a reasoning model built on Gemini Flash; the Gemini Robotics 2 vision-language-action model, which controls a robot's full body from fingertips to feet; and a smaller on-device version that runs locally without a cloud connection (24:40, 25:36).
Why is robot manipulation harder to simulate than walking or running?
Locomotion mostly involves predictable contact with flat, rigid surfaces, which simulators model well. Manipulation, especially with deformable objects like cloth, involves friction and contact dynamics that break current physics simulators, according to Keerthana Gopalakrishnan (08:55, 09:23).
How much training data does it take to teach a robot a new task on new hardware?
Google DeepMind's trusted-tester program showed that around 200 examples were enough to teach new tasks across different robot bodies, since the underlying foundation model already carries general physical understanding and multi-embodiment control data to build on (51:47).
Are robot hands better than grippers now?
Keerthana Gopalakrishnan says multi-fingered hands have passed grippers as the frontier of dexterity research: within about a year, Google DeepMind went from gripper-only demos to a model that ties trash bags and performs fine multi-finger control (57:45, 58:14).
When will humanoid robots be broadly useful?
The AI 2027 forecast predicts humanoid robots become broadly useful by mid-2027 and that a robot economy forms by 2028 (00:28). Gopalakrishnan does not endorse a date herself, arguing robotics is still in its GPT-2 era because cross-embodiment generalization remains unreliable (20:13).
What is cross-embodiment in robotics?
Cross-embodiment means a single AI model can control many different robot bodies, not just the one it trained on. Gopalakrishnan says this remains unresolved, since a model that only works on its home robot is not yet a truly generic brain (20:41).
The full read, in cards
Go deeper
- AI 2027 — a forecast predicting humanoid robots become broadly useful by mid-2027 and a robot economy forms by 2028
Mentioned
Keerthana Gopalakrishnan · Google DeepMind · Gemini Robotics 2 · Gemini Robotics ER2 · Boston Dynamics · Apptronik · Agile Robots · Dr. Jim Fan · NVIDIA · Gemini 3.5 Flash













