One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids

Hello, and welcome back to the Cognitive Revolution!

Today I'm excited to welcome Keerthana Gopalakrishnan, Staff Research Scientist at Google DeepMind and Research Lead for Gemini Robotics back for her 4th annual appearance on the show. One of the biggest questions in AI today is: how soon will general-purpose robotics become broadly useful? AI is already affecting the world in major ways, even in purely digital form, but the most sci-fi forecasts for the AI future predict that robotics will soon hit key tipping points. AI 2027, for example, predicts that humanoid robots will become useful in mid 2027, and that by 2028, a whole "robot economy" could take shape, where robots build more robot factories, which in turn produce more robots, leading to unprecedented exponential economic growth and creating serious risk of AI takeover. So, how does that vision line up with the reality of robotics research today? We begin with a discussion of the extremely viral Robot Olympics held in China this summer, as I was really curious to get Keerthana's take on what it means that humanoid robots can now run faster than the fastest humans. Keerthana's take, which inverts the usual US-China dichotomy in AI, was that while the videos are impressive, skills like running on a flat track are relatively easy to train in simulation, and more to the point, footspeed is not really a limiting factor in the utility that robots can provide, and as such her team at Google is more focused on practical value. From there, we get into Gemini Robotics 2, a suite of 3 models that Google released this summer. Gemini Robotics ER - or Embodied Reasoning - 2, which Keerthana describes as a "system two" for robotics control, is based on Gemini Flash and is available via the API, allowing developers to define the affordances available to the model as tools, just like we do with digital agents. I played around with it and found it remarkably accessible. The other two models – Gemini Robotics 2 and Gemini Robotics On-Device 2 – translate higher level tool calls to robot actions, and are now capable of controlling the whole robot, from fingertips to toes, on a wide range of form factors, but are currently available to trusted testers only. Having understood how these models work, we again zoom out, and I ask Keerthana how she understands progress in robotics overall. On the one hand, like so many other AI researchers recently, she says that she's been surprised by the pace of progress, but at the same time, still feels that robotics remain in its GPT-2 era. While we've seen a number of recent demos of robots learning new tasks from just a few, or even just a single human demonstration, Keerthana argues that the range of tasks you can teach this way, and the generalization profile of robotics models across different robot bodies, still isn't strong enough to deliver the versatility and reliability that real-world use cases demand. From there, we move on to talk about recent progress in hardware, where robots are likely to be deployed at scale first, whether capabilities or safety, alignment, and adversarial robustness will ultimately be the limiting factor for the consumer market, and consider the future of the field, including whether we can continue to run LLM-derived robotics models in a loop as today's systems mostly do, or will need some sort of predictive world modelling, as we humans use, to create systems that can smoothly interact with the world. On that question, and on the timeline for key utility milestones, Keerthana isn't one to speculate too wildly – for her, these are all empirical questions to be resolved by research & experimentation – but what is clear is that robotics will either need to hit key tipping points soon, or some of the most aggressive timelines for AI transformation will be pushed back. For now, I hope you enjoy this very-well-grounded update on the state of robotics, with Keerthana Gopalakrishnan, from Google DeepMind.

Watch now!

Thank you for being part of The Cognitive Revolution,
Nathan Labenz

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to The Cognitive Revolution.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.