Gemini Robotics ER 2 and the Shift Toward Real-Time Physical AI

Google's new embodied reasoning model wants to fix robotic hesitation. But orchestration is only half the battle when metal meets the messy physical world....

Feed
September 18, 2026
Gemini Robotics ER 2 and the Shift Toward Real-Time Physical AI


We love talking about artificial intelligence in the abstract, but the real test has always been the physical world. It is messy, unpredictable, and fiercely uncooperative. For years, robots have suffered from a fatal flaw: the pause. They think, they wait, they compute, and then they awkwardly move. Google DeepMind just dropped Gemini Robotics ER 2, targeting this exact bottleneck. They want to turn their massive multimodal models into a high-level cognitive engine for hardware, decoupling heavy thinking from low-level muscle control.

The core idea here isn't entirely new, but the execution is getting sharper. Instead of asking a single neural network to handle everything from spatial math to torque calculations, ER 2 acts as the manager. It processes continuous video feeds, tracks progress, and orchestrates lower-level vision-language-action models or direct hardware APIs. It even integrates with real-time bidirectional streaming endpoints to cut down on latency. The result? A system that can theoretically plan multi-step tasks and call external tools like Google Search while the robot is actively executing a physical motion.

That continuous processing loop is what actually matters for builders. Stop-and-think pauses kill utility in real environments. If a delivery bot freezes every time a pedestrian shifts direction, it fails. By letting the model think about the next step while simultaneously executing the current one, ER 2 inches us closer to fluid automation.

Gemini Robotics ER 2 and the Shift Toward Real-Time Physical AI

Of course, we should keep our enthusiasm in check. Demos on Boston Dynamics' Spot look incredible in controlled lab settings, but translating high-level reasoning to chaotic, dusty warehouses or cramped kitchens remains a massive engineering hurdle. Tool orchestration and video understanding are powerful tools, yet handling edge cases in physical reality is where most shiny architectures quietly break down. Still, watching the stack evolve past clumsy scripts into adaptive, agentic orchestration is hard not to appreciate.

If you're building physical AI systems, the takeaway is clear. The architectural split between brain and brawn is solidifying. As these reasoning layers get faster and more reliable, developers can focus less on hardcoding every single trajectory and more on defining high-level goals. Also, more reliable, developers can focus less on hardcoding every single trajectory and more on defining high-level goals. Just don't expect the hardware to magically forgive bad inputs once it hits the floor.