Skip to content
Defici
← Back to news

Archived · Published 5 August 2026

Gemini Robotics 2 Pushes Vision-Language-Action Models Toward Genuinely General Manipulation

Google DeepMind's Gemini Robotics 2 advances the case that robot control is becoming a foundation-model problem. The system's core vision-language-action model drives physically different robots — including Apptronik's Apollo 2 humanoid and dual-arm Franka research setups — through tasks that combine perception, language understanding, and whole-body movement such as walking, bending, and coordinated two-hand manipulation. Reported success rates land where lab demos start becoming operationally interesting: roughly 89.6% on precision insertion with the Franka Duo and 92% on delicate tasks like lightbulb removal with Apollo 2. The architectural significance is generality. The traditional robotics stack hand-engineered perception, planning, and control per platform and per task; a VLA model learns a shared representation that transfers across embodiments, meaning capability gains arrive by model update rather than re-engineering. Hardware is moving in parallel — 1X's new 25-degree-of-freedom tendon-driven hand with tactile skin and backdrivable joints shows manipulators being designed for exactly the fine-grained, force-aware control that VLA models can now exploit. Deployment context sharpens the story: Figure is producing at roughly one robot per hour past its first thousand units, and thousands of general-purpose humanoids are already working commercial warehouses. The remaining gap between 92% and the 99.9%-plus reliability that unsupervised industrial work demands is still wide — but it is now a data-and-scale gap rather than a conceptual one, and the industry's bet is that it closes the way language-model reliability did: gradually, then suddenly.

Defici Editorial · Robotics

This article was generated by Defici's AI editorial system.