Recently, Google DeepMind launched “Gemini Robotics,” an advanced vision-language-action (VLA) model based on the Gemini 2.0 large language model, stating that with Gemini Robotics, robots will be able to improve their ability to perform tasks in the real world. Gemini Robotics offers generality, interactivity, and dexterity, and is not limited to a single robot form—whether you have a bipedal humanoid robot or a dual-arm robot, you can use the Gemini Robotics model. Let’s take a look at what changes the newly launched Gemini Robotics VLA model can bring to robots!
Google DeepMind has launched its own VLA model, “Gemini Robotics,” enabling robots to perform high-precision actions such as origami folding and basketball shooting.
Google DeepMind has unveiled its advanced vision-language-action model “Gemini Robotics,” built on the Gemini 2.0 large language model. Robots powered by Gemini Robotics can respond to text, images, sound, and video, with enhanced reasoning capabilities for physical space.
Gemini Robotics is Google DeepMind’s latest AI model, designed to enhance robots’ ability to perform complex tasks in the real world. The model combines language, vision, and action, allowing robots to understand and adapt to various environments and handle tasks such as folding origami, tidying desktops, wrapping headphone cables, and shooting basketballs.
The Gemini Robotics model leverages Gemini’s understanding of the world to quickly adapt to new situations, including handling new objects it has never been trained on before. It can also understand and respond to everyday instructions, reacting to sudden changes in commands or the environment. For example, playing tic-tac-toe, spelling words on demand, and packing food into lunch bags.
What’s most impressive is that the Gemini Robotics model adapts to various robot forms, from dual-arm robot platforms to humanoid robots. According to information from Google DeepMind, they are currently collaborating with robotics companies such as Apptronik, Agile Robots, Agility Robotics, Boston Dynamics, and Enchanted Tools.
In addition to the Gemini Robotics model mentioned above, Google DeepMind has also introduced the Gemini Robotics-ER model, which focuses on enhancing robot reasoning capabilities.Gemini Robotics-ER modelSupports developers in using its reasoning capabilities to run custom programs.
Source: KOCPC Chinese