Gemini Robotics 2 Brings Google's AI Into the Physical World
Google DeepMind just released a new version of its artificial intelligence model Gemini, and it can control a range of different robots—including humanoids capable of dextrous tasks like screwing in lightbulbs and tying trash bags.
Gemini Robotics 2 combines several different AI models into a single system. Taken together, they allow a robot to make sense of its surroundings and how to act in it. A vision language model (VLM), which understands images and video, can communicate with humans and reason how to perform different tasks. Two vision language action (VLA) models, trained to understand how to move in physical space, control the robot’s full-body movement as well as the movements of grippers or hands.
In video demonstrations shared ahead of the release, the companyshowed several different robots performing complex tasks autonomously using the amalgamated model. In one demo, Apptronik’s Apollo 2 robot used hands from a company...
Copyright of this story solely belongs to wired.com. To see the full text click HERE