
Google DeepMind shared new developments in robotics and visual language models (VLM) on Thursday. The tech giant’s artificial intelligence (AI) research arm has used advanced vision models to develop new capabilities for robots. In a new study, DeepMind highlighted that Gemini 1.5 Pro and its long context window allowed the department to make improvements in robot navigation and understanding of the real world. Earlier this year, Nvidia also announced new AI technology that supports the advanced capabilities of humanoid robots.
Google DeepMind uses Gemini artificial intelligence to improve robots
Google DeepMind announced in a post on X (formerly Twitter) that it is using Gemini 1.5 Pro’s 2 million token context window to train its bot. A context window can be understood as a knowledge window that an AI model can see and use to process tangential information about the topic in question.
For example, if a user asks an AI model about the “most popular flavors of ice cream,” the AI model will examine ice cream keywords and flavors to find information relevant to that question. If this information window is too small, the AI can only react to the names of different types of ice cream. However, for larger quantities, the AI can also look at the number of articles about each ice cream flavor to see which ones are mentioned the most and to estimate the “popularity factor”.
DeepMind uses this long context window to train robots in real-world environments. The goal of this department is to find out whether robots can remember the details of their environment and help users when they have questions about the environment when it is unclear. In a video shared on Instagram, the AI department demonstrated how the bot can direct users to a whiteboard when asked where they can draw.
“With 1.5 Pro’s 1 million field length, our robots can use human instructions, video tours and common sense thinking to find their way through space,” Google DeepMind said in the post.
In a study published in arXiv, a non-peer-reviewed online journal, DeepMind described the technology behind the breakthrough. In addition to Gemini, we also use our Robotic Transformer 2 (RT-2) model. It is a Vision Language Action (VLA) model that learns from web and bot data. Use computer vision to process real-world environments and use this information to create datasets. This data set can then be processed by generative AI to analyze text commands and produce the desired results.
Google DeepMind is currently using this architecture to train robots in a broad category called Multimodal Instruction Navigation (MIN), which includes both environmental exploration and instruction-guided navigation. If the demonstration shared by the department holds up, the technology could advance robotics even further.
Bhupendra Singh Chundawat is a seasoned technology journalist with over 22 years of experience in the media industry. He specializes in covering the global technology landscape, with a deep focus on manufacturing trends and the geopolitical impact on tech companies. Currently serving as the Editor at Udaipur Kiran, his insights are shaped by decades of hands-on reporting and editorial leadership in the fast-evolving world of technology.

