Research·Set 06
Multimodal Learning and Embodied AI
How AI systems learn from many kinds of data at once — and what changes when they have a body.
-
What is a modality in AI? Give three examples of different modalities.
In AI, a modality is a type of data a model can take in and process. Humans experience the world through several senses, and in machine learning, modality means the kind of data a model works with, like audio, images, or text, and each type has its own properties that need different processing methods. For example, a picture is a grid of pixel values, while a sentence is a sequence of words.
Three examples:
- Text: written language, such as articles, captions, or chat messages
- Images/video: visual data from photos or cameras
- Audio: speech, music, or environmental sounds
Other modalities include depth data from lidar sensors, radar signals, and touch (tactile) data from robot sensors.
-
What is multimodal learning?
Multimodal learning is when a single AI system learns from two or more modalities at once instead of just one. It is a subfield of machine learning where models learn from several data types together, such as text, images, audio, and video, so they can use the complementary information each type provides. The main idea is that different modalities give different clues. A model that handles both images and text, for instance, can understand an image’s context better and make more accurate predictions.
In practice, these models usually have separate parts for processing each modality, and their outputs are then combined through methods like concatenation or attention mechanisms.
-
Where is multimodal learning used? Can you find one real application and identify the types of information it combines?
Multimodal learning is used in voice assistants, medical diagnosis (combining scans with patient records), chatbots that can read images, and self-driving cars.
Real application: Waymo self-driving cars. Waymo’s autonomous vehicles combine several types of sensor data through a technique called sensor fusion:
- Camera images: cameras pick out visual details like the color of a traffic light or a temporary road sign
- Lidar point clouds: lidar is strong at measuring depth and detecting the 3D shape of objects
- Radar: radar works well in bad weather and when tracking fast-moving objects
- Audio: external microphones help the car detect important sounds, like approaching emergency vehicles or railroad crossings
Combining these makes the system more reliable than any single sensor. For example, if the camera sees a stop sign, lidar can help the car figure out that it is actually just a reflection in a shop window or an image on the back of a bus.
-
What is embodied AI?
Embodied AI is AI that has a physical body and interacts with the real world, instead of living only as software on a computer. These systems use sensors to perceive their surroundings (like vision or touch) and actuators to take actions (like moving or picking up objects). Examples include robots, drones, and autonomous vehicles.
The key difference from regular AI is learning through action. Unlike purely software-based models, embodied AI senses the world, acts on it, and learns from the physical consequences. This matters because tasks that are easy for humans, such as picking up an oddly shaped object or navigating a messy room, remain very hard for AI. The idea goes back decades; roboticist Rodney Brooks, former director of MIT CSAIL, was a key figure who pushed for robots that learn from interacting with the world.
-
How are multimodal learning and embodied AI connected?
They depend on each other. A robot in the real world is naturally multimodal: it has to see (cameras), sometimes hear (microphones), feel (touch sensors), understand instructions (language), and then produce actions. Embodied AI can’t work well without multimodal learning to make sense of all that input.
A clear example is Google DeepMind’s RT-2. It is a vision-language-action (VLA) model that learns from both web data and robot data, and turns that knowledge into instructions for controlling a robot. In other words, it takes a camera image plus a text command and outputs robot movements. This let robots handle new situations; for example, a robot could sort files alphabetically by reading their labels and placing them in the right spots. More broadly, VLA models have replaced what used to be a multi-stage perception, planning, and control pipeline with a single learned system.
So you can think of it this way: multimodal learning gives AI the ability to understand the world through many senses, and embodied AI gives it a body to act in that world.
Links for more information
- Waymo, The Waymo Driver Handbook: How Perception Works — waymo.com
- Google DeepMind, RT-2: translating vision and language into action — deepmind.google
- Saturn Cloud, “What is Multimodal Learning?” — saturncloud.io
- Envisioning, “Embodied AI” — envisioning.com
- IEEE Signal Processing Society, Multimodal Learning short course — resourcecenter.ieee.org