A modality is a type of information or a channel through which information reaches an AI system. Humans use several modalities at once, such as sight, hearing and touch, and AI systems can also take in data in different forms. Three common examples are text (written language), images (or video) and audio (speech or sound).
Multimodal learning is when an AI model learns from several types of data at the same time and combines them, instead of using just one. This is similar to how people understand the world: we combine what we see, hear and read to get a fuller picture. A model that looks at a photo together with its caption can understand more than a model that only sees the image or only reads the text. Making the different modalities line up with each other is one of the main challenges in this field.
It is used in many areas, such as medical imaging, social media analysis, emotion recognition, voice assistants and robotics.
One real application is self-driving cars. A car has to combine camera images (what the road looks like), lidar and radar readings (how far away objects are), GPS and map data (where it is) and sometimes audio (for example, sirens). No single sensor is reliable in every situation: cameras struggle in the dark or in fog, while radar can detect objects but cannot read a traffic sign. Combining the modalities makes the system more robust.
Embodied AI is AI that has a physical body (or a simulated one) and learns by acting in an environment, not just by processing data on a screen. Examples are robots, drones and virtual agents in simulated worlds. The agent perceives its surroundings, decides what to do, acts and then sees how its actions changed the world. Unlike a chatbot that only produces text, an embodied agent has to deal with real-world consequences, such as picking up an object without dropping it.
An embodied agent lives in a world full of different kinds of information, so it needs multimodal learning to make sense of it. A household robot, for example, must see (vision), understand an instruction like "bring me the red cup" (language), and then move its arm (action). Recent research has developed vision-language-action models that combine exactly these three modalities. In this way, multimodal learning works like the "brain" that understands the world, and embodied AI is the "body" that acts in it and gathers new experience to learn from.