Week 05 · Research Update

Multimodal Learning and Embodied AI

1. What is a "modality" in AI?

Imagine you are trying to understand what is happening in a room. You could:

  • Read a note that says what's happening (that's text)
  • Look at a photo of the room (that's an image)
  • Listen to a recording from the room (that's audio)

Each one of these is called a modality. A modality is simply one type, or one format, of information. It's like a channel that information can travel through. Humans have many senses (eyes, ears, skin), and each sense is like a different modality of information about the world. AI borrows this same idea.

Three simple examples of modalities:

  • Text: written words, like a sentence, a book, or a chat message.
  • Image: a picture, like a photo of a cat or an X-ray scan.
  • Audio: sound, like someone's voice, music, or a car engine noise.

Other modalities exist too (video, touch, temperature, motion) but text, image, and audio are the three easiest to picture.

If an AI model only works with one of these types of data, we call it a unimodal model. For example, an old-style image classifier that only ever looks at pictures is unimodal, it never reads any text and never listens to any sound.

2. What is multimodal learning?

Now, multimodal learning means teaching an AI model to understand more than one modality at the same time. For example, understanding text and images together, or video and sound together.

This is useful because most normal AI systems only handle one kind of data. A text-only model like GPT can only read words, it cannot see a photo. An image-only model can only look at pictures, it cannot hear a sound or read a sentence. But in real life, information almost never comes in just one neat package. Think about watching a movie: you need both the video and the sound to really understand the story. Imagine watching a scary movie with no sound, or listening to it with your eyes closed. You would miss half the meaning. Multimodal learning is built exactly to solve this kind of problem, letting one AI model combine information from more than one source at once, just like your brain naturally does.

So how does a computer actually do this?

  • Step 1: Encoding (understand each piece separately first): Each type of data gets its own reader. A picture gets read by a part of the model that is good with images. A sentence gets read by a different part of the model that is good with language. It's like having a translator for each language before a meeting.
  • Step 2: Fusion (mix the pieces together): After each piece has been read, the model combines, or fuses, all of that information into one shared understanding. This can happen early on (mixing raw, unprocessed information), in the middle (mixing partly-processed information), or late (mixing final decisions from each piece).
  • Step 3: Prediction (make a decision using everything together): Now that everything is combined, the model uses this rich, mixed understanding to do a task like answering a question, describing a picture, or deciding what to do next.

There's a fun example that shows why mixing senses matters so much: it's called the McGurk effect. Scientists found that what you hear someone say can actually change depending on what you see their mouth doing; Your brain automatically blends sound and sight together, sometimes even hearing a different sound than what was actually played, just because of what your eyes saw at the same time. This shows that even human brains are constantly doing multimodal learning without us noticing.

3. Where is multimodal learning actually used? A real example

Multimodal learning is used in lots of places: voice assistants that understand both your spoken words and a picture you show them, hospitals that combine medical scans with written doctor notes, and more.

But let's look closely at one very clear real example: self-driving cars.

Self-driving cars are a perfect example because the car cannot just use ONE type of sensor, it needs to combine several different types of information to drive safely, kind of like how you use your eyes, ears, and sense of balance all together when riding a bike.

Here is what a self-driving car combines:

  • Cameras: these work like eyes. They give lots of color and detail, so the car can recognize what something is (a person? a stop sign? a dog?). The weakness: cameras get confused in bad light, rain, or fog.
  • LiDAR: this is a sensor that shoots out laser beams and measures how long they take to bounce back, which tells the car exactly how far away something is, in 3D. The weakness: it's expensive, and can also struggle in heavy rain or fog.
  • Radar: this uses radio waves. It's very good at working in bad weather and telling how fast something is moving. The weakness: it gives a blurrier picture than cameras or LiDAR.
  • Ultrasonic sensors: good for very short distances, like when parking.
  • GPS and motion sensors: these tell the car where it is and how it's moving.

None of these sensors is good enough by itself, each one is missing something. So the car's AI does exactly what we talked about above: it fuses all of this different information together (this is called sensor fusion) so that the weaknesses of one sensor get covered by the strengths of another. For example, radar still works well in fog even though the camera can't see well in fog, so combining them keeps the car safe even when one sensor is struggling.

Real examples of this fusion happening inside the car:

  • Adaptive cruise control: mixes radar and camera data to keep a safe following distance from the car in front.
  • Lane-keeping assist: mixes camera and LiDAR data to see the lane lines and any obstacles.
  • Pedestrian detection: mixes camera and radar data so the car can reliably notice a person crossing the road.
  • High-definition mapping: mixes GPS, LiDAR, and motion sensor data to know exactly where the car is on the road.

So the types of information (modalities) being combined here are mainly: visual images (from the cameras), 3D distance/shape data (from LiDAR), speed/motion data (from radar), and location data (from GPS). All of these very different data types get fused into one single, trustworthy understanding of the road, that is multimodal learning, being used to literally keep people safe while driving.

4. What is embodied AI?

Normal AI usually lives only inside a computer; It reads data, thinks, and gives an answer on a screen, but it never actually touches the real world. Embodied AI is different: it means putting an AI brain inside something with a physical body, usually a robot, so that it can actually sense the world around it and physically act inside it, not just think about it on a screen.

Giving a computer a body is a bit like giving a brain arms, legs, eyes, and ears, so it can walk around, pick things up, and actually change the world. Instead of only reading about the world through text or pictures someone gives it.

More formally: embodied AI studies AI agents (robots) that learn and act by physically interacting with an environment, instead of only learning from a fixed, already-collected pile of data.

The key idea is a loop: the robot senses the world → decides what to do → acts → this action changes what it senses next → and the cycle repeats. This loop is what makes embodied AI special, it's an ongoing back-and-forth between the robot and its surroundings.

How does a robot actually do this? It uses sensors, cameras (to see), microphones (to hear), and sometimes touch sensors, to understand what's happening around it. Then it uses this understanding to make a decision, and finally it uses motors (its muscles) to actually do something, like moving an arm or driving forward.

The newest and most exciting version of embodied AI today is something called a vision-language-action (VLA) model. This is a single AI system that can: see through a camera, understand a spoken or written instruction (language), and then directly move a robot's motors to complete that instruction, all in one connected system.

5. How are multimodal learning and embodied AI connected?

A robot's senses are always multimodal, and embodied AI is basically what happens when you give a multimodal AI a body to act with.

A robot needs to understand its surroundings. To do that, it uses several different sensors at once, a camera (sight), a microphone (hearing), and sometimes touch sensors. Understanding all of these together, at the same time, is exactly what multimodal learning is about, mixing different types of data into one understanding. This is the exact same idea as the self-driving car example above, just now inside a robot's body.

But embodied AI goes one step further than plain multimodal learning: it adds action. The newest robots (like the vision-language-action models mentioned above) take in a camera image (one modality) plus a spoken or written instruction (a second modality), and then produce a physical motor movement as the output. That's three different types of information (seeing, understanding language, and moving) all combined in one system.

  • Multimodal learning = combining different types of information into one understanding.
  • Embodied AI = taking that combined understanding and giving it a body, so it can also act in the world and learn from what happens next.

This connection matters a lot for the future of AI, because an AI that can only browse the internet and write text is useful, but an AI that can also walk through a warehouse, pick up objects, or help in surgery is a much bigger leap, and building that kind of robot requires combining computer vision, language understanding, and physical motor control all together, which is fundamentally a multimodal problem.

In short: multimodal learning gives an embodied AI its combined senses, and embodiment gives multimodal learning a body and a feedback loop to actually use those senses in the real world.

References

  1. Deepgram. Multimodal AI (AI Glossary). Link
  2. Voxel51. Embodied AI (Glossary). Link
  3. arXiv (Yurtsever, Lambert, Carballo, Takeda). Self-Driving Cars: A Survey. Link
  4. ThinkRobotics. Sensor Fusion for Autonomous Vehicles. Link

Further reading and tools

  1. YouTube: Multimodal AI Overview
  2. Sony AI: Multimodal Embodied Attribute Learning by Robots
  3. Nature: Research Article
  4. NVIDIA: Embodied AI Glossary