← All weekly research updates

Multimodal Learning and Embodied AI

1. What is a modality in AI? Give three examples of different modalities.

A modality is a type of data that an AI system can understand and process. Texts, images, audio, video are examples of modalities, which are different types of data. Modalities resemble our different senses that we use to understand the world.

GMI Cloud. (n.d.). Modality. https://www.gmicloud.ai/en/glossary/modality

2. What is multimodal learning?

In education, multimodal learning is the simultaneous use of various teaching methods by an instructor. Similarly, in machine learning, multimodal learning is an AI modal’s learning from and processing of different modalities. It may involve reinfocement learning to match or relate phenomena observed in different data types, such as matching the textual representation of a cup to an image of a cup. Multimodal learning has challenges including the representation and alignment of different data types, which are listed in Tutorial on MultiModal Machine Learning by Liang and Morency.

Baltrušaitis, T., Ahuja, C., & Morency, L.-P. (2017). Multimodal machine learning: A survey and taxonomy. arXiv. https://doi.org/10.48550/arXiv.1705.09406

Do, H. (2023, December 16). Introduction to multimodal learning. Medium. https://medium.com/@hannah.hj.do/introduction-to-multimodal-learning-3226ed49d577

Liang, P., & Morency, L.-P. (2023a). Tutorial on multimodal machine learning [Presentation slides]. ICML 2023. https://drive.google.com/file/d/1qIYBuYrSW2-e95DL7LndfLFqGkIWFG21/view

3. Where is multimodal learning used? Can you find one real application and identify the types of information it combines?

Multimodal learning is used in training AI chatbots, such as ChatGPT, where you can train the modal with images and written descriptions about those images. Here, visual and textual information are combined and related. As a result, if we upload a photo from our room and ask the modal to identify the blue objects within the photo, it can create a response based on those two modalities. Another example of the usage of multimodal training is a self-driving vehicle, which combines the textual information from the traffic signs with visual information from the environment it is in, acting accordingly.

4. What is embodied AI?

Embodied AI refers to a physical system that is controlled by AI to interact with the real world. Humanoid robots and autonomous vehicles are examples of it. They may use cameras, audio sensors, LiDAR (usage of laser light pulses to detect distances and map 3D objects), and tactile sensors to perceive the real world and make decisions accordingly.

National Oceanic and Atmospheric Administration. (2026, September 23). What is lidar? National Ocean Service. https://oceanservice.noaa.gov/facts/lidar.html

NVIDIA. (n.d.). What is embodied AI? https://www.nvidia.com/en-us/glossary/embodied-ai/

5. How are multimodal learning and embodied AI connected?

Embodied AI systems use multimodal learning to operate. They make a decision based on the visuals, texts, audios, and videos used in its training, as well as those it encounters in action, learning through their own experiences. For example, a humanoid robot may be trained by videos of robots holding a cup, and then it may be trained by practicing holding a cup, where it combines data from its visual sensors to identify the object and uses its tactile sensors to learn the weight of the cup, friction on the cup’s surface, and resulting pressure when it holds the cup. Also, in order for it to identify the cup, it may need to previously be trained by matching the sound of the cup (if it is instructed by speech) or the textual representation of the cup (if it is instructed by text) to the visual representation of the cup. Therefore, assessing and learning information from multiple modalities is at the core of embodied AI.