← back to main site
Research Update

Week 5 — Multimodal Learning and Embodied AI

Topic presented by Dr. Bilal Taha · Baraa Nsour
Q1

What is a modality in AI? Give three examples of different modalities.

Modalities are types of data which a machine can learn from. Different modalities represent different ways of coding. Examples of these include text, images, audio, and more.

Q2

What is multimodal learning?

Multimodal learning is using multiple forms of data input to train an AI such that its outputs are more comprehensive.

Q3

Where is multimodal learning used? Can you find one real application and identify the types of information it combines?

Multimodal learning is used in healthcare, for instance, to analyze medical images. It is used when AI must be trained on several different data input types.

GPT-4o was the introduction of multimodal capabilities into ChatGPT.

Q4

What is embodied AI?

It is a branch of AI which refers to the integration of AI into the physical world, and physical systems, such as humanoid robots.

Q5

How are multimodal learning and embodied AI connected?

The multimodal learning, from the slight bit I understand of it, allows the AI to cognitively perceive the world around it and understand it, while the embodied AI allows the AI to actually ACT physically.

Further reading

Modality GMI Cloud glossary — what counts as a modality and how models take each one in. (Q1) What is multimodal AI? IBM — how models trained on several kinds of input differ from single-input ones. (Q2) Embodied AI NVIDIA glossary — AI that senses and acts through a physical body, such as a robot. (Q4)