A modality is a type of data or sensory input an AI can process. Three examples are text, images, and audio (video is another).
It is when an AI system combines information from several modalities instead of just one. This makes it more reliable, because if one type of data is missing or unreliable, it can lean on the others.
In healthcare. IBM Research is building tools that combine medical images (mammography, ultrasound, MRI) with clinical data about the patient to help estimate a diagnosis. The two types of information combined are images and patient clinical data.
It is AI built into physical systems like robots or self-driving cars, which use sensors and computer vision to perceive their surroundings, reason about them, and act in the real world.
Robots need to combine what they read, see, and do to complete tasks. MIT's HiP shows this. It uses a multimodal approach so a robot can plan and carry out a task from language, visual, and action information. So multimodal learning is what lets embodied AI understand and act in the physical world.
MIT News - Multiple AI Models Help Robots Execute Complex Plans