Research Update — Week 6
A modality is one type or channel of information that an AI system receives, represents, or uses. Humans use modalities too - we see with our eyes, hear with our ears, and feel through touch. In AI, these are three examples:
Multimodal learning is a type of machine learning in which an AI model learns from, combines, and connects information from two or more modalities. Instead of learning only from text or only from images, it uses the strengths of several information types together. An example of how combining modalities is useful:
Multimodal learning is used in search engines, voice assistants, medical AI, autonomous vehicles, accessibility tools, online content moderation, robots, and generative AI systems.
Real application: Google Lens
A user can take or upload a photo and add text to the search—for instance, take a picture of a green dress and type "in black," or photograph an object and ask what it is. Essentially, it combines images, text queries, and web information to supply relevant results. This combines visual and text modalities.
Embodied AI is AI that has a body (physical or simulated) and can interact with an environment. The body could be a robot arm, robot, robot, drone, autonomous car, or a character in a virtual world. Unlike an ordinary text-only chatbot, embodied AI does not only analyze information on a screen. It must cope with physical space, movement, uncertainty, objects, other people, and the consequences of its actions.
Multimodal learning is often a major part of embodied AI because a robot needs information from several senses to understand its surroundings. It may combine:
Multimodal learning helps the system create a richer understanding of the environment. Embodied AI uses that understanding to make decisions and take physical actions.
← Previous: Week 5 — Distributed Systems · Next: Week 7 — What is AI? →