multimodal learning and embodied AI questions
1. What is a modality in AI? Give three examples of different modalities.
Answer: Modality is a form of data that an AI can process. Three examples: visual (e.g. pictures and videos), textual (like a book or an article), and audio (usually speech for translations and summaries).
2. What is multimodal learning?
Answer: Multimodal learning is when an AI model is trained to process different types of data at the same time. An AI that went through multimodal learning can, for example, it can use both pictures and text to generate output.
3. Where is multimodal learning used? Can you find one real application and identify the types of information it combines?
Answer: One application of multimodal learning in AI is a feature called Google Lens. In its current version, it combines the textual information (question from the user, info from the web) with visual information (image captured by the camera) and can answer the prompt based on the image or perform a visual web search based on the prompt.
4. What is embodied AI?
Answer: Embodied AI are AI systems that are integrated with a physical "body", like robots and self-driving cars. Famous examples include Boston Dynamics robots, and Waymos.
5. How are multimodal learning and embodied AI connected?
Answer: They are connected because embodied AI requires the system to be able to accept inputs from different modalities to operate. For example, a Waymo cannot drive based on textual information. It has to continuously process visual and sensory data to navigate the environment and monitor the state and location of the car.
references:
- https://medium.com/@ganeshkannappan/modalities-in-ai-understanding-the-power-of-diverse-data-7de99ddb9343
- https://www.v7labs.com/blog/multimodal-deep-learning-guide
- https://www.gmicloud.ai/en/glossary/modality
- https://www.nvidia.com/en-us/glossary/embodied-ai/
- https://www.ionos.com/digitalguide/websites/web-development/embodied-ai/