Multimodal AI
A multimodal AI is a system able to process and combine different types of information —text, images, audio, video— within a single model, instead of handling only one modality. Thanks to this you can, for example, show it a photo and ask about it, have it describe a chart, summarize a video or generate an image from a sentence. It is an important step toward more useful and natural assistants, because people also communicate by mixing words, images and sounds. The latest models from the major providers are multimodal by default. It greatly widens the use cases —accessibility, education, analyzing documents with images— though it inherits the same reliability limits as the rest A REST API is the most common style for building web APIs: it uses web addresses (URLs) and standard HTTP methods to request and send data, usually in JSON format. More in the glossary → of generative AI Generative AI creates new content —text, images, audio or code— from a prompt, instead of only analyzing existing data. More in the glossary → .