Loading timeline…
20232020s
Multimodal Foundation Models
Models trained jointly on text, images, video, and audio data streams.
Why It Was Important
Instead of having completely separate models for language and vision, 2020s architectures integrated them deeply from the initial training loss. This allowed models to natively understand a video, transcribe the audio, identify the actors, and output a Python script summarizing the plot simultaneously.
Who Invented It
Google (Gemini), OpenAI (GPT-4V), Meta
The premier AI conglomerates.
Applications
- Universal Assistants
- Medical diagnostics (MRI + Notes)
- Robotics perception