Loading timeline…
20222020s
Mixture of Experts (MoE)
A sparse network architecture selectively activating only specialized sub-networks per token.
Why It Was Important
Training massive dense models became computationally prohibitive. MoE (heavily rumored to be the backbone of GPT-4 and confirmed in Mistral's Mixtral) routes incoming text to various specialized 'expert' neural networks depending on context. This approach massively increases parameter counts without dramatically slowing inference speeds.
Who Invented It
Google Shazeer et al. (Revitalized in 2020s scale)
Hardware/software co-designers chasing supreme efficiency.
Applications
- GPT-4 Architecture
- Mixtral 8x7B
- High-efficiency scaling
Key Papers
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Noam Shazeer et al. · ICLR 2017
Videos
What is Mixture of Experts?
IBM Technology