Loading timeline…
20202020s
GPT-3
A 175-billion parameter LLM that proved language models possessed 'emergent' zero-shot capabilities.
Why It Was Important
Scaling up GPT-2 by over 100x proved the 'scaling laws' theory correct: simply making a Transformer larger and feeding it more internet data fundamentally altered its capabilities. GPT-3 could write code probabilistically, translate languages, and answer facts without ever being explicitly programmed to do so, changing the global tech trajectory.
Who Invented It
OpenAI Researchers
A large team executing brutal compute-scaling infrastructure.
Applications
- Zero-shot Prompting
- Generative Writing
- Coding Assistants
Key Papers
- Language Models are Few-Shot Learners
Tom B. Brown · NeurIPS 2020
Videos
How Large Language Models Work
IBM Technology