Loading timeline…
20202020s
AI Safety & Alignment
The formalized research field ensuring superintelligent systems follow human intent without disastrous side effects.
Why It Was Important
As timelines for AGI sharply decreased, organizations heavily funded mathematical alignment framing. Concepts like 'instrumental convergence' (the idea an AI will gain power as a side effect of achieving its goal) drove billions of dollars into interpretability research, 'red teaming', and containment strategies.
Who Invented It
Alignment Community (Anthropic, Redwood, Alignment Forum)
Philosophers, mathematicians, and engineers focused on existential risk.
Applications
- Superalignment (SSI)
- Model Red-teaming
- Sleeper Agent defense
Key Papers
- Concrete Problems in AI Safety
Dario Amodei et al. · arXiv · 2016