Loading timeline…
20172010s
Chris Olah
Pioneer in mechanistic interpretability—visualizing the internal workings of neural networks.
Organizations
Google BrainOpenAIAnthropic
Major Achievements
- •Wrote foundational blog posts demystifying complex neural architectures like LSTMs.
- •Pioneered research actively mapping the inner weights of networks to human concepts (Feature Visualization).
- •Founded the interpretability team at Anthropic to diagnose AI behaviors safely.
Key Papers
- Concrete Problems in AI Safety
Dario Amodei et al. · arXiv · 2016
- Zoom In: An Introduction to Circuits
Chris Olah et al. · Distill · 2020