Loading timeline…
20172010s
Paul Christiano
Alignment researcher who developed reinforcement learning from human feedback, the technique that made chat models usable.
Organizations
OpenAIAlignment Research CenterUS AI Safety Institute
Major Achievements
- •Led the 2017 work on deep reinforcement learning from human preferences, establishing RLHF.
- •RLHF became the method behind InstructGPT and ChatGPT, turning raw language models into instruction-followers.
- •Founded the Alignment Research Center to work on scalable oversight of advanced AI systems.
- •Appointed head of AI safety at the US AI Safety Institute.
Key Papers
- Deep Reinforcement Learning from Human Preferences
Paul F. Christiano et al. · NeurIPS 2017
- Concrete Problems in AI Safety
Dario Amodei et al. · arXiv · 2016