Loading timeline…
20172010s
Reinforcement Learning from Human Feedback (RLHF)
A safety and fine-tuning mechanism for aligning AI behavior with human intent.
Why It Was Important
Pioneered by OpenAI and DeepMind, RLHF trains a 'reward model' based on human rankers choosing the best model outputs safely. It became the critical mechanism to convert chaotic, toxic Next-Word-Predictors into helpful, polite conversational assistants like ChatGPT and Claude.
Who Invented It
Paul Christiano et al.
AI alignment researchers focusing on scalable oversight.
Applications
- ChatGPT
- AI Alignment
- Conversational Agents
Key Papers
- Deep Reinforcement Learning from Human Preferences
Paul F. Christiano et al. · NeurIPS 2017
Videos
Reinforcement Learning from Human Feedback (RLHF) Explained
IBM Technology
New course with Google Cloud: Reinforcement Learning from Human Feedback (RLHF)
DeepLearningAI