Loading timeline…
20232020s
Reinforcement Learning from AI Feedback (RLAIF)
Replacing human trainers with AI trainers to scale model alignment securely.
Why It Was Important
As models became vastly smarter than the humans paid to rate their answers (e.g., verifying a PhD-level physics proof), humans could no longer provide accurate reinforcement labels. AI companies used stronger AIs to grade the outputs of weaker AIs, creating an infinitely scalable, automated self-improvement loop.
Who Invented It
Anthropic / OpenAI
Researchers circumventing human evaluation bottlenecks.
Applications
- Constitutional AI
- Scalable Oversight
- Automated code-testing
Key Papers
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Harrison Lee · arXiv · 2023