Loading timeline…
20222020s
Constitutional AI
A mechanism for training helpful and harmless AI via a set of written principles.
Why It Was Important
Pioneered by Anthropic (creators of Claude), this method relies on RLAIF (Reinforcement Learning from AI Feedback). Instead of relying strictly on human labelers (who carry bias and tire quickly), another robust AI grades thousands of outputs continually against a static 'Constitution' of rules, vastly scaling model alignment safety.
Who Invented It
Anthropic Research Team
Former OpenAI safety researchers who splintered to focus on highly aligned models.
Applications
- Claude Model Alignment
- Safe Corporate AI
- Bias minimization
Key Papers
- Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai · arXiv · 2022