Loading timeline…
20172010s
John Schulman
OpenAI co-founder who invented PPO, the workhorse algorithm of deep reinforcement learning, and co-led the RLHF training behind ChatGPT.
Organizations
OpenAIUC BerkeleyAnthropicThinking Machines Lab
Major Achievements
- •Invented Trust Region Policy Optimization (TRPO, 2015) and Proximal Policy Optimization (PPO, 2017), building on his PhD work under Pieter Abbeel at UC Berkeley.
- •Co-founded OpenAI in 2015 as one of its founding research scientists.
- •Co-led the creation of ChatGPT and its reinforcement learning from human feedback (RLHF) training pipeline.
- •Continued alignment research at Anthropic before becoming chief scientist of Thinking Machines Lab.
Key Papers
- Proximal Policy Optimization Algorithms
John Schulman et al. · arXiv · 2017
- Trust Region Policy Optimization
John Schulman et al. · ICML 2015
- Training Language Models to Follow Instructions with Human Feedback
Long Ouyang · NeurIPS 2022