Loading timeline…
20062000s
Csaba Szepesvári
Reinforcement learning theorist, co-author of UCT and of the bandit results modern exploration rests on.
Organizations
University of AlbertaGoogle DeepMindInstitute for Computer Science and Control (SZTAKI), Budapest
Major Achievements
- •Co-introduced UCT (2006) with Levente Kocsis, giving Monte Carlo tree search its selection rule.
- •Wrote 'Algorithms for Reinforcement Learning', a standard concise reference for the field.
- •Contributed foundational theory on bandit problems and the exploration–exploitation trade-off.
- •Continued this work as a research scientist at DeepMind alongside his Alberta professorship.
Key Papers
- Bandit Based Monte-Carlo Planning
Levente Kocsis, Csaba Szepesvári · ECML 2006