Reinforcement learning and control theory are two adjacent scientific fields
that focus on optimizing the controller of unknown dynamical systems using feedback. While both fields have common roots in dynamic programming, they have evolved with distinct methodologies, goals, and cultures. Despite decades of mutual influence, a significant gap persists between the two communities.
This tutorial introduces adaptive control, actor-critic reinforcement algorithms, and a new original way to combine these two paradigms for data-driven decision making on a classical locomotion control problem. Our aim is to provide keys to understand the core differences between both approaches, and insights to help experts in each field better understand and engage with the tools and approaches of the other.
Claire Vernade is Full Professor of Foundations of Machine Learning at the University of Technology Nuremberg (UTN). Her research focuses on sequential decision making under uncertainty, spanning reinforcement learning, online learning, and statistical machine learning. She develops theoretically grounded learning and decision-making algorithms for adaptive and interactive systems, with an emphasis on bridging mathematical foundations and scalable machine learning methods. Before joining UTN in 2025, she was a Group Leader at the University of Tübingen and a Senior Research Scientist at Google DeepMind. She is the recipient of an Emmy Noether Programme grant and an ERC Starting Grant.
Onno Eberhard is a Ph.D. student in Computer Science at the Max Planck
Institute for Intelligent Systems and the University of Tübingen. He holds an M.Sc. in Machine Learning from the University of Tübingen and a B.Sc. in Electrical Engineering from the University of Duisburg-Essen, and has gained professional experience at Google Research and Siemens. His research focuses on the theoretical foundations of reinforcement learning, particularly regarding
partially observable environments, recurrent memory, and its intersections with control theory.
Martha White is a Professor of Computing Science at the University of Alberta and a Fellow of Amii, which is one of the top machine learning centres in the world. She holds a Canada CIFAR AI Chair, a Tier 2 Canada Research Chair in Reinforcement Learning, received IEEE’s “AIs 10 to Watch: The Future of AI” award in 2020 and was inducted into the College of New Scholars by the Royal Society of Canada in 2024. She has authored more than 80 papers in top journals and conferences. Martha is an associate editor for JMLR and TMLR, server on the RLC board and has served as co-program chair for ICLR and for RLC. Her research focus is on developing reinforcement learning algorithms that learn to adapt continually, with a focus on process control and more sustainable systems.
Florian Dörfler is a Professor at the Automatic Control Laboratory at ETH Zürich. He received his Ph.D. degree in Mechanical Engineering from the University of California at Santa Barbara in 2013. From 2013 to 2014 he was an Assistant Professor at the University of California Los Angeles. His research interests are centered around automatic control, system theory, optimization, and learning. His particular foci are on network systems, data-driven settings, and applications to power systems. He is a recipient of the Rössler Prize, the highest scientific award at ETH Zürich across all disciplines, as well as the distinguished career awards by IFAC (Manfred Thoma Medal) and EUCA (European Control Award). He and his team received best paper distinctions in the top venues of control, power systems, power electronics, circuits and systems.
Csaba Szepesvári is a Canada CIFAR AI Chair, Professor of Computing Science
at the University of Alberta, and Team Lead for the Foundations team at DeepMind. He is a Fellow of the Association for the Advancement of Artificial Intelligence and an IEEE Fellow. He received his PhD in 1999 from József Attila University in Szeged, Hungary, in probability and statistics. His research
advances the foundations of learning-based artificial intelligence, especially reinforcement learning, bandit algorithms, and planning under uncertainty. He is the author or co-author of three books, including Bandit Algorithms, published by Cambridge University Press in 2020. He is also the co-inventor of UCT, an algorithm that helped ignite the modern development of Monte Carlo tree search and made simulation-based planning a central tool in game AI and decision making under uncertainty.
Miroslav Krstic is a professor and serves as senior associate vice chancellor for research at UC San Diego. He is the recipient of the IEEE Brockett Award and Bode Prize, ASME Oldenburger Medal, SIAM Reid Prize, Bellman Award, and other recognitions, including the Chestnut prize and several IFAC TC awards. He is a member of the Serbian Academy of Sciences and Arts, Academia Europaea, fellow of IEEE, IFAC, SIAM, ASME, AIAA, and other societies, and Fellow-Ambassador of CNRS. Krstic is the current editor-in-chief of IEEE Transactions on Automatic Control, a former EiC of Systems & Control Letters, and former senior editor in Automatica. He is a coauthor of 19 books and several hundred papers on various nonlinear, adaptive, and infinite-dimensional control subjects.
Michael Mühlebach leads the research group learning and dynamical systems at the Max Planck Institute for Intelligent Systems in Tübingen, Germany. His group conducts fundamental research in machine learning, reinforcement learning, and large-scale optimization. He won numerous awards including an Emmy Noether and Branco Weiss fellowship, as well as an ETH Medal and the HILTI prize for innovative research. He is also a member of the editorial board of Foundations and Trends in Machine Learning.