The driver and the engineer: Reinforcement learning and robust control

The driver and the engineer: Reinforcement learning and robust control
复制标题

驾驶员和工程师:强化学习和鲁棒控制

DOI:
10.23919/acc45564.2020.9147347
复制
发表时间:
2020
期刊:
2020 American Control Conference (ACC)
影响因子:
--
通讯作者:
J. Doyle
J. Doyle
中科院分区:
--
文献类型:
--
作者:
Natalie M Bernat;Jiexin Chen;N. Matni;J. Doyle

文献摘要

被引文献

相似文献

强化学习 (RL) 和其他人工智能方法是数据驱动控制设计的令人兴奋的方法,但 RL 强调最大化预期性能,这与鲁棒控制理论 (RCT) 形成鲜明对比,后者重点强调模型不确定性和最坏情况的影响。本文认为,这些方法可能是互补的,大致类似于一级方程式赛车中的车手和工程师的方法。每一个都是不可或缺的,但角色却截然不同。如果 RL 在安全关键型应用中占据主导地位,RCT 仍可能在工厂设计中发挥作用,并且在诊断和减轻由于组件或环境变化或故障而导致性能下降的影响方面也发挥作用。虽然许多 RCT 研究强调控制器的综合,就像 RL 一样,但实际上,RCT 的影响可能已经更大,因为它使用硬限制和鲁棒性能权衡来提供对工厂设计的深入了解,广泛解释为除了核心工厂动态之外,还包括传感器、执行器、通信以及计算机选择和布局。如果我们的系统要更加高效和稳健,更多的自动化最终可能需要更多的严谨性和理论,而不是更少。在这里,我们使用最简单的玩具模型来说明当控制不仅困难而且不可能时,RCT 如何潜在地增强 RL 寻找机制解释,以及使它们更兼容数据驱动的问题。尽管很简单,但问题仍然存在。我们还讨论了这些想法与更现实的挑战的相关性。
Reinforcement learning (RL) and other AI methods are exciting approaches to data-driven control design, but RL's emphasis on maximizing expected performance contrasts with robust control theory (RCT), which puts central emphasis on the impact of model uncertainty and worst case scenarios. This paper argues that these approaches are potentially complementary, roughly analogous to that of a driver and an engineer in, say, formula one racing. Each is indispensable but with radically different roles. If RL takes the driver seat in safety critical applications, RCT may still play a role in plant design, and also in diagnosing and mitigating the effects of performance degradation due to changes or failures in component or environments. While much RCT research emphasizes synthesis of controllers, as does RL, in practice RCT's impact has perhaps already been greater in using hard limits and tradeoffs on robust performance to provide insight into plant design, interpreted broadly as including sensor, actuator, communications, and computer selection and placement in addition to core plant dynamics. More automation may ultimately require more rigor and theory, not less, if our systems are going to be both more efficient and robust. Here we use the simplest possible toy model to illustrate how RCT can potentially augment RL in finding mechanistic explanations when control is not merely hard, but impossible, and issues in making them more compatibly data-driven. Despite the simplicity, questions abound. We also discuss the relevance of these ideas to more realistic challenges.