A Tutorial on Derivative-Free Policy Learning Methods for Interpretable Controller Representations

A Tutorial on Derivative-Free Policy Learning Methods for Interpretable Controller Representations
复制标题

DOI:
10.23919/acc55779.2023.10156412
复制
发表时间:
2023-05
期刊:
2023 American Control Conference (ACC)
影响因子:
--
通讯作者:
J. Paulson;Farshud Sorourifar;A. Mesbah
J. Paulson;Farshud Sorourifar;A. Mesbah
中科院分区:
其他
文献类型:
--
作者:
J. Paulson;Farshud Sorourifar;A. Mesbah

文献摘要

相似文献

本文提供了学习复杂系统的控制策略表示的最新进展的教程概述。我们的重点是通过解决一个依赖于当前状态和一些可调参数的优化问题来确定控制策略。我们将这样的策略称为可解释的,因为一旦设置了参数,从业者就可以直接理解每个单独的组件,即目标函数编码期望的目标,约束函数执行系统的规则。我们讨论了如何以这种方式看待各种常用的控制策略,例如线性二次调节器,(非线性)模型预测控制和近似动态规划。传统上,这些控制策略中出现的参数是通过手工、专家知识或简单的试错实验来调优的,这可能非常耗时,并且在实践中会导致次优结果。为此,我们描述了如何使用贝叶斯优化框架,这是一类有效的无导数噪声函数优化方法,可以有效地自动化这一过程。除了回顾相关文献并展示这些新方法在说明性示例问题上的有效性外,我们还对该领域的未来研究提出了展望。
This paper provides a tutorial overview of recent advances in learning control policy representations for complex systems. We focus on control policies that are determined by solving an optimization problem that depends on the current state and some adjustable parameters. We refer to such policies as interpretable in the sense that each of the individual components can be directly understood by practitioners once the parameters are set, i.e., the objective function encodes the desired goal and the constraint functions enforce the rules of the system. We discuss how various commonly used control policies can be viewed in this manner such as the linear quadratic regulator, (nonlinear) model predictive control, and approximate dynamic programming. Traditionally, the parameters that appear in these control policies have been tuned by hand, expert knowledge, or simple trial-and-error experimentation, which can be time consuming and lead to suboptimal results in practice. To this end, we describe how the Bayesian optimization framework, which is a class of efficient derivative-free optimization methods for noisy functions, can be used to efficiently automate this process. In addition to reviewing relevant literature and demonstrating the effectiveness of these new methods on an illustrative example problem, we also offer perspectives on future research in this area.