Improving the Robustness of Reinforcement Learning Policies With ${\mathcal {L}_{1}}$ Adaptive Control

Improving the Robustness of Reinforcement Learning Policies With ${\mathcal {L}_{1}}$ Adaptive Control
复制标题

${\mathcal {L}_{1}}$自适应控制提高强化学习策略的鲁棒性

DOI:
10.1109/lra.2022.3169309
复制
发表时间:
2021-12
影响因子:
5.2
通讯作者:
Y. Cheng;Penghui Zhao;F. Wang;D. Block;N. Hovakimyan
Y. Cheng;Penghui Zhao;F. Wang;D. Block;N. Hovakimyan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Y. Cheng;Penghui Zhao;F. Wang;D. Block;N. Hovakimyan

文献摘要

相似文献

强化学习(RL)控制策略可能会失败,在一个新的/扰动的环境,这是不同的训练环境,由于动态变化的存在。对于具有连续状态和动作空间的控制系统,我们提出了一种附加方法,通过用${\mathcal {L}_{1}}$自适应控制器(${\mathcal {L}_{1}}$AC)来增强预训练的RL策略。利用${\mathcal {L}_{1}}$AC的快速估计和动态变化的主动补偿能力,所提出的方法可以提高RL策略的鲁棒性,该策略在仿真器或真实的世界中训练,而不考虑广泛的动态变化。数值和真实世界的实验经验证明了所提出的方法在鲁棒RL政策使用无模型和基于模型的方法训练的有效性。
A reinforcement learning (RL) control policy could fail in a new/perturbed environment that is different from the training environment, due to the presence of dynamic variations. For controlling systems with continuous state and action spaces, we propose an add-on approach to robustifying a pre-trained RL policy by augmenting it with an ${\mathcal {L}_{1}}$ adaptive controller (${\mathcal {L}_{1}}$AC). Leveraging the capability of an ${\mathcal {L}_{1}}$AC for fast estimation and active compensation of dynamic variations, the proposed approach can improve the robustness of an RL policy which is trained either in a simulator or in the real world without consideration of a broad class of dynamic variations. Numerical and real-world experiments empirically demonstrate the efficacy of the proposed approach in robustifying RL policies trained using both model-free and model-based methods.