Improving Robustness via Risk Averse Distributional Reinforcement Learning

Improving Robustness via Risk Averse Distributional Reinforcement Learning
复制标题

通过风险规避分布式强化学习提高鲁棒性

DOI:
--
复制
发表时间:
2020
期刊:
Conference on Learning for Dynamics & Control
影响因子:
--
通讯作者:
Yongxin Chen
Yongxin Chen
中科院分区:
--
文献类型:
--
作者:
Rahul Singh;Qinsheng Zhang;Yongxin Chen

文献摘要

被引文献

相似文献

阻碍强化学习在现实世界应用中取得成功的一个主要障碍是训练策略缺乏鲁棒性,无论是对模型不确定性还是外部干扰。当在模拟而不是真实的世界环境中训练策略时,鲁棒性至关重要。在这项工作中,我们提出了一个风险意识的算法来学习强大的政策,以弥合模拟训练和现实世界的实现之间的差距差距。我们的算法是基于最近发现的分布式RL框架。我们将CVaR风险度量纳入基于样本的分布政策梯度(SDPG)中,以学习风险规避政策,以实现对一系列系统干扰的鲁棒性。我们验证了风险感知SDPG在多个环境中的鲁棒性。
One major obstacle that precludes the success of reinforcement learning in real-world applications is the lack of robustness, either to model uncertainties or external disturbances, of the trained policies. Robustness is critical when the policies are trained in simulations instead of real world environment. In this work, we propose a risk-aware algorithm to learn robust policies in order to bridge the gap between simulation training and real-world implementation. Our algorithm is based on recently discovered distributional RL framework. We incorporate CVaR risk measure in sample based distributional policy gradients (SDPG) for learning risk-averse policies to achieve robustness against a range of system disturbances. We validate the robustness of risk-aware SDPG on multiple environments.