A Maximum Divergence Approach to Optimal Policy in Deep Reinforcement Learning

A Maximum Divergence Approach to Optimal Policy in Deep Reinforcement Learning
复制标题

DOI:
10.1109/tcyb.2021.3104612
复制
发表时间:
2021-09
影响因子:
11.8
通讯作者:
Zhi-Xuan Yang;Hong Qu;Mingsheng Fu;Wang Hu;Yongze Zhao
Zhi-Xuan Yang;Hong Qu;Mingsheng Fu;Wang Hu;Yongze Zhao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhi-Xuan Yang;Hong Qu;Mingsheng Fu;Wang Hu;Yongze Zhao

文献摘要

相似文献

基于熵正则化的无模型强化学习算法在控制任务中取得了良好的性能。这些算法考虑使用策略的熵正则化项来学习随机策略。这项工作提供了一个新的视角,旨在明确地学习表示的内在信息的状态转换,以获得多模态随机政策,用于处理探索和开发之间的权衡。研究了一类具有最大散度的马尔可夫决策过程,称为散度马尔可夫决策过程。发散MDP的目标是找到一个最优的随机策略,最大化的期望折扣总回报和发散项的总和,其中发散函数学习状态转移的隐式信息。因此,它可以提供更好的小康随机策略,以提高在高维连续设置的鲁棒性和性能。在此框架下,得到最优性方程,并基于发散策略迭代方法提出了一种求解大规模连续问题的发散actor-critic算法.实验结果表明,与其他方法相比,该方法具有更好的性能和鲁棒性,特别是在复杂的环境中。DivAC的代码可以在https://github.com/yzyvl/DivAC中找到。
Model-free reinforcement learning algorithms based on entropy regularized have achieved good performance in control tasks. Those algorithms consider using the entropy-regularized term for the policy to learn a stochastic policy. This work provides a new perspective that aims to explicitly learn a representation of intrinsic information in state transition to obtain a multimodal stochastic policy, for dealing with the tradeoff between exploration and exploitation. We study a class of Markov decision processes (MDPs) with divergence maximization, called divergence MDPs. The goal of the divergence MDPs is to find an optimal stochastic policy that maximizes the sum of both the expected discounted total rewards and a divergence term, where the divergence function learns the implicit information of state transition. Thus, it can provide better-off stochastic policies to improve both in robustness and performance in a high-dimension continuous setting. Under this framework, the optimality equations can be obtained, and then a divergence actor–critic algorithm is developed based on the divergence policy iteration method to address large-scale continuous problems. The experimental results, compared to other methods, show that our approach achieved better performance and robustness in the complex environment particularly. The code of DivAC can be found in https://github.com/yzyvl/DivAC.