An Improved Trust-Region Method for Off-Policy Deep Reinforcement Learning

An Improved Trust-Region Method for Off-Policy Deep Reinforcement Learning
复制标题

DOI:
10.1109/ijcnn54540.2023.10191837
复制
发表时间:
2023-06
期刊:
2023 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Hepeng Li;Xiangnan Zhong;Haibo He
Hepeng Li;Xiangnan Zhong;Haibo He
中科院分区:
其他
文献类型:
--
作者:
Hepeng Li;Xiangnan Zhong;Haibo He

文献摘要

相似文献

增强学习(RL)是培训代理与复杂环境互动的强大工具。特别是,信任区域方法被广泛用于无模型RL中的策略优化。但是,这些方法由于其上的性质而遭受了较高的样本复杂性,这需要与每个更新的环境进行互动。为了解决这个问题,已经提出了非政策信任区域方法,但是与其他非政策外DRL方法相比,它们在高度连续控制问题中的成功有限。为了提高信任区域策略优化的绩效和样本效率,我们提出了一个非政策信任区域RL算法。我们的算法基于基于信任区域策略优化的封闭式解决方案的理论结果,并有效地优化了复杂的非线性策略。我们证明了算法比先前的信任区域DRL方法的优越性,并表明它在具有联系(Mujoco)环境的多关节动力学(Mujoco)环境中在一系列连续控制任务上取得了出色的性能非政策算法。
Reinforcement learning (RL) is a powerful tool for training agents to interact with complex environments. In particular, trust-region methods are widely used for policy optimization in model-free RL. However, these methods suffer from high sample complexity due to their on-policy nature, which requires interactions with the environment for each update. To address this issue, off-policy trust-region methods have been proposed, but they have shown limited success in highdimensional continuous control problems compared to other off-policy DRL methods. To improve the performance and sample efficiency of trust-region policy optimization, we propose an off-policy trust-region RL algorithm. Our algorithm is based on a theoretical result on a closed-form solution to trust-region policy optimization and is effective in optimizing complex nonlinear policies. We demonstrate the superiority of our algorithm over prior trust-region DRL methods and show that it achieves excellent performance on a range of continuous control tasks in the Multi-Joint dynamics with Contact (MuJoCo) environment, comparable to state-of-the-art off-policy algorithms.