Geometric Value Iteration: Dynamic Error-Aware KL Regularization for Reinforcement Learning

Geometric Value Iteration: Dynamic Error-Aware KL Regularization for Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Toshinori Kitamura;Lingwei Zhu;Takamitsu Matsubara
Toshinori Kitamura;Lingwei Zhu;Takamitsu Matsubara
中科院分区:
其他
文献类型:
--
作者:
Toshinori Kitamura;Lingwei Zhu;Takamitsu Matsubara

文献摘要

相似文献

最近关于熵正则化强化学习(RL)方法的文献的繁荣表明,Kullback-Leibler(KL)正则化通过在温和的假设下消除错误而为RL算法带来了优势。然而,现有的分析集中在固定的正则化与一个恒定的权重系数,并没有考虑的情况下,允许系数动态变化。本文研究了动态系数格式,给出了其第一渐近误差界。基于动态系数误差界,我们提出了一种有效的方案来调整系数根据误差的大小,有利于更强大的学习。为了补充这一发展,我们提出了一种新的算法,几何值迭代(GVI),具有动态的错误感知KL系数设计的目的是减轻错误对性能的影响。我们的实验表明,GVI可以有效地利用学习速度和鲁棒性之间的平衡,均匀平均一个常数KL系数。GVI和深度网络的组合即使在没有目标网络的情况下也表现出稳定的学习行为,其中具有恒定KL系数的算法将大幅振荡甚至无法收敛。
The recent boom in the literature on entropy-regularized reinforcement learning (RL) approaches reveals that Kullback-Leibler (KL) regularization brings advantages to RL algorithms by canceling out errors under mild assumptions. However, existing analyses focus on fixed regularization with a constant weighting coefficient and do not consider cases where the coefficient is allowed to change dynamically. In this paper, we study the dynamic coefficient scheme and present the first asymptotic error bound. Based on the dynamic coefficient error bound, we propose an effective scheme to tune the coefficient according to the magnitude of error in favor of more robust learning. Complementing this development, we propose a novel algorithm, Geometric Value Iteration (GVI), that features a dynamic error-aware KL coefficient design with the aim of mitigating the impact of errors on performance. Our experiments demonstrate that GVI can effectively exploit the trade-off between learning speed and robustness over uniform averaging of a constant KL coefficient. The combination of GVI and deep networks shows stable learning behavior even in the absence of a target network, where algorithms with a constant KL coefficient would greatly oscillate or even fail to converge.