Deep Reinforcement Learning with Double Q-Learning

Deep Reinforcement Learning with Double Q-Learning
复制标题

DOI:
10.1609/aaai.v30i1.10295
复制
发表时间:
2015-09
期刊:
--
影响因子:
--
通讯作者:
H. V. Hasselt;A. Guez;David Silver
H. V. Hasselt;A. Guez;David Silver
中科院分区:
其他
文献类型:
--
作者:
H. V. Hasselt;A. Guez;David Silver

文献摘要

被引文献

相似文献

众所周知,流行的Q学习算法在某些条件下会高估动作值。以前不知道在实践中,这种高估是否常见,它们是否会损害业绩,以及它们是否可以普遍预防。在本文中,我们肯定地回答了所有这些问题。特别是,我们首先展示了最近的DQN算法,它将Q学习与深度神经网络相结合,在Atari 2600领域的一些游戏中存在严重的高估。然后,我们证明了在表格设置中引入的双Q学习算法背后的思想可以推广到大规模函数逼近。我们提出了一个具体的适应DQN算法,并表明,由此产生的算法不仅减少了观察到的高估,假设,但这也导致更好的性能在几个游戏。
The popular Q-learning algorithm is known to overestimate action values under certain conditions. It was not previously known whether, in practice, such overestimations are common, whether they harm performance, and whether they can generally be prevented. In this paper, we answer all these questions affirmatively. In particular, we first show that the recent DQN algorithm, which combines Q-learning with a deep neural network, suffers from substantial overestimations in some games in the Atari 2600 domain. We then show that the idea behind the Double Q-learning algorithm, which was introduced in a tabular setting, can be generalized to work with large-scale function approximation. We propose a specific adaptation to the DQN algorithm and show that the resulting algorithm not only reduces the observed overestimations, as hypothesized, but that this also leads to much better performance on several games.