Accelerating Deep Q Network by Weighting Experiences

Accelerating Deep Q Network by Weighting Experiences
复制标题

通过加权经验加速 Deep Q 网络

DOI:
10.1007/978-3-030-04167-0_19
复制
发表时间:
2018
期刊:
Proceedings of the 25th International Conference on Neural Information
影响因子:
--
通讯作者:
and
and
中科院分区:
--
文献类型:
--
作者:
Kazuhiro Murakami;Koichi Moriyama;Atsuko Mutoh;Tohgoroh Matsui;and

文献摘要

相似文献

深度Q网络(Deep Q Network,DQN)是一种强化学习方法,它使用深度神经网络来近似Q函数。文献表明,DQN可以选择比人类更好的反应。然而,DQN需要很长的时间来学习适当的动作,通过使用从其记忆中采样的状态、动作、奖励和下一状态的元组,称为“经验”。DQN对它们进行均匀和随机的采样,但是经验是倾斜的,导致学习缓慢,因为频繁的经验被冗余地采样,而不频繁的经验则没有。这项工作减轻了问题的加权经验的基础上,他们的频率和操纵他们的抽样概率。在视频游戏环境中,所提出的方法比DQN更快地学习适当的响应。
Deep Q Network (DQN) is a reinforcement learning methodlogy that uses deep neural networks to approximate the Q-function. Literature reveals that DQN can select better responses than humans. However, DQN requires a lengthy period of time to learn the appropriate actions by using tuples of state, action, reward and next state, called “experience”, sampled from its memory. DQN samples them uniformly and randomly, but the experiences are skewed resulting in slow learning because frequent experiences are redundantly sampled but infrequent ones are not. This work mitigates the problem by weighting experiences based on their frequency and manipulating their sampling probability. In a video game environment, the proposed method learned the appropriate responses faster than DQN.