Multisource Transfer Double DQN Based on Actor Learning

Multisource Transfer Double DQN Based on Actor Learning
复制标题

DOI:
10.1109/tnnls.2018.2806087
复制
发表时间:
2018-03
影响因子:
10.4
通讯作者:
Jie Pan;X. Wang;Yuhu Cheng;Qiang Yu
Jie Pan;X. Wang;Yuhu Cheng;Qiang Yu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jie Pan;X. Wang;Yuhu Cheng;Qiang Yu

文献摘要

被引文献

相似文献

深度强化学习(Deep reinforcement learning, RL)综合运用了强化学习中“试错”和“奖惩”的心理机制,以及深度学习中强大的特征表达和非线性映射。目前,它在人工智能和机器学习领域发挥着至关重要的作用。由于RL agent需要不断地与周围环境相互作用,深度Q网络(deep Q network, DQN)不可避免地需要学习大量的网络参数,这导致学习效率较低。本文提出了一种基于行动者学习的多源迁移双DQN (MTDDQN)。将迁移学习技术与深度强化学习相结合,使强化学习代理(RL agent)收集、总结、迁移包括策略模拟、特征回归在内的行动知识,用于相关任务的训练。DQN中存在动作过估计,即最大Q值对应的动作的下概率极限不为零。因此,采用双DQN训练传递网络,消除动作高估引起的误差积累。此外,为了避免负迁移,即确保源任务和目标任务之间的强相关性,采用了多源迁移学习机制。在游戏厅学习环境平台上对Atari2600游戏进行测试,通过与DQN和双DQN等主流方法进行比较,评估MTDDQN的可行性和性能。实验证明,MTDDQN不仅实现了类人actor学习迁移能力,而且在目标任务上达到了预期的学习效率和测试精度。
Deep reinforcement learning (RL) comprehensively uses the psychological mechanisms of “trial and error” and “reward and punishment” in RL as well as powerful feature expression and nonlinear mapping in deep learning. Currently, it plays an essential role in the fields of artificial intelligence and machine learning. Since an RL agent needs to constantly interact with its surroundings, the deep Q network (DQN) is inevitably faced with the need to learn numerous network parameters, which results in low learning efficiency. In this paper, a multisource transfer double DQN (MTDDQN) based on actor learning is proposed. The transfer learning technique is integrated with deep RL to make the RL agent collect, summarize, and transfer action knowledge, including policy mimic and feature regression, to the training of related tasks. There exists action overestimation in DQN, i.e., the lower probability limit of action corresponding to the maximum Q value is nonzero. Therefore, the transfer network is trained by using double DQN to eliminate the error accumulation caused by action overestimation. In addition, to avoid negative transfer, i.e., to ensure strong correlations between source and target tasks, a multisource transfer learning mechanism is applied. The Atari2600 game is tested on the arcade learning environment platform to evaluate the feasibility and performance of MTDDQN by comparing it with some mainstream approaches, such as DQN and double DQN. Experiments prove that MTDDQN achieves not only human-like actor learning transfer capability, but also the desired learning efficiency and testing accuracy on target task.