Human-level control through deep reinforcement learning

Human-level control through deep reinforcement learning
复制标题

DOI:
10.1038/nature14236
复制
发表时间:
2015-02-26
期刊:
影响因子:
64.8
通讯作者:
Hassabis, Demis
Hassabis, Demis
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Mnih, Volodymyr;Kavukcuoglu, Koray;Hassabis, Demis

文献摘要

被引文献

相似文献

强化学习的理论提供了一个规范的帐户,深深植根于心理学和神经科学的角度对动物行为,代理如何优化他们对环境的控制。然而,为了在接近现实世界复杂性的情况下成功地使用强化学习,智能体面临着一项艰巨的任务:他们必须从高维感官输入中获得环境的有效表示,并使用这些来将过去的经验推广到新的情况。值得注意的是,人类和其他动物似乎通过强化学习和分层感觉处理系统的和谐结合来解决这个问题4'5,大量神经数据证明了前者,揭示了多巴胺能神经元发出的相位信号与时间之间的显着相似之处差异强化学习算法”。虽然强化学习代理在各种领域取得了一些成功,但它们的适用性以前仅限于可以手工制作有用特征的领域,或者具有完全观察到的低维状态空间的领域。在这里,我们使用训练深度神经网络的最新进展来开发一种新型人工智能体,称为深度Q网络,它可以使用端到端强化学习直接从高维感官输入中学习成功的策略。我们在经典Atari 2600游戏的挑战领域测试了这个代理。我们证明了深度Q网络代理,只接收像素和游戏分数作为输入,能够超越所有以前的算法的性能,并在49场比赛中达到与专业人类游戏测试人员相当的水平,使用相同的算法,网络架构和超参数。这项工作弥合了高维感官输入和动作之间的鸿沟,从而产生了第一个能够学习擅长各种挑战性任务的人工智能体。
The theory of reinforcement learning provides a normative account', deeply rooted in psychological' and neuroscientifie perspectives on animal behaviour, of how agents may optimize their control of an environment. To use reinforcement learning successfully in situations approaching real-world complexity, however, agents are confronted with a difficult task: they must derive efficient representations of the environment from high-dimensional sensory inputs, and use these to generalize past experience to new situations. Remarkably, humans and other animals seem to solve this problem through a harmonious combination of reinforcement learning and hierarchical sensory processing systems4'5, the former evidenced by a wealth of neural data revealing notable parallels between the phasic signals emitted by dopaminergic neurons and temporal difference reinforcement learning algorithms'. While reinforcement learning agents have achieved some successes in a variety of domains", their applicability has previously been limited to domains in which useful features can be handcrafted, or to domains with fully observed, low-dimensional state spaces. Here we use recent advances in training deep neural networks'" to develop a novel artificial agent, termed a deep Q-network, that can learn successful policies directly from high-dimensional sensory inputs using end-to-end reinforcement learning. We tested this agent on the challenging domain of classic Atari 2600 games". We demonstrate that the deep Q-network agent, receiving only the pixels and the game score as inputs, was able to surpass the performance of all previous algorithms and achieve a level comparable to that of a professional human games tester across a set of 49 games, using the same algorithm, network architecture and hyperparameters. This work bridges the divide between high-dimensional sensory inputs and actions, resulting in the first artificial agent that is capable of learning to excel at a diverse array of challenging tasks.