Obstacle Tower: A Generalization Challenge in Vision, Control, and Planning

Obstacle Tower: A Generalization Challenge in Vision, Control, and Planning
复制标题

DOI:
10.24963/ijcai.2019/373
复制
发表时间:
2019-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Arthur Juliani;A. Khalifa;Vincent-Pierre Berges;Jonathan Harper;Hunter Henry;A. Crespi;J. Togelius;Danny Lange
Arthur Juliani;A. Khalifa;Vincent-Pierre Berges;Jonathan Harper;Hunter Henry;A. Crespi;J. Togelius;Danny Lange
中科院分区:
其他
文献类型:
--
作者:
Arthur Juliani;A. Khalifa;Vincent-Pierre Berges;Jonathan Harper;Hunter Henry;A. Crespi;J. Togelius;Danny Lange

文献摘要

被引文献

相似文献

最近人工智能研究的快速发展部分是由于快速和具有挑战性的模拟环境的存在。这些环境通常采用游戏的形式;任务范围从简单的棋盘游戏到竞争性的电子游戏。我们提出了一个新的基准-障碍塔:一个高保真,3D,第三人称,程序生成的环境。障碍物塔中的智能体必须学会同时解决低级控制和高级规划问题,同时从像素和稀疏奖励信号中学习。与Arcade Learning Environment等其他基准不同,《Obstacle Tower》中对代理表现的评估是基于代理在未知环境中表现良好的能力。在本文中,我们概述了环境,并提供了一组由当前最先进的深度强化学习方法以及人类玩家产生的基线结果。这些算法无法产生能够执行接近人类水平的代理。
The rapid pace of recent research in AI has been driven in part by the presence of fast and challenging simulation environments. These environments often take the form of games; with tasks ranging from simple board games, to competitive video games. We propose a new benchmark - Obstacle Tower: a high fidelity, 3D, 3rd person, procedurally generated environment. An agent in Obstacle Tower must learn to solve both low-level control and high-level planning problems in tandem while learning from pixels and a sparse reward signal. Unlike other benchmarks such as the Arcade Learning Environment, evaluation of agent performance in Obstacle Tower is based on an agent's ability to perform well on unseen instances of the environment. In this paper we outline the environment and provide a set of baseline results produced by current state-of-the-art Deep RL methods as well as human players. These algorithms fail to produce agents capable of performing near human level.