Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation

Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
arXiv: Learning
影响因子:
--
通讯作者:
Niels Justesen;R. Torrado;Philip Bontrager;A. Khalifa;J. Togelius;S. Risi
Niels Justesen;R. Torrado;Philip Bontrager;A. Khalifa;J. Togelius;S. Risi
中科院分区:
其他
文献类型:
--
作者:
Niels Justesen;R. Torrado;Philip Bontrager;A. Khalifa;J. Togelius;S. Risi

文献摘要

被引文献

相似文献

深度强化学习(RL)在许多领域都表现出了令人印象深刻的结果,直接从高维感觉流中学习。然而,当神经网络在固定的环境中训练时,例如视频游戏中的单个级别,它们通常会过拟合,并且无法泛化到新的级别。当RL模型过拟合时,即使对环境进行轻微的修改也会导致代理性能低下。本文探讨了如何在训练过程中程序生成的水平可以提高通用性。我们发现,对于一些游戏程序水平的生成,使泛化到新的水平在同一分布。此外,通过根据代理的性能操纵级别的难度,可以用更少的数据实现更好的性能。学习行为的一般性也在一组人类设计的水平上进行评估。结果表明,人类设计的水平高度依赖于水平生成器的设计,概括的能力。我们应用降维和聚类技术来可视化生成器的水平分布,并分析它们在多大程度上可以产生与人类设计的水平相似的水平。
Deep reinforcement learning (RL) has shown impressive results in a variety of domains, learning directly from high-dimensional sensory streams. However, when neural networks are trained in a fixed environment, such as a single level in a video game, they will usually overfit and fail to generalize to new levels. When RL models overfit, even slight modifications to the environment can result in poor agent performance. This paper explores how procedurally generated levels during training can increase generality. We show that for some games procedural level generation enables generalization to new levels within the same distribution. Additionally, it is possible to achieve better performance with less data by manipulating the difficulty of the levels in response to the performance of the agent. The generality of the learned behaviors is also evaluated on a set of human-designed levels. The results suggest that the ability to generalize to human-designed levels highly depends on the design of the level generators. We apply dimensionality reduction and clustering techniques to visualize the generators' distributions of levels and analyze to what degree they can produce levels similar to those designed by a human.