Learning to Plan from Raw Data in Grid-based Games

Learning to Plan from Raw Data in Grid-based Games
复制标题

在基于网格的游戏中学习根据原始数据进行规划

DOI:
--
复制
发表时间:
2018
期刊:
Global Conference on Artificial Intelligence
影响因子:
--
通讯作者:
O. Winther
O. Winther
中科院分区:
--
文献类型:
--
作者:
Andrea Dittadi;Thomas Bolander;O. Winther

文献摘要

被引文献

相似文献

一个自主学习在其环境中行动的智能体必须获取领域动态的模型。这可能是一项具有挑战性的任务,尤其是在现实世界的领域中,那里的观测数据是高维且有噪声的。尽管在自动规划中动态通常是给定的,但也有一些动作模式学习方法,它们学习符号规则(例如STRIPS或PDDL)以供传统规划器使用。然而,这些算法依赖于对环境观测的逻辑描述。相比之下,用于游戏的深度强化学习的最新方法从像素观测中学习。然而,它们通常不会获取环境模型,而是一个用于单步动作选择的策略。即使学习到了一个模型,它也无法推广到训练领域中未见过的实例。在这里,我们提出一种基于神经网络的方法,该方法从视觉观测中学习领域动态的一种近似、紧凑、隐式表示,它可用于使用标准搜索算法进行规划,并能推广到新的领域实例。所学习的模型由子模块组成,每个子模块都隐式地表示传统意义上的一个动作模式。我们在标准领域“推箱子”的视觉版本上评估我们的方法,并表明通过在一个单一实例上进行训练,它学习到一个转换模型,该模型可成功用于解决游戏的新关卡。
An agent that autonomously learns to act in its environment must acquire a model of the domain dynamics. This can be a challenging task, especially in real-world domains, where observations are high-dimensional and noisy. Although in automated planning the dynamics are typically given, there are action schema learning approaches that learn sym- bolic rules (e.g. STRIPS or PDDL) to be used by traditional planners. However, these algorithms rely on logical descriptions of environment observations. In contrast, recent methods in deep reinforcement learning for games learn from pixel observations. However, they typically do not acquire an environment model, but a policy for one-step action selec- tion. Even when a model is learned, it cannot generalize to unseen instances of the training domain. Here we propose a neural network-based method that learns from visual obser- vations an approximate, compact, implicit representation of the domain dynamics, which can be used for planning with standard search algorithms, and generalizes to novel domain instances. The learned model is composed of submodules, each implicitly representing an action schema in the traditional sense. We evaluate our approach on visual versions of the standard domain Sokoban, and show that, by training on one single instance, it learns a transition model that can be successfully used to solve new levels of the game.