LeDeepChef: Deep Reinforcement Learning Agent for Families of Text-Based Games

LeDeepChef: Deep Reinforcement Learning Agent for Families of Text-Based Games
复制标题

LeDeepChef:用于基于文本的游戏系列的深度强化学习代理

DOI:
10.1609/aaai.v34i05.6228
复制
发表时间:
2019
期刊:
ArXiv
影响因子:
--
通讯作者:
T. Hofmann
T. Hofmann
中科院分区:
--
文献类型:
--
作者:
Leonard Adolphs;T. Hofmann

文献摘要

参考文献

被引文献

相似文献

虽然强化学习(RL)方法在最近的历史中在各个领域取得了重大成就,但自然语言任务基本上没有受到影响,因为它们的组成和组合性质使它们难以优化。随着基于文本的游戏(TBGs)领域的兴起,研究者们试图弥合这一鸿沟。受雅达利游戏中强化学习算法成功的启发,我们的想法是在受限的游戏世界中开发新方法,然后逐渐转向更复杂的环境。之前在tbg领域的工作主要集中在解决个人游戏。然而,我们考虑的任务是设计一个不只是在单个游戏中成功的代理,而是在共享相同主题的整个游戏系列中表现良好的代理。在这项工作中,我们展示了我们的深度强化学习代理——ledeepchef——它展示了在不同环境和任务描述下从未见过的同类游戏的泛化能力。该智能体参加了微软研究院的第一次TextWorld问题:语言和强化学习挑战,并在最终测试集中超越了所有竞争对手。挑战中的所有游戏都有相同的主题,即在现代房屋环境中烹饪,但在房间的布置、呈现的物品和特定目标(烹饪食谱)方面存在显著差异。为了构建一个在整个游戏系列中获得高分的代理,我们使用了一个演员-评论家框架,并通过使用来自分层强化学习的想法和在配方数据库上训练的专门模块来修剪动作空间。
While Reinforcement Learning (RL) approaches lead to significant achievements in a variety of areas in recent history, natural language tasks remained mostly unaffected, due to the compositional and combinatorial nature that makes them notoriously hard to optimize. With the emerging field of Text-Based Games (TBGs), researchers try to bridge this gap. Inspired by the success of RL algorithms on Atari games, the idea is to develop new methods in a restricted game world and then gradually move to more complex environments. Previous work in the area of TBGs has mainly focused on solving individual games. We, however, consider the task of designing an agent that not just succeeds in a single game, but performs well across a whole family of games, sharing the same theme. In this work, we present our deep RL agent—LeDeepChef—that shows generalization capabilities to never-before-seen games of the same family with different environments and task descriptions. The agent participated in Microsoft Research's First TextWorld Problems: A Language and Reinforcement Learning Challenge and outperformed all but one competitor on the final test set. The games from the challenge all share the same theme, namely cooking in a modern house environment, but differ significantly in the arrangement of the rooms, the presented objects, and the specific goal (recipe to cook). To build an agent that achieves high scores across a whole family of games, we use an actor-critic framework and prune the action-space by using ideas from hierarchical reinforcement learning and a specialized module trained on a recipe database.
DOI: --
发表时间: 2012
期刊: --
影响因子: --
作者:
Shay B. Cohen;Michael Collins
通讯作者: Shay B. Cohen;Michael Collins