How to Avoid Being Eaten by a Grue: Structured Exploration Strategies for Textual Worlds

How to Avoid Being Eaten by a Grue: Structured Exploration Strategies for Textual Worlds
复制标题

如何避免被格鲁吃掉:文本世界的结构化探索策略

DOI:
--
复制
发表时间:
2020
期刊:
arXiv.org
影响因子:
--
通讯作者:
Mark O. Riedl
Mark O. Riedl
中科院分区:
--
文献类型:
--
作者:
Prithviraj Ammanabrolu;Ethan Tien;Matthew J. Hausknecht;Mark O. Riedl

文献摘要

被引文献

相似文献

基于文本的游戏是长谜题或任务,其特点是一系列稀疏且可能具有欺骗性的奖励。它们提供了一个理想的平台来开发使用组合大小的自然语言状态动作空间感知世界并对其采取行动的代理。标准的强化学习智能体缺乏有效探索这些空间的能力,并且常常难以克服瓶颈——智能体无法通过只是因为它们没有足够多次看到正确的动作序列来得到充分的强化。我们引入了 Q*BERT,一种通过回答问题来学习构建世界知识图谱的智能体,从而提高样本效率。为了克服瓶颈,我们进一步引入了 MC!Q*BERT 代理,它使用基于知识图谱的内在动机来检测瓶颈,并采用新颖的探索策略来有效地学习一系列策略模块来克服瓶颈。我们提出了一项消融研究和结果,展示了我们的方法如何在九种文本游戏中超越当前最先进的技术,其中包括流行的游戏 Zork,在该游戏中,学习代理首次突破了玩家被 Grue 吃掉的瓶颈。
Text-based games are long puzzles or quests, characterized by a sequence of sparse and potentially deceptive rewards. They provide an ideal platform to develop agents that perceive and act upon the world using a combinatorially sized natural language state-action space. Standard Reinforcement Learning agents are poorly equipped to effectively explore such spaces and often struggle to overcome bottlenecks---states that agents are unable to pass through simply because they do not see the right action sequence enough times to be sufficiently reinforced. We introduce Q*BERT, an agent that learns to build a knowledge graph of the world by answering questions, which leads to greater sample efficiency. To overcome bottlenecks, we further introduce MC!Q*BERT an agent that uses an knowledge-graph-based intrinsic motivation to detect bottlenecks and a novel exploration strategy to efficiently learn a chain of policy modules to overcome them. We present an ablation study and results demonstrating how our method outperforms the current state-of-the-art on nine text games, including the popular game, Zork, where, for the first time, a learning agent gets past the bottleneck where the player is eaten by a Grue.