Situated Mapping of Sequential Instructions to Actions with Single-step Reward Observation

Situated Mapping of Sequential Instructions to Actions with Single-step Reward Observation
复制标题

DOI:
10.18653/v1/p18-1193
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Alane Suhr;Yoav Artzi
Alane Suhr;Yoav Artzi
中科院分区:
其他
文献类型:
--
作者:
Alane Suhr;Yoav Artzi

文献摘要

被引文献

相似文献

我们提出了一种学习方法,将与上下文相关的顺序指令映射到动作。我们使用基于注意力的模型来解决话语和状态依赖性问题,该模型既考虑互动的历史和世界状态。为了从开始和目标状态训练而无需参加示威活动,我们提出了Sestra,这是一种学习算法,利用单步奖励观察和立即预期的奖励最大化。我们对Scone域进行评估,并在使用高级逻辑表示的方法上显示了9.8%-25.3%的绝对精度提高。
We propose a learning approach for mapping context-dependent sequential instructions to actions. We address the problem of discourse and state dependencies with an attention-based model that considers both the history of the interaction and the state of the world. To train from start and goal states without access to demonstrations, we propose SESTRA, a learning algorithm that takes advantage of single-step reward observations and immediate expected reward maximization. We evaluate on the SCONE domains, and show absolute accuracy improvements of 9.8%-25.3% across the domains over approaches that use high-level logical representations.