PixL2R: Guiding Reinforcement Learning Using Natural Language by Mapping Pixels to Rewards

PixL2R: Guiding Reinforcement Learning Using Natural Language by Mapping Pixels to Rewards
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Prasoon Goyal;S. Niekum;R. Mooney
Prasoon Goyal;S. Niekum;R. Mooney
中科院分区:
其他
文献类型:
--
作者:
Prasoon Goyal;S. Niekum;R. Mooney

文献摘要

相似文献

强化学习(RL),特别是在稀疏奖励设置中,通常需要与环境进行大量的交互,从而限制了其对复杂问题的适用性。为了解决这个问题,一些先前的方法使用自然语言来指导代理的探索。然而,这些方法通常对环境的结构化表示进行操作,和/或假设自然语言命令中的某种结构。在这项工作中,我们提出了一个模型,直接映射像素奖励,给定一个自由形式的自然语言描述的任务,然后可以用于策略学习。我们在Meta-World机器人操作领域的实验表明,基于语言的奖励显着提高了策略学习的样本效率,无论是在稀疏和密集的奖励设置。
Reinforcement learning (RL), particularly in sparse reward settings, often requires prohibitively large numbers of interactions with the environment, thereby limiting its applicability to complex problems. To address this, several prior approaches have used natural language to guide the agent's exploration. However, these approaches typically operate on structured representations of the environment, and/or assume some structure in the natural language commands. In this work, we propose a model that directly maps pixels to rewards, given a free-form natural language description of the task, which can then be used for policy learning. Our experiments on the Meta-World robot manipulation domain show that language-based rewards significantly improves the sample efficiency of policy learning, both in sparse and dense reward settings.