Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward Tasks

Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward Tasks
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Robert McCarthy;S. Redmond
Robert McCarthy;S. Redmond
中科院分区:
其他
文献类型:
--
作者:
Robert McCarthy;S. Redmond

文献摘要

被引文献

相似文献

基于模型的强化学习对于实际机器人应用来说是一种很有前景的学习策略,因为它比无模型的强化学习提高了数据效率。然而,当前最先进的基于模型的方法依赖于成形的奖励信号,这可能难以设计和实现。为了解决这个问题,我们提出了一种针对稀疏奖励多目标任务的简单的基于模型的方法,无需复杂的奖励工程。这种方法被称为想象的后见之明体验重放,通过将想象的数据纳入策略更新来最大限度地减少现实世界的交互。为了改善稀疏奖励环境中的探索,该策略通过标准的事后经验重放进行训练,并赋予基于好奇心的内在奖励。经过评估,在基准 OpenAI Gym Fetch Robotics 任务中,与最先进的无模型方法相比,该方法平均数据效率提高了一个数量级。
Model-based reinforcement learning is a promising learning strategy for practical robotic applications due to its improved data-efficiency versus model-free counterparts. However, current state-of-the-art model-based methods rely on shaped reward signals, which can be difficult to design and implement. To remedy this, we propose a simple model-based method tailored for sparse-reward multi-goal tasks that foregoes the need for complicated reward engineering. This approach, termed Imaginary Hindsight Experience Replay, minimises real-world interactions by incorporating imaginary data into policy updates. To improve exploration in the sparse-reward setting, the policy is trained with standard Hindsight Experience Replay and endowed with curiosity-based intrinsic rewards. Upon evaluation, this approach provides an order of magnitude increase in data-efficiency on average versus the state-of-the-art model-free method in the benchmark OpenAI Gym Fetch Robotics tasks.