Energy-Based Hindsight Experience Prioritization

Energy-Based Hindsight Experience Prioritization
复制标题

基于能源的事后经验优先级排序

DOI:
--
复制
发表时间:
2018
期刊:
Conference on Robot Learning
影响因子:
--
通讯作者:
Volker Tresp
Volker Tresp
中科院分区:
--
文献类型:
--
作者:
Rui Zhao;Volker Tresp

文献摘要

被引文献

相似文献

在后见之明经验重播(HER)中,强化学习代理通过将其所取得的任何成就视为虚拟目标来进行训练。然而,在以前的工作中,经验被随机重播,没有考虑哪一集可能是最有价值的学习。在本文中,我们开发了一个基于能量的框架,用于在机器人操作任务中优先考虑后见之明的经验。我们的方法是受到物理学中的功-能原理的启发。我们定义弹道能量函数为目标物体在弹道上的跃迁能量之和。我们假设,在机器人学中,重放具有高轨迹能量的情节对于强化学习更有效。为了验证我们的假设,我们设计了一个基于目标状态轨迹能量的后见之明体验优先排序框架。轨道能量函数考虑了势能、动能和转动能。我们评估了我们的基于能量的优先级(EBP)方法在四个具有挑战性的机器人操作任务上的仿真。我们的实验结果表明,我们提出的方法在不增加计算时间的情况下,在所有四个任务上的性能和样本效率方面都超过了最先进的方法。显示实验结果的视频可在此HTTPS URL上获得
In Hindsight Experience Replay (HER), a reinforcement learning agent is trained by treating whatever it has achieved as virtual goals. However, in previous work, the experience was replayed at random, without considering which episode might be the most valuable for learning. In this paper, we develop an energy-based framework for prioritizing hindsight experience in robotic manipulation tasks. Our approach is inspired by the work-energy principle in physics. We define a trajectory energy function as the sum of the transition energy of the target object over the trajectory. We hypothesize that replaying episodes that have high trajectory energy is more effective for reinforcement learning in robotics. To verify our hypothesis, we designed a framework for hindsight experience prioritization based on the trajectory energy of goal states. The trajectory energy function takes the potential, kinetic, and rotational energy into consideration. We evaluate our Energy-Based Prioritization (EBP) approach on four challenging robotic manipulation tasks in simulation. Our empirical results show that our proposed method surpasses state-of-the-art approaches in terms of both performance and sample-efficiency on all four tasks, without increasing computational time. A video showing experimental results is available at this https URL