Safety Augmented Value Estimation From Demonstrations (SAVED): Safe Deep Model-Based RL for Sparse Cost Robotic Tasks

Safety Augmented Value Estimation From Demonstrations (SAVED): Safe Deep Model-Based RL for Sparse Cost Robotic Tasks
复制标题

DOI:
10.1109/lra.2020.2976272
复制
发表时间:
2019-05
影响因子:
5.2
通讯作者:
Brijen Thananjeyan;A. Balakrishna;Ugo Rosolia;Felix Li;R. McAllister;Joseph Gonzalez;S. Levine;F. Borrelli;Ken Goldberg
Brijen Thananjeyan;A. Balakrishna;Ugo Rosolia;Felix Li;R. McAllister;Joseph Gonzalez;S. Levine;F. Borrelli;Ken Goldberg
中科院分区:
计算机科学2区
文献类型:
--
作者:
Brijen Thananjeyan;A. Balakrishna;Ugo Rosolia;Felix Li;R. McAllister;Joseph Gonzalez;S. Levine;F. Borrelli;Ken Goldberg

文献摘要

相似文献

强化学习(RL)对于机器人来说是具有挑战性的,因为人工设计的难度很大,密集的代价函数会导致意外行为,而动态不确定性会使探索和约束满足具有挑战性。我们通过一种新的基于模型的强化学习算法来解决这些问题,该算法名为安全增强的示范价值估计(SAVED),它使用仅识别任务完成的监督和一组适度的次优示范来限制探索和在处理复杂约束的同时高效地学习。然后,我们将SAVE算法与3种最先进的基于模型和无模型的RL算法在6个标准仿真基准上进行了比较,这些基准涉及导航和操作以及达芬奇手术机器人上的物理打结任务。结果表明,SAVE在成功率、约束满足度和样本效率等方面都优于以往的方法,使得在不到一小时的时间内直接在真实机器人上安全地学习控制策略是可行的。对于机器人上的任务,基线成功的时间少于$\Text{5}\$,而保存的任务在前50个训练迭代中的成功率超过$\Text{75}\$。有关代码和补充材料,请访问https://tinyurl.com/saved-rl.
Reinforcement learning (RL) for robotics is challenging due to the difficulty in hand-engineering a dense cost function, which can lead to unintended behavior, and dynamical uncertainty, which makes exploration and constraint satisfaction challenging. We address these issues with a new model-based reinforcement learning algorithm, Safety Augmented Value Estimation from Demonstrations (SAVED), which uses supervision that only identifies task completion and a modest set of suboptimal demonstrations to constrain exploration and learn efficiently while handling complex constraints. We then compare SAVED with 3 state-of-the-art model-based and model-free RL algorithms on 6 standard simulation benchmarks involving navigation and manipulation and a physical knot-tying task on the da Vinci surgical robot. Results suggest that SAVED outperforms prior methods in terms of success rate, constraint satisfaction, and sample efficiency, making it feasible to safely learn a control policy directly on a real robot in less than an hour. For tasks on the robot, baselines succeed less than $\text{5}\%$ of the time while SAVED has a success rate of over $\text{75}\%$ in the first 50 training iterations. Code and supplementary material is available at https://tinyurl.com/saved-rl.