Learning From Sparse Demonstrations

Learning From Sparse Demonstrations
复制标题

DOI:
10.1109/tro.2022.3191592
复制
发表时间:
2020-08
影响因子:
7.8
通讯作者:
Wanxin Jin;T. Murphey;D. Kulić;Neta Ezer;Shaoshuai Mou
Wanxin Jin;T. Murphey;D. Kulić;Neta Ezer;Shaoshuai Mou
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wanxin Jin;T. Murphey;D. Kulić;Neta Ezer;Shaoshuai Mou

文献摘要

被引文献

相似文献

在这篇文章中,我们开发了连续庞特里亚金可微编程(连续PDP)的方法,它使机器人能够从几个稀疏演示的关键帧中学习目标函数。标记有一些时间戳的关键帧是期望的任务空间输出,机器人期望顺序地遵循这些输出。关键帧的时间戳可以与机器人的实际执行时间不同。该方法联合找到一个目标函数和时间扭曲函数,使得机器人的所得轨迹顺序地遵循关键帧,具有最小的差异损失。连续PDP通过有效地求解机器人轨迹相对于未知参数的梯度,使用投影梯度下降来最小化差异损失。该方法首先在模拟机器人手臂上进行评估,然后应用于6自由度四旋翼机,以学习未建模环境中运动规划的目标函数。结果表明,该方法的效率,它能够处理关键帧和机器人执行之间的时间错位,以及客观学习到看不见的运动条件的推广。
In this article, we develop the method of continuous Pontryagin differentiable programming (Continuous PDP), which enables a robot to learn an objective function from a few sparsely demonstrated keyframes. The keyframes, labeled with some time stamps, are the desired task-space outputs, which a robot is expected to follow sequentially. The time stamps of the keyframes can be different from the time of the robot's actual execution. The method jointly finds an objective function and a time-warping function such that the robot's resulting trajectory sequentially follows the keyframes with minimal discrepancy loss. The Continuous PDP minimizes the discrepancy loss using projected gradient descent by efficiently solving the gradient of the robot trajectory with respect to the unknown parameters. The method is first evaluated on a simulated robot arm and then applied to a 6-DoF quadrotor to learn an objective function for motion planning in unmodeled environments. The results show the efficiency of the method, its ability to handle time misalignment between keyframes and robot execution, and the generalization of objective learning into unseen motion conditions.