Learning from Demonstration for Real-Time User Goal Prediction and Shared Assistive Control

Learning from Demonstration for Real-Time User Goal Prediction and Shared Assistive Control
复制标题

从实时用户目标预测和共享辅助控制的演示中学习

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Conference on Robotics and Automation
影响因子:
--
通讯作者:
H. Admoni
H. Admoni
中科院分区:
--
文献类型:
--
作者:
Calvin Z. Qiao;Maram Sakr;Katharina Muelling;H. Admoni

文献摘要

参考文献

被引文献

相似文献

在共享自主中,用户输入与辅助运动混合以完成机器人通常不知道用户目标的任务。人与机器人之间的透明度对于有效协作至关重要。先前的工作已经为机器人提供了推断用户目标的方法;然而,它们通常依赖于机器人与物体之间的距离,这可能与用户的实时控制意图没有直接关联,从而导致控制感较低。在这里,我们提出了一种实时目标预测方法,该方法由通过演示学习(LfD)生成的辅助运动驱动,允许更多反应性辅助行为。 LfD 生成的辅助运动与基于目标预测的用户输入混合,以实现目标任务。 LfD 策略是离线学习的,并与不同的用户一起使用。为了评估我们提出的方法,我们将其与使用距离成本和直接控制方法(即操纵杆)的最先进的基于部分可观察马尔可夫决策过程(POMDP)的方法进行了比较。进行了一项试点研究 (N = 6),以控制 6 自由度 Kinova Mico 机械臂执行三项任务:(1) 伸手抓取、(2) 倾倒和 (3) 使用三种控制方法返回物体。我们在比较研究中使用了客观和主观测量。结果表明,与基于 POMDP 的方法相比,我们的方法具有最短的任务完成时间、所有三种控制方法中最少的操纵杆控制输入量,以及用户输入和辅助运动之间的角度差异显着更低。此外,它在用户偏好和感知速度评级方面获得了最高的主观得分,在控制感和机器人做了我想要的评级方面获得了第二高的分数。
In shared autonomy, the user input is blended with the assistive motion to accomplish a task where the user goal is typically unknown to the robot. Transparency between the human and robot is essential for effective collaboration. Prior works have provided methods for the robot to infer the user goal; however, they are usually dependent on the distance between the robot and object, which may not be directly associated with the real-time user control intention and thus cause low control feelings. Here, we propose a real-time goal prediction method driven by assistive motion generated by learning from demonstration (LfD) allowing more reactive assistive behaviors. This LfD-generated assistive motion is blended with the user input based on goal predictions to achieve targeted tasks. The LfD policy was learned offline and used with different users. To evaluate our proposed method, we compared it with a state-of-the-art Partially Observable Markov Decision Process (POMDP) based method using a distance cost, and a direct control method (i.e., joystick). A pilot study (N = 6) was conducted to control a 6-DoF Kinova Mico robotic arm to carry out three tasks: (1) reaching-and-grasping, (2) pouring, and (3) object-returning with the three control methods. We used both objective and subjective measures in the comparative study. Results show that our method has the shortest task completion time, the lowest amount of joystick control inputs among all three control methods, as well as a significantly lower angular difference between the user input and assistive motion compared to the POMDP-based method. Besides, it obtains the highest subjective score in the user preference and perceived speed ratings, and the second-highest in the control feeling and the robot did what I wanted ratings.
调节人类输入以实现动态环境中的共享自治
DOI: 10.1109/ro-man46459.2019.8956304
发表时间: 2019
期刊: --
影响因子: --
作者:
Mower C
通讯作者: Mower C