AMI: Adaptive Motion Imitation Algorithm Based on Deep Reinforcement Learning

AMI: Adaptive Motion Imitation Algorithm Based on Deep Reinforcement Learning
复制标题

DOI:
10.1109/icra46639.2022.9812121
复制
发表时间:
2022-05
期刊:
2022 International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
N. Taghavi;Moath H. A. Alqatamin;D. Popa
N. Taghavi;Moath H. A. Alqatamin;D. Popa
中科院分区:
其他
文献类型:
--
作者:
N. Taghavi;Moath H. A. Alqatamin;D. Popa

文献摘要

相似文献

在本文中,我们开发了一种新的自适应运动模仿算法(AMI)的机器人系统。虽然AMI可用于各种人机交互场景,但我们对机器人康复特别感兴趣,其中机器人扮演演示和练习具有挑战性的运动理疗的角色。在治疗过程中,机器人首先向患者演示一个参考轨迹,该轨迹需要在练习过程中重复,然后根据患者的能力将其运动调整为循环速度和幅度。使用该算法,机器人系统学习人类用户的上身运动,并基于从用户学习的轨迹执行独特的、相似的和更容易的运动。AMI中的自适应基于深度强化学习,具有在机器人操作系统(ROS)环境中实现的深度确定性策略梯度。实验数据收集从11个用户在上半身的人-机器人模仿会话与社会机器人芝诺被用来表明,该算法可以学习参考肘关节轨迹的用户在离线的方式后,只有几个周期。最后,我们还实现了该算法在线使用巴克斯特机器人,以证明其学习和回放性能。
In this paper, we develop a novel adaptive motion imitation algorithm (AMI) for robotic systems. Although AMI can be used in a variety of human-robot interaction scenarios, we are particularly interested in robotic rehabilitation where the robot plays the role of demonstrating and practicing challenging motion physiotherapy. During therapy, the robot first demonstrates a reference trajectory to the patient that needs to be repeated during practice and then adapts its motion to a cyclic speed and amplitude based on the patient's abilities. Using this algorithm, the robotic system learns an upper-body motion of the human user and performs a unique, similar, and easier motion based on the learned trajectory from the user. Adaptation in the AMI is based on deep reinforcement learning with deep deterministic policy gradient implemented in the Robot Operating System (ROS) environment. Experimental data collected from 11 users during upper body human-robot imitation sessions with social robot Zeno was used to show that the algorithm can learn reference elbow joint trajectories of the user in an off-line manner after just a few cycles. Finally, we also implemented the algorithm online using the Baxter robot to demonstrate its learning and playback performance.