Learning task-parametrized assistive strategies for exoskeleton robots by multi-task reinforcement learning

Learning task-parametrized assistive strategies for exoskeleton robots by multi-task reinforcement learning
复制标题

通过多任务强化学习学习外骨骼机器人任务参数化辅助策略

DOI:
--
复制
发表时间:
2017
期刊:
IEEE International Conference on Robotics and Automation
影响因子:
--
通讯作者:
J. Morimoto
J. Morimoto
中科院分区:
--
文献类型:
--
作者:
Masashi Hamaya;Takamitsu Matsubara;T. Noda;T. Teramae;J. Morimoto

文献摘要

被引文献

相似文献

最近的研究表明,强化学习具有很大的潜力,通过用户和机器人之间的物理交互来生成外骨骼中的辅助策略。以前的方法集中在特定于任务的辅助策略上,对于每一个任务(情况/上下文),用户需要与机器人交互以学习适当的辅助策略。因此,学习到的策略不能推广到新的任务。由于采样成本对于诸如外骨骼的人在回路系统是昂贵的,因此必须启用泛化。在本文中,我们提出了学习任务参数化辅助策略的外骨骼机器人。我们的方法采用了一种辅助策略,该策略取决于任务参数和状态变量,可以从不同任务的多组人机交互数据中学习,并且即使对于看不见的任务也可以推广,给定任务参数而无需额外学习。为了减轻用户在多任务学习过程中的负担,我们开发了一个数据高效的多任务强化学习框架。为了验证我们的方法的有效性,我们开发了一个实验平台与外骨骼机器人。我们进行了一系列的实验,其实验结果表明,我们的方法可以学习这样一个任务参数化的辅助策略,并推广到看不见的任务,以减少用户的肌电图信号(EMG)在任务。
Recent studies suggest that reinforcement learning has great potential for generating assistive strategies in exoskeletons through physical interactions between a user and a robot. Previous methods focused on a task-specific assistive strategy, where for every single task (situation/context), the user needs to interact with a robot to learn an appropriate assistive strategy. Therefore, the learned strategies cannot be generalized for a new task. Since the sampling cost is expensive for such human-in-the-loop systems as exoskeletons, generalization must be enabled. In this paper, we propose to learn task-parametrized assistive strategies for exoskeleton robots. Our method employs an assistive strategy, which depends on the task parameter and the state variable, that can be learned from multiple sets of human-robot interaction data across different tasks and generalized even for an unseen task, given the task parameter without additional learning. To alleviate the user's burden in the learning process across multiple tasks, we exploit a data-efficient multi-task reinforcement learning framework. To verify the effectiveness of our method, we developed an experimental platform with an exoskeleton robot. We conducted a series of experiments whose experimental results show that our method can learn such a task-parametrized assistive strategy and be generalized for unseen tasks to reduce the user's electromyography signals (EMGs) during tasks.