Training an Actor-Critic Reinforcement Learning Controller for Arm Movement Using Human-Generated Rewards.

Training an Actor-Critic Reinforcement Learning Controller for Arm Movement Using Human-Generated Rewards.
复制标题

DOI:
10.1109/tnsre.2017.2700395
复制
发表时间:
2017-10
期刊:
IEEE transactions on neural systems and rehabilitation engineering : a publication of the IEEE Engineering in Medicine and Biology Society
影响因子:
--
通讯作者:
Kirsch RF
Kirsch RF
中科院分区:
其他
文献类型:
--
作者:
Jagodnik KM;Thomas PS;van den Bogert AJ;Branicky MS;Kirsch RF

文献摘要

被引文献

相似文献

功能性电刺激(FES)采用神经假体向因脊髓损伤(SCI)而瘫痪的个体的神经和肌肉施加电流,以恢复自主运动。神经假体控制器计算刺激模式以产生期望的动作。迄今为止,没有现有的控制器能够有效地使其控制策略适应于随时间变化的广泛的可能的生理手臂特性、伸展运动和用户偏好。强化学习(RL)是一种控制策略,可以将人类奖励信号作为输入,以允许人类用户塑造控制器行为。在这项研究中,10名神经系统完好的人类参与者分配主观的数字奖励来训练RL控制器,评估使用平面肌肉骨骼人类手臂模拟执行的目标导向的达到任务的动画。将使用人类训练器实现的RL控制器学习与使用由算法生成的类人奖励完成的学习进行比较;指标包括达到指定目标的成功率;达到目标所需的时间;以及目标超调。这两组控制器学习效率高,差异最小,显著优于标准控制器。奖励积极性和一致性被发现是无关的学习成功。这些结果表明,人的奖励可以有效地用来训练RL为基础的FES控制器。
Functional Electrical Stimulation (FES) employs neuroprostheses to apply electrical current to the nerves and muscles of individuals paralyzed by spinal cord injury (SCI) to restore voluntary movement. Neuroprosthesis controllers calculate stimulation patterns to produce desired actions. To date, no existing controller is able to efficiently adapt its control strategy to the wide range of possible physiological arm characteristics, reaching movements, and user preferences that vary over time. Reinforcement learning (RL) is a control strategy that can incorporate human reward signals as inputs to allow human users to shape controller behavior. In this study, ten neurologically intact human participants assigned subjective numerical rewards to train RL controllers, evaluating animations of goal-oriented reaching tasks performed using a planar musculoskeletal human arm simulation. The RL controller learning achieved using human trainers was compared with learning accomplished using human-like rewards generated by an algorithm; metrics included success at reaching the specified target; time required to reach the target; and target overshoot. Both sets of controllers learned efficiently and with minimal differences, significantly outperforming standard controllers. Reward positivity and consistency were found to be unrelated to learning success. These results suggest that human rewards can be used effectively to train RL-based FES controllers.