Online Reinforcement Learning Control for the Personalization of a Robotic Knee Prosthesis

Online Reinforcement Learning Control for the Personalization of a Robotic Knee Prosthesis
复制标题

DOI:
10.1109/tcyb.2019.2890974
复制
发表时间:
2020-06-01
影响因子:
11.8
通讯作者:
Huang, He (Helen)
Huang, He (Helen)
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wen, Yue;Si, Jennie;Huang, He (Helen)

文献摘要

被引文献

相似文献

机器人假肢比被动假体提供了更大的功能,但是我们面临着调整大量控制参数的挑战,以便为单个截肢者用户个性化设备。传统的控制设计或最新的机器人技术不容易解决此问题。强化学习(RL)自然具有吸引力。 Alphazero最近的前所未有的成功表明RL是可行的大规模问题解决者。但是,假体调整问题与几个未解决的问题有关,例如它没有已知且稳定的模型,问题的连续状态和控制可能会导致维度的诅咒,并且人类验证系统不断受到测量噪声,环境变化,被人体造成的人体变化。在本文中,我们证明了直接启发式动态编程(一种近似动态编程(ADP)方法)的可行性,以自动调整12个机器人膝盖假体参数以满足人类用户的需求。我们在两个受试者(一个健全的受试者和一个截肢者)上测试了ADP-tuner,以固定的速度在跑步机上行走。 ADP-Tuner学会了平均300步态循环或步行10分钟的步态运动学。当我们将先前学习的ADP控制器转移到具有同一主题的新学习会议上时,我们观察到了改善的ADP调音性能。据我们所知,我们个性化机器人假体的方法是将在线ADP学习控制控制到涉及人类受试者的临床问题的首次实施。
Robotic prostheses deliver greater function than passive prostheses, but we face the challenge of tuning a large number of control parameters in order to personalize the device for individual amputee users. This problem is not easily solved by traditional control designs or the latest robotic technology. Reinforcement learning (RL) is naturally appealing. The recent, unprecedented success of AlphaZero demonstrated RL as a feasible, large-scale problem solver. However, the prosthesis-tuning problem is associated with several unaddressed issues such as that it does not have a known and stable model, the continuous states and controls of the problem may result in a curse of dimensionality, and the human-prosthesis system is constantly subject to measurement noise, environmental change and human-body-caused variations. In this paper, we demonstrated the feasibility of direct heuristic dynamic programming, an approximate dynamic programming (ADP) approach, to automatically tune the 12 robotic knee prosthesis parameters to meet individual human users' needs. We tested the ADP-tuner on two subjects (one able-bodied subject and one amputee subject) walking at a fixed speed on a treadmill. The ADP-tuner learned to reach target gait kinematics in an average of 300 gait cycles or 10 min of walking. We observed improved ADP tuning performance when we transferred a previously learned ADP controller to a new learning session with the same subject. To the best of our knowledge, our approach to personalize robotic prostheses is the first implementation of online ADP learning control to a clinical problem involving human subjects.