Robotic Knee Tracking Control to Mimic the Intact Human Knee Profile Based on Actor-Critic Reinforcement Learning

Robotic Knee Tracking Control to Mimic the Intact Human Knee Profile Based on Actor-Critic Reinforcement Learning
复制标题

DOI:
10.1109/jas.2021.1004272
复制
发表时间:
2022-01-01
影响因子:
11.8
通讯作者:
Huang, He Helen
Huang, He Helen
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wu, Ruofan;Yao, Zhikai;Huang, He Helen

文献摘要

被引文献

相似文献

我们解决了一个国家的最先进的强化学习(RL)控制方法,自动配置机器人假体阻抗参数,使端到端,连续的运动,旨在为经股截肢受试者。具体来说,我们的演员评论家为基础的RL提供跟踪控制的机器人膝关节假体,以模仿完整的膝盖轮廓,这是一个显着的进步,从我们以前的RL为基础的自动调整假体控制参数,集中在调节控制与设计师规定的机器人膝关节轮廓作为目标。除了提出基于直接启发式动态规划(dHDP)的跟踪控制算法,我们提供了一个控制性能保证,包括约束输入的情况下。我们表明,我们提出的跟踪控制具有几个重要的属性,如学习网络的权重收敛,贝尔曼(次)最优的成本去值函数和控制输入,和实际的人-机器人系统的稳定性。我们进一步提供了一个系统的模拟建议的跟踪控制使用一个现实的人-机器人系统模拟器,OpenSim,模拟如何dHDP使平地行走,行走在不同的地形和不同的步伐。这些结果表明,我们提出的基于dHDP的跟踪控制不仅在理论上是合适的,而且实际上是有用的。
We address a state-of-the-art reinforcement learning (RL) control approach to automatically configure robotic prosthesis impedance parameters to enable end-to-end, continuous locomotion intended for transfemoral amputee subjects. Specifically, our actor-critic based RL provides tracking control of a robotic knee prosthesis to mimic the intact knee profile, This is a significant advance from our previous RL based automatic tuning of prosthesis control parameters which have centered on regulation control with a designer prescribed robotic knee profile as the target. In addition to presenting the tracking control algorithm based on direct heuristic dynamic programming (dHDP), we provide a control performance guarantee including the case of constrained inputs. We show that our proposed tracking control possesses several important properties, such as weight convergence of the learning networks, Bellman (sub) optimality of the cost-to-go value function and control input, and practical stability of the human-robot system. We further provide a systematic simulation of the proposed tracking control using a realistic human-robot system simulator, the OpenSim, to emulate how the dHDP enables level ground walking, walking on different terrains and at different paces. These results show that our proposed dHDP based tracking control is not only theoretically suitable, but also practically useful.