Reinforcement Learning Control of Robotic Knee With Human-in-the-Loop by Flexible Policy Iteration

Reinforcement Learning Control of Robotic Knee With Human-in-the-Loop by Flexible Policy Iteration
复制标题

基于柔性策略迭代的人在环机器人膝关节强化学习控制

DOI:
10.1109/tnnls.2021.3071727
复制
发表时间:
2021-05-06
影响因子:
10.4
通讯作者:
Huang, He
Huang, He
中科院分区:
计算机科学1区
文献类型:
--
作者:
Gao, Xiang;Si, Jennie;Huang, He

文献摘要

被引文献

相似文献

我们受到人类-机器人系统面临的真正挑战的激励,开发在数据级别高效并具有性能保证的新设计,例如在系统级别上的稳定性和最佳性。现有的理论上考虑系统性能的近似/自适应动态规划(ADP)结果并不容易为该问题提供实用的学习控制算法,而解决数据效率问题的强化学习(RL)算法通常不能为受控系统提供性能保证。本研究通过在策略迭代算法中引入创新特征来填补这些重要的空白。我们引入了灵活策略迭代(FPI),它可以灵活地有机地将经验回放和来自先前经验的补充值整合到RL控制器中。我们证明了系统级的性能,包括近似值函数的收敛、解的(次)最优性和系统的稳定性。通过对人-机器人系统的真实模拟,验证了该方法的有效性。值得注意的是,我们在这项研究中面临的问题可能很难通过基于经典控制理论的设计方法来解决,因为在线或离线都几乎不可能获得定制的人-机器人系统的数学模型。我们所得到的结果也表明了RL控制在解决高维控制输入的现实和挑战性问题方面的巨大潜力。
We are motivated by the real challenges presented in a human-robot system to develop new designs that are efficient at data level and with performance guarantees, such as stability and optimality at system level. Existing approximate/adaptive dynamic programming (ADP) results that consider system performance theoretically are not readily providing practically useful learning control algorithms for this problem, and reinforcement learning (RL) algorithms that address the issue of data efficiency usually do not have performance guarantees for the controlled system. This study fills these important voids by introducing innovative features to the policy iteration algorithm. We introduce flexible policy iteration (FPI), which can flexibly and organically integrate experience replay and supplemental values from prior experience into the RL controller. We show system-level performances, including convergence of the approximate value function, (sub)optimality of the solution, and stability of the system. We demonstrate the effectiveness of the FPI via realistic simulations of the human-robot system. It is noted that the problem we face in this study may be difficult to address by design methods based on classical control theory as it is nearly impossible to obtain a customized mathematical model of a human-robot system either online or offline. The results we have obtained also indicate the great potential of RL control to solving realistic and challenging problems with high-dimensional control inputs.