Finite State Machine-Based Motion-Free Learning of Biped Walking

Finite State Machine-Based Motion-Free Learning of Biped Walking
复制标题

DOI:
10.1109/access.2021.3055241
复制
发表时间:
2021-01-01
期刊:
影响因子:
3.9
通讯作者:
Lee, Yoonsang
Lee, Yoonsang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kang, Gyoo-Chul;Lee, Yoonsang

文献摘要

被引文献

相似文献

最近,深度强化学习(DRL)被广泛用于为物理模拟角色创建控制器。在基于DRL的方法中,使用运动捕捉片段作为跟踪参考的角色控制的模仿学习已经在控制各种自然运动的运动技能方面显示出成功的结果。但是,输出运动往往受到接近参考运动的约束,因此学习各种运动样式需要许多运动剪辑。本文提出了一种基于有限状态机(FSM)的学习策略的DRL方法,该方法以自由运动的方式(不使用任何运动数据)来控制模拟角色产生期望步态参数所指定的步态。控制策略基于每一步开始时的角色状态和用户指定的步态参数,例如期望的步长或最大摆动脚高度,学习输出每个FSM状态的目标姿势和状态之间的转换定时。基于有限状态机的策略学习与基本控制器中嵌入的简单线性平衡反馈相结合,对学习策略的性能有积极的协同作用。学习的策略允许模拟角色按照不断变化的步态参数的指示行走,同时响应外部扰动。我们通过交互控制、外部推动、比较和消融研究来证明我们的方法的有效性。
Recently, deep reinforcement learning (DRL) is commonly used to create controllers for physically simulated characters. Among DRL-based approaches, imitation learning for character control using motion capture clips as tracking references has shown successful results in controlling various motor skills with natural movement. However, the output motion tends to be constrained close to the reference motion, and thus the learning of various styles of motion requires many motion clips. In this paper, we present a DRL method for learning a finite state machine (FSM) based policy in a motion-free manner (without the use of any motion data), which controls a simulated character to produce a gait as specified by the desired gait parameters. The control policy learns to output the target pose for each FSM state and transition timing between states, based on the character state at the beginning of each step and the user-specified gait parameters, such as the desired step length or maximum swing foot height. The combination of FSM-based policy learning and simple linear balance feedback embedded in the base controller has a positive synergistic effect on the performance of the learned policy. The learned policy allows the simulated character to walk as instructed by the continuously changing the gait parameters while responding to external perturbations. We demonstrate the effectiveness of our approach through interactive control, external push, comparison, and ablation studies.