Terrain Adaptive Walking of Biped Neuromuscular Virtual Human Using Deep Reinforcement Learning

Terrain Adaptive Walking of Biped Neuromuscular Virtual Human Using Deep Reinforcement Learning
复制标题

DOI:
10.1109/access.2019.2927606
复制
发表时间:
2019
期刊:
影响因子:
3.9
通讯作者:
Jianpeng Wang;Wenhu Qin;Libo Sun
Jianpeng Wang;Wenhu Qin;Libo Sun
中科院分区:
计算机科学3区
文献类型:
--
作者:
Jianpeng Wang;Wenhu Qin;Libo Sun

文献摘要

被引文献

相似文献

已经有一些基于生物力学的控制系统实现了更逼真的虚拟人体运动。但与直接由比例微分作动器驱动的传统控制系统相比,其适应环境变化的能力较弱。在我们的方法中,我们构建了一个由低级脊柱反射层和高级策略控制层组成的分层神经肌肉虚拟人(NMVH)运动控制系统。脊柱反射层使用一个反馈网络将感觉信息映射到刺激,刺激肌肉产生关节扭矩。策略控制层包含深度神经网络,为脊柱反射层提供学习动作策略,实现地形自适应运动技能。采用粒子群优化算法对反馈网络的增益因子进行优化,找出虚拟人在平坦地形上自主行走的基本策略。采用近端策略优化算法对策略控制层的深度神经网络进行训练,学习如何调节动作以适应不断变化的地形。在Matlab中的仿真结果表明,虚拟人能够平稳行走,更好地适应给定的地形变化。实验结果表明,该控制系统提高了神经肌肉虚拟人的地形适应性行走能力。
There have been some biomechanics-based control systems that have achieved better realistic virtual human motion. Yet their abilities to adapt the changing environments are weaker than the traditional control systems with characters driven by proportional derivative actuators directly. In our method, we build a hierarchical neuromuscular virtual human (NMVH) motion control system that consists of a low-level spine reflex layer and a high-level policy control layer. The spine reflex layer uses a feedback net to map sensory information to excitations, which stimulate muscles to generate joint torques. The policy control layer includes a deep neural network, which provides a learned action policy to spine reflex layer for achieving terrain-adaptive motion skills. The particle swarm optimization algorithm is used to optimize the gain factors of the feedback net for finding out a basic policy to make the virtual human walk on the flat terrain autonomously. The proximal policy optimization algorithm is employed to train the deep neural network in policy control layer for learning how to modulate the actions to adapt to the changing terrain. The simulation results in Matlab show that virtual human can walk smoothly and better adapt to the given terrain changes. It demonstrates that our control system improves the terrain-adaptive walking skill of the neuromuscular virtual human.