Learning Natural Locomotion Behaviors for Humanoid Robots Using Human Bias

Learning Natural Locomotion Behaviors for Humanoid Robots Using Human Bias
复制标题

DOI:
10.1109/lra.2020.2972879
复制
发表时间:
2020-02
影响因子:
5.2
通讯作者:
Chuanyu Yang;Kai Yuan;Shuai Heng;T. Komura;Zhibin Li
Chuanyu Yang;Kai Yuan;Shuai Heng;T. Komura;Zhibin Li
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chuanyu Yang;Kai Yuan;Shuai Heng;T. Komura;Zhibin Li

文献摘要

相似文献

这封信提出了一种新的学习框架,它利用模仿学习、深度强化学习和控制理论的知识来实现人形运动,对于类人来说是自然的、动态的和稳健的。我们提出了引入人类偏见的新方法,即运动捕捉数据和特殊的多专家网络结构。我们使用了多专家网络结构来平滑融合行为特征,并使用了任务和模仿奖励的增广奖励设计。我们的奖励设计是可组合的,可调的,并可通过使用传统类人控制的基本概念来解释。我们严格验证了学习框架,并对其进行了基准测试,该框架在各种测试场景中始终产生健壮的运动行为。此外,我们还展示了在存在干扰(如地形不规则性和外部推力)的情况下学习稳健和通用政策的能力。
This letter presents a new learning framework that leverages the knowledge from imitation learning, deep reinforcement learning, and control theories to achieve human-style locomotion that is natural, dynamic, and robust for humanoids. We proposed novel approaches to introduce human bias, i.e. motion capture data and a special Multi-Expert network structure. We used the Multi-Expert network structure to smoothly blend behavioral features, and used the augmented reward design for the task and imitation rewards. Our reward design is composable, tunable, and explainable by using fundamental concepts from conventional humanoid control. We rigorously validated and benchmarked the learning framework which consistently produced robust locomotion behaviors in various test scenarios. Further, we demonstrated the capability of learning robust and versatile policies in the presence of disturbances, such as terrain irregularities and external pushes.