Emergence of human-comparable balancing behaviours by deep reinforcement learning

Emergence of human-comparable balancing behaviours by deep reinforcement learning
复制标题

DOI:
10.1109/humanoids.2017.8246900
复制
发表时间:
2017-11
期刊:
2017 IEEE-RAS 17th International Conference on Humanoid Robotics (Humanoids)
影响因子:
--
通讯作者:
Chuanyu Yang;Taku Komura;Zhibin Li
Chuanyu Yang;Taku Komura;Zhibin Li
中科院分区:
其他
文献类型:
--
作者:
Chuanyu Yang;Taku Komura;Zhibin Li

文献摘要

相似文献

本文提出了一种基于深度强化学习的分层框架,该框架自然地获取能够执行平衡行为(例如人形机器人的脚踝推出)的控制策略,而无需对控制器进行明确的人类设计。只有训练神经网络的奖励是根据物理原理和数量专门制定的,因此是可解释的。通过深度强化学习成功出现的与人类相似的行为证明了在统一框架中使用基于人工智能的方法进行人形运动控制的可行性。此外,通过强化学习学习到的平衡策略比基于零矩点的方法提供了更大范围的干扰抑制,这表明了使用基于学习的控制来探索最优性能的研究方向。
This paper presents a hierarchical framework based on deep reinforcement learning that naturally acquires control policies that are capable of performing balancing behaviours such as ankle push-offs for humanoid robots, without explicit human design of controllers. Only the reward for training the neural network is specifically formulated based on the physical principles and quantities, and hence explainable. The successful emergence of human-comparable behaviours through the deep reinforcement learning demonstrates the feasibility of using an AI-based approach for humanoid motion control in a unified framework. Moreover, the balance strategies learned by reinforcement learning provides a larger range of disturbance rejection than that of the zero moment point based methods, suggesting a research direction of using learning-based controls to explore the optimal performance.