Recurrent Deterministic Policy Gradient Method for Bipedal Locomotion on Rough Terrain Challenge

Recurrent Deterministic Policy Gradient Method for Bipedal Locomotion on Rough Terrain Challenge
复制标题

DOI:
10.1109/icarcv.2018.8581309
复制
发表时间:
2017-10
期刊:
2018 15th International Conference on Control, Automation, Robotics and Vision (ICARCV)
影响因子:
--
通讯作者:
Doo Re Song;Chuanyu Yang;C. McGreavy;Zhibin Li
Doo Re Song;Chuanyu Yang;C. McGreavy;Zhibin Li
中科院分区:
其他
文献类型:
--
作者:
Doo Re Song;Chuanyu Yang;C. McGreavy;Zhibin Li

文献摘要

相似文献

本文提出了一个深度学习框架,该框架能够基于我们对递归确定性策略梯度(RDPG)的新解释来解决部分可观察的运动任务。分别研究了环境部分可观性和子轨迹采样引起的采样误差测度偏差及其方差。在我们的RDPG为基础的学习框架中引入了三个主要的改进:尾步骤引导的时间差,初始化的隐藏状态使用过去的子轨迹,截断的时间反向传播,并注入外部经验的其他代理。所提出的学习框架被实现来解决OpenAI的健身房模拟环境中的双足步行者挑战,其中只有部分状态信息可用。我们的模拟研究表明,RDPG智能体产生的自主行为对各种障碍物具有高度适应性,使智能体能够有效地长距离穿越崎岖地形,成功率高于领先竞争者。
This paper presents a deep learning framework that is capable of solving partially observable locomotion tasks based on our novel interpretation of Recurrent Deterministic Policy Gradient (RDPG). We study on bias of sampled error measure and its variance induced by the partial observability of environment and subtrajectory sampling, respectively. Three major improvements are introduced in our RDPG based learning framework: tail-step bootstrap of temporal difference, initialisation of hidden state using past subtrajectory, truncation of temporal backpropagation, and injection of external experiences learned by other agents. The proposed learning framework was implemented to solve the Bipedal-Walker challenge in OpenAI's gym simulation environment where only partial state information is available. Our simulation study shows that the autonomous behaviors generated by the RDPG agent are highly adaptive to a variety of obstacles and enables the agent to effectively traverse rugged terrains for long distance with higher success rate than leading contenders.