Path Integral Policy Improvement with Population Adaptation

Path Integral Policy Improvement with Population Adaptation
复制标题

走人口适应的综合政策改进之路

DOI:
10.1109/tcyb.2020.2983923
复制
发表时间:
2020
影响因子:
11.8
通讯作者:
and F. Matsuno
and F. Matsuno
中科院分区:
计算机科学1区
文献类型:
--
作者:
K. Yamamoto;R. Ariizumi;T. Hayakawa;and F. Matsuno

文献摘要

相似文献

路径积分策略改进(PI 2)是一种有效的强化学习算法,特别是当目标系统是高维动力系统时。然而,PI2及其现有的扩展具有可调节的参数,其效率显著依赖于这些参数。本文提出了一个扩展的PI 2,自动调整所有的关键参数。三种不同类型的模拟腿式机器人的运动采集任务进行了测试所提出的算法的有效性。结果表明,该方法不仅可以消除用户的负担,以设置适当的参数,但优化性能显着提高。对于其中一个运动,进行了真实的机器人实验,以证明运动的有效性。
Path integral policy improvement (PI2) is known to be an efficient reinforcement learning algorithm, particularly, if the target system is a high-dimensional dynamical system. However, PI2, and its existing extensions, have adjustable parameters, on which the efficiency depends significantly. This article proposes an extension of PI2that adjusts all of the critical parameters automatically. Motion acquisition tasks for three different types of simulated legged robots were performed to test the efficacy of the proposed algorithm. The results show that the proposed method cannot only eliminate the burden on the user to set the parameters appropriately but also improve the optimization performance significantly. For one of the acquired motions, a real robot experiment was conducted to show the validity of the motion.