Differential dynamic programming with temporally decomposed dynamics

Differential dynamic programming with temporally decomposed dynamics
复制标题

DOI:
10.1109/humanoids.2015.7363430
复制
发表时间:
2015-12
期刊:
2015 IEEE-RAS 15th International Conference on Humanoid Robots (Humanoids)
影响因子:
--
通讯作者:
Akihiko Yamaguchi;C. Atkeson
Akihiko Yamaguchi;C. Atkeson
中科院分区:
其他
文献类型:
--
作者:
Akihiko Yamaguchi;C. Atkeson

文献摘要

相似文献

我们探讨了动态的时间分解,以提高政策学习与未知的动态。对于动态未知的策略学习,有无模型方法和基于模型的方法,但这两种方法都存在问题:一般来说,无模型方法的泛化能力较差,而基于模型的方法往往受到假设模型结构的限制,或者需要收集许多样本来建立模型。我们考虑动态的时间分解,使学习模型更容易。为了获得一个政策,我们应用微分动态规划(DDP)。我们的方法的一个特点是,我们认为分解的动态,即使没有采取行动,这使我们能够更灵活地分解动态。因此,学习的动力学变得更加准确。我们的DDP是一个一阶梯度下降算法与随机评价函数。在具有学习模型的DDP中,通常存在许多局部最大值。为了避免它们,我们考虑多准则评价函数。除了随机评价函数外,我们还使用了参考值函数。通过浇注模拟实验验证了该方法的有效性,并在实验中建立了复杂的动力学模型。结果表明,我们可以优化DDP的动作,同时学习动力学模型。
We explore a temporal decomposition of dynamics in order to enhance policy learning with unknown dynamics. There are model-free methods and model-based methods for policy learning with unknown dynamics, but both approaches have problems: in general, model-free methods have less generalization ability, while model-based methods are often limited by the assumed model structure or need to gather many samples to make models. We consider a temporal decomposition of dynamics to make learning models easier. To obtain a policy, we apply differential dynamic programming (DDP). A feature of our method is that we consider decomposed dynamics even when there is no action to be taken, which allows us to decompose dynamics more flexibly. Consequently learned dynamics become more accurate. Our DDP is a first-order gradient descent algorithm with a stochastic evaluation function. In DDP with learned models, typically there are many local maxima. In order to avoid them, we consider multiple criteria evaluation functions. In addition to the stochastic evaluation function, we use a reference value function. This method was verified with pouring simulation experiments where we created complicated dynamics. The results show that we can optimize actions with DDP while learning dynamics models.