Combining learned controllers to achieve new goals based on linearly solvable MDPs

Combining learned controllers to achieve new goals based on linearly solvable MDPs
复制标题

DOI:
10.1109/icra.2014.6907631
复制
发表时间:
2014-09
期刊:
2014 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
E. Uchibe;K. Doya
E. Uchibe;K. Doya
中科院分区:
其他
文献类型:
--
作者:
E. Uchibe;K. Doya

文献摘要

相似文献

学习复杂的行为通常需要密集的人工调整和昂贵的计算优化,因为我们必须求解一个非线性的Hamilton-Jacobi-Bellman(HJB)方程。最近,Todorov提出了一类所谓的线性可解马尔可夫决策过程(LMDP),它将一个非线性HJB方程转化为一个线性微分方程。简化的HJB方程的线性度使我们可以应用叠加从一组学习的原始控制器中推导出新的复合控制器。然而,他的方法是一种基于模型的方法,并且没有在真实的领域中进行评估。提出了一种类似于最小二乘时差(LSTD)学习的无模型方法。在该方法中,指数变换后的代价函数可以作为LSTD中的贴现因子。将提出的方法应用于四足机器人行走行为的学习,并在真实机器人实验中进行了评估。每个基元任务的目标是到达环境中特定的目标位置,而复合任务的目标是接近基元目标位置所代表的任意区域。实验结果表明,该组合策略可以作为新任务的良好初始策略。
Learning complicated behaviors usually involves intensive manual tuning and expensive computational optimization because we have to solve a nonlinear Hamilton-Jacobi-Bellman (HJB) equation. Recently, Todorov proposed a class of the so-called Linearly solvable Markov Decision Process (LMDP) which converts a nonlinear HJB equation to a linear differential equation. Linearity of the simplified HJB equation allows us to apply superposition to derive a new composite controller from a set of learned primitive controllers. However, his method was a model-based approach and it was not evaluated in a real domain. This study proposes a model-free method which is similar to the Least Squares Temporal Difference (LSTD) learning. In this method, the exponentially transformed cost function can be regarded as the discount factor in LSTD. Our proposed method is applied to learning walking behaviors with the quadruped robot to evaluate in real robot experiments. The goal of each primitive task is to go to the specific target position in the environment and that of the composite task is to approach arbitrary region represented by the primitives' target positions. Experimental results show that the composite policy can be used as a good initial policy for the new task.