Inverse reinforcement learning using Dynamic Policy Programming

Inverse reinforcement learning using Dynamic Policy Programming
复制标题

DOI:
10.1109/devlrn.2014.6982985
复制
发表时间:
2014-12
期刊:
4th International Conference on Development and Learning and on Epigenetic Robotics
影响因子:
--
通讯作者:
E. Uchibe;K. Doya
E. Uchibe;K. Doya
中科院分区:
其他
文献类型:
--
作者:
E. Uchibe;K. Doya

文献摘要

相似文献

在动态策略规划框架下,提出了一种新的基于密度比估计的无模型逆强化学习方法。我们证明了最优策略与基准策略之比的对数由状态依赖成本和价值函数表示。我们的建议是用密度比估计方法估计保单的密度比,用正则化的最小二乘法估计状态依赖成本和满足该关系的价值函数。该方法避免了求解配分函数等积分运算。一个简单的网格世界导航、汽车驾驶和摆动的数值模拟显示了它比传统方法的优越性。
This paper proposes a novel model-free inverse reinforcement learning method based on density ratio estimation under the framework of Dynamic Policy Programming. We show that the logarithm of the ratio between the optimal policy and the baseline policy is represented by the state-dependent cost and the value function. Our proposal is to use density ratio estimation methods to estimate the density ratio of policies and the least squares method with regularization to estimate the state-dependent cost and the value function that satisfies the relation. Our method can avoid computing the integral such as evaluating the partition function. A simple numerical simulation of a grid world navigation, a car driving, and a pendulum swing-up shows its superiority over conventional methods.