Fast Inverse Reinforcement Learning with Interval Consistent Graph for Driving Behavior Prediction

Fast Inverse Reinforcement Learning with Interval Consistent Graph for Driving Behavior Prediction
复制标题

用于驾驶行为预测的区间一致图的快速逆强化学习

DOI:
--
复制
发表时间:
2017
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
K. Hitomi
K. Hitomi
中科院分区:
--
文献类型:
--
作者:
M. Shimosaka;Junichi Sato;Kazuhito Takenaka;K. Hitomi

文献摘要

被引文献

相似文献

最大熵逆强化学习(MaxEnt IRL)是一种有效的学习人类行为的潜在回报的方法,但由于计算成本呈指数增长,它在高维状态空间中是难以处理的。近年来,一些工作在大的状态空间中的图近似MaxEnt IRL提供了成功的结果,但是,状态空间模型的类型是相当有限的。在这项工作中,我们将它们扩展到更通用的大型状态空间模型,图中的马尔可夫决策过程的时间间隔一致性得到保证。我们验证了我们提出的方法在驾驶行为预测的背景下。使用实际驾驶数据的实验结果证实了我们的算法在预测性能和计算成本方面优于其他现有的IRL框架。
Maximum entropy inverse reinforcement learning (MaxEnt IRL) is an effective approach for learning the underlying rewards of demonstrated human behavior, while it is intractable in high-dimensional state space due to the exponential growth of calculation cost. In recent years, a few works on approximating MaxEnt IRL in large state spaces by graphs provide successful results, however, types of state space models are quite limited. In this work, we extend them to more generic large state space models with graphs where time interval consistency of Markov decision processes are guaranteed. We validate our proposed method in the context of driving behavior prediction. Experimental results using actual driving data confirm the superiority of our algorithm in both prediction performance and computational cost over other existing IRL frameworks.