Modeling Human Driving Behavior Through Generative Adversarial Imitation Learning

Modeling Human Driving Behavior Through Generative Adversarial Imitation Learning
复制标题

DOI:
10.1109/tits.2022.3227738
复制
发表时间:
2022-12-16
影响因子:
8.5
通讯作者:
Kochenderfer, Mykel J.
Kochenderfer, Mykel J.
中科院分区:
工程技术1区
文献类型:
--
作者:
Bhattacharyya, Raunak;Wulfe, Blake;Kochenderfer, Mykel J.

文献摘要

被引文献

相似文献

自动驾驶汽车安全验证中的一个悬而未决的问题是在模拟中建立人类驾驶行为的可靠模型。这项工作提出了一种从现实世界驾驶演示数据中学习神经驾驶策略的方法。我们将人类驾驶建模为一个顺序决策问题,其特征是非线性和随机性以及未知的潜在成本函数。模仿学习是一种在成本函数未知或难以指定时生成智能行为的方法。基于逆强化学习 (IRL) 的工作,生成对抗性模仿学习 (GAIL) 旨在提供有效的模仿,即使是对于具有大型或连续状态和动作空间的问题,例如模拟人类驾驶。本文介绍了如何使用 GAIL 进行基于学习的驾驶员建模。由于驾驶员建模本质上是一个多智能体问题,需要对智能体之间的交互进行建模,因此本文描述了 GAIL 的参数共享扩展(称为 PS-GAIL)来解决多智能体驾驶员建模问题。此外,GAIL 与领域无关,因此很难在学习过程中编码与驾驶相关的特定知识。本文描述了奖励增强模仿学习(RAIL),它修改奖励信号以向代理提供特定领域的知识。最后,人类的示范取决于 GAIL 可能无法捕获的潜在因素。本文描述了 Burn-InfoGAIL,它可以解开演示中潜在的可变性。使用现实世界高速公路驾驶数据集 NGSIM 进行模仿学习实验。实验表明,对 GAIL 的这些修改可以成功地模拟高速公路驾驶行为,准确地复制人类演示,并在驾驶代理之间的交互中产生真实的、紧急的交通流行为。
An open problem in autonomous vehicle safety validation is building reliable models of human driving behavior in simulation. This work presents an approach to learn neural driving policies from real world driving demonstration data. We model human driving as a sequential decision making problem that is characterized by non-linearity and stochasticity, and unknown underlying cost functions. Imitation learning is an approach for generating intelligent behavior when the cost function is unknown or difficult to specify. Building upon work in inverse reinforcement learning (IRL), Generative Adversarial Imitation Learning (GAIL) aims to provide effective imitation even for problems with large or continuous state and action spaces, such as modeling human driving. This article describes the use of GAIL for learning-based driver modeling. Because driver modeling is inherently a multi-agent problem, where the interaction between agents needs to be modeled, this paper describes a parameter-sharing extension of GAIL called PS-GAIL to tackle multi-agent driver modeling. In addition, GAIL is domain agnostic, making it difficult to encode specific knowledge relevant to driving in the learning process. This paper describes Reward Augmented Imitation Learning (RAIL), which modifies the reward signal to provide domain-specific knowledge to the agent. Finally, human demonstrations are dependent upon latent factors that may not be captured by GAIL. This paper describes Burn-InfoGAIL, which allows for disentanglement of latent variability in demonstrations. Imitation learning experiments are performed using NGSIM, a real-world highway driving dataset. Experiments show that these modifications to GAIL can successfully model highway driving behavior, accurately replicating human demonstrations and generating realistic, emergent behavior in the traffic flow arising from the interaction between driving agents.