Adversarial Mixture Density Networks: Learning to Drive Safely from Collision Data

Adversarial Mixture Density Networks: Learning to Drive Safely from Collision Data
复制标题

DOI:
10.1109/itsc48978.2021.9564916
复制
发表时间:
2021-07
期刊:
2021 IEEE International Intelligent Transportation Systems Conference (ITSC)
影响因子:
--
通讯作者:
Sampo Kuutti;Saber Fallah;R. Bowden
Sampo Kuutti;Saber Fallah;R. Bowden
中科院分区:
其他
文献类型:
--
作者:
Sampo Kuutti;Saber Fallah;R. Bowden

文献摘要

相似文献

模仿学习已被广泛用于基于预先记录的数据学习自动驾驶的控制策略。然而,已经证明,当遇到训练分布之外的状态时,基于模仿学习的策略容易受到复合错误的影响。此外,这些代理已被证明很容易被敌对的道路使用者利用,目的是制造碰撞。为了克服这些缺点,我们引入了对抗性混合密度网络(AMDN),它从不同的数据集中学习两个分布。第一个是从自然主义的人类驾驶数据集学习到的安全动作的分布。第二个是表示可能导致碰撞的不安全动作的分布,该分布是从碰撞数据集学习的。在训练期间,我们利用这两个分布来提供基于这两个分布的相似性的额外损失。在碰撞数据集上进行训练时,根据安全动作分布与不安全动作分布的相似性对安全动作分布进行惩罚,从而得到更稳健、更安全的控制策略。我们在遵循用例的车辆上演示了所提出的AMDN方法,并在自然测试环境和对抗性测试环境下进行了评估。我们表明,尽管AMDN很简单,但与纯模仿学习或标准混合密度网络方法相比,它在学习控制策略的安全性方面提供了显著的好处。
Imitation learning has been widely used to learn control policies for autonomous driving based on pre-recorded data. However, imitation learning based policies have been shown to be susceptible to compounding errors when encountering states outside of the training distribution. Further, these agents have been demonstrated to be easily exploitable by adversarial road users aiming to create collisions. To overcome these shortcomings, we introduce Adversarial Mixture Density Networks (AMDN), which learns two distributions from separate datasets. The first is a distribution of safe actions learned from a dataset of naturalistic human driving. The second is a distribution representing unsafe actions likely to lead to collision, learned from a dataset of collisions. During training, we leverage these two distributions to provide an additional loss based on the similarity of the two distributions. By penalising the safe action distribution based on its similarity to the unsafe action distribution when training on the collision dataset, a more robust and safe control policy is obtained. We demonstrate the proposed AMDN approach in a vehicle following use-case, and evaluate under naturalistic and adversarial testing environments. We show that despite its simplicity, AMDN provides significant benefits for the safety of the learned control policy, when compared to pure imitation learning or standard mixture density network approaches.