Preceding vehicle following algorithm with human driving characteristics

Preceding vehicle following algorithm with human driving characteristics
复制标题

符合人类驾驶特点的前车跟随算法

DOI:
10.1177/0954407020981546
复制
发表时间:
2021-01
期刊:
Proceedings of the Institution of Mechanical Engineers, Part D: Journal of Automobile Engineering
影响因子:
--
通讯作者:
Hong Bao
Hong Bao
中科院分区:
其他
文献类型:
--
作者:
feng pan;Hong Bao

文献摘要

参考文献

相似文献

本文提出了一种利用强化学习(RL)训练智能体执行具有人类驾驶特征的车辆跟随任务的新方法。我们参考逆强化学习的理想来设计强化学习模型的奖励函数。将车辆跟随过程中需要权衡的因素矢量化为奖励向量,并将奖励函数定义为奖励向量与权重的内积。收集人类驾驶员的驾驶数据并进行分析,得到真实的奖励函数。由于状态空间和动作空间是连续的,采用确定性策略梯度算法对RL模型进行训练。我们调整了奖励函数的权重向量,使强化学习模型的价值向量不断接近人类驾驶员的价值向量。经过几十轮的训练,我们选择了与人类驾驶员的值向量最接近的策略,并在PanoSim模拟环境中进行了测试。结果显示了智能体安全、平稳地跟随前车的理想性能。
This paper proposes a new approach of using reinforcement learning (RL) to train an agent to perform the task of vehicle following with human driving characteristics. We refer to the ideal of inverse reinforcement learning to design the reward function of the RL model. The factors that need to be weighed in vehicle following were vectorized into reward vectors, and the reward function was defined as the inner product of the reward vector and weights. Driving data of human drivers was collected and analyzed to obtain the true reward function. The RL model was trained with the deterministic policy gradient algorithm because the state and action spaces are continuous. We adjusted the weight vector of the reward function so that the value vector of the RL model could continuously approach that of a human driver. After dozens of rounds of training, we selected the policy with the nearest value vector to that of a human driver and tested it in the PanoSim simulation environment. The results showed the desired performance for the task of an agent following the preceding vehicle safely and smoothly.
DOI: --
发表时间: 1966
期刊: --
影响因子: --
作者:
H. Diamond;W. Lawrence
通讯作者: H. Diamond;W. Lawrence
DOI: --
发表时间: 2014-06
期刊: --
影响因子: --
作者:
David Silver;Guy Lever;N. Heess;T. Degris;Daan Wierstra;Martin A. Riedmiller
通讯作者: David Silver;Guy Lever;N. Heess;T. Degris;Daan Wierstra;Martin A. Riedmiller
DOI: 10.1007/978-981-13-8285-7_1
发表时间: 2020-12
期刊: Deep Reinforcement Learning in Unity
影响因子: --
作者:
Mohit Sewak
通讯作者: Mohit Sewak
DOI: --
发表时间: 2016-02
期刊: ArXiv
影响因子: --
作者:
Shai Shalev-Shwartz;Nir Ben-Zrihem;Aviad Cohen;A. Shashua
通讯作者: Shai Shalev-Shwartz;Nir Ben-Zrihem;Aviad Cohen;A. Shashua
DOI: --
发表时间: 1998-07
期刊: --
影响因子: --
作者:
J. Randløv;P. Alstrøm
通讯作者: J. Randløv;P. Alstrøm