Reinforcement Learning based cooperative longitudinal control for reducing traffic oscillations and improving platoon stability

Reinforcement Learning based cooperative longitudinal control for reducing traffic oscillations and improving platoon stability
复制标题

DOI:
10.1016/j.trc.2022.103744
复制
发表时间:
2022-06-09
影响因子:
8.3
通讯作者:
Chen, Danjue
Chen, Danjue
中科院分区:
工程技术1区
文献类型:
--
作者:
Jiang, Liming;Xie, Yuanchang;Chen, Danjue

文献摘要

被引文献

相似文献

走走停停的交通对交通运营的效率和安全构成了重大挑战。在这项研究中,合作纵向控制的基础上软演员批评(SAC)强化学习(RL)提出了解决这个问题。奖励函数经过精心设计,以考虑车辆合作并实现三个主要目标:安全性,效率和振荡阻尼。提出了一个全球性的性能指标振荡阻尼RL和其他基线模型进行评估。根据可以共享机动信息的前车数量,提出了两个模型RL-1和RL-2,并使用HighD和仿真数据与人类驱动(HD)和自适应巡航控制(ACC)模型进行了比较。据发现,从额外的前面的车辆的信息,RL-2可以更有效地抑制冲击波。具体而言,RL-1和RL-2分别将流量振荡降低15%-36%和15%-42%,而HD将振荡放大14- 37%。ACC模型也可以抑制冲击波,但不如RL-1和RL-2有效。基于使用商业Model X车辆收集的数据,进一步评估了两种RL控制方法。与商用Model X ACC车辆在某些受控设置下相比,所提出的RL方法可以通过产生更小的振荡增长、超调和平均加/减速度变化来更好地抑制停-走波,这表明它们可以在新的但类似的环境中很好地推广。最后,RL方法进行评估,考虑一个车队的车辆具有不同的RL渗透率。结果表明,它们在抑制冲击波方面始终优于HD和ACC。
Stop-and-go traffic poses significant challenges to the efficiency and safety of traffic operations. In this study, a cooperative longitudinal control based on Soft Actor Critic (SAC) Reinforcement Learning (RL) is proposed to address this issue. The reward function is carefully designed to consider vehicle cooperation and to achieve three main objectives: safety, efficiency, and oscillation dampening. A global performance metric for oscillation dampening is proposed to evaluate the developed RL and other baseline models. Depending on the number of preceding vehicles that can share maneuver information, two models RL-1 and RL-2 are proposed and compared with human driven (HD) and an adaptive cruise control (ACC) model using the HighD and simulated data. It is found that with information from additional preceding vehicles, RL-2 can dampen shockwaves more efficiently. Specifically, RL-1 and RL-2 decrease traffic oscillation by 15%-36% and 15%-42%, respectively, while HD amplifies the oscillation by 14-37%. The ACC model can also dampen shockwaves but is not as effective as RL-1 and RL-2. The two RL control methods are further evaluated based on data collected using a commercial Model X vehicle. Compared with the commercial Model X ACC vehicle in some controlled settings, the proposed RL methods can better dampen the stop-and-go waves by generating smaller oscillation growth, overshooting, and average acceleration/deceleration rate change, suggesting that they can generalize well in a new but similar environment. Finally, the RL methods are evaluated considering a platoon of vehicles with different RL penetration rates. The results show that they consistently outperform HD and ACC in dampening shockwaves.