A deep reinforcement learning based distributed control strategy for connected automated vehicles in mixed traffic platoon

A deep reinforcement learning based distributed control strategy for connected automated vehicles in mixed traffic platoon
复制标题

DOI:
10.1016/j.trc.2023.104019
复制
发表时间:
2023-03
期刊:
Transportation Research Part C: Emerging Technologies
影响因子:
--
通讯作者:
Haotian Shi;Danjue Chen;Nan Zheng;Xin Wang;Yang Zhou;Bin Ran
Haotian Shi;Danjue Chen;Nan Zheng;Xin Wang;Yang Zhou;Bin Ran
中科院分区:
其他
文献类型:
--
作者:
Haotian Shi;Danjue Chen;Nan Zheng;Xin Wang;Yang Zhou;Bin Ran

文献摘要

相似文献

提出了一种新的分布式纵向控制策略,用于连接自动车辆(CAV)在混合交通环境中的CAV和人类驾驶车辆(HDVs),将高维队列信息。对于混合交通,传统的CAV控制方法集中于微观轨迹信息,其在处理HDV随机性(例如,反应时间长;各种驾驶风格)和混合交通异质性。与传统方法不同,我们的方法首次将连续HDV作为一个整体进行表征(即,AHDV)来降低HDV的随机性,并利用其宏观特征来控制后续CAV。新的控制策略利用车队信息来预测混合交通场景下下游引起的干扰和交通特征,大大优于传统的方法。特别是,该控制算法基于深度强化学习(DRL)来实现车辆跟随控制效率,并通过将其嵌入训练环境中来进一步解决聚合车辆跟随行为的随机性。为了更好地利用宏观交通特征,混合交通的一般队列被归类为CAV-HDVs-CAV模式,并通过相应的DRL状态来描述。宏观交通流特性建立在纽韦尔跟驰模型上,以捕捉聚集HDV的联合行为特征。仿真实验验证了我们提出的策略。结果表明,所提出的控制方法具有出色的性能方面的振荡阻尼,生态驱动,和泛化能力。
This paper proposes an innovative distributed longitudinal control strategy for connected automated vehicles (CAVs) in the mixed traffic environment of CAV and human-driven vehicles (HDVs), incorporating high-dimensional platoon information. For mixed traffic, the traditional CAV control method focuses on microscopic trajectory information, which may not be efficient in handling the HDV stochasticity (e.g., long reaction time; various driving styles) and mixed traffic heterogeneities. Different from traditional methods, our method, for the first time, characterizes consecutive HDVs as a whole (i.e., AHDV) to reduce the HDV stochasticity and utilize its macroscopic features to control the following CAVs. The new control strategy takes advantage of platoon information to anticipate the disturbances and traffic features induced downstream under mixed traffic scenarios and greatly outperforms the traditional methods. In particular, the control algorithm is based on deep reinforcement learning (DRL) to fulfill car-following control efficiency and further address the stochasticity for the aggregated car following behavior by embedding it in the training environment. To better utilize the macroscopic traffic features, a general platoon of mixed traffic is categorized as a CAV-HDVs-CAV pattern and described by corresponding DRL states. The macroscopic traffic flow properties are built upon the Newell car-following model to capture the characteristics of aggregated HDVs' joint behaviors. Simulated experiments are conducted to validate our proposed strategy. The results demonstrate that the proposed control method has outstanding performances in terms of oscillation dampening, eco-driving, and generalization capability.