A Multi-Agent Reinforcement Learning Approach for Safe and Efficient Behavior Planning of Connected Autonomous Vehicles

A Multi-Agent Reinforcement Learning Approach for Safe and Efficient Behavior Planning of Connected Autonomous Vehicles
复制标题

DOI:
10.1109/tits.2023.3336670
复制
发表时间:
2020-03
影响因子:
8.5
通讯作者:
Songyang Han;Shangli Zhou;Jiangwei Wang;Lynn Pepin;Caiwen Ding;Jie Fu;Fei Miao
Songyang Han;Shangli Zhou;Jiangwei Wang;Lynn Pepin;Caiwen Ding;Jie Fu;Fei Miao
中科院分区:
工程技术1区
文献类型:
--
作者:
Songyang Han;Shangli Zhou;Jiangwei Wang;Lynn Pepin;Caiwen Ding;Jie Fu;Fei Miao

文献摘要

相似文献

无线技术的最新进步使联网的自动驾驶车辆(CAV)能够通过车对车(V2V)通信收集关于其环境的信息。在这项工作中,我们设计了一个基于信息共享的多智能体强化学习(MAIL)框架,以利用额外的信息进行决策,以提高交通效率和安全性。我们提出的安全参与者-批评者算法有两个新技术:截断的$\数学{q}$-函数和安全动作映射。截断的$\mathcal{q}$-函数利用来自相邻CAV的共享信息,使得$\mathcal{q}$-函数的联合状态和动作空间不会在大规模CAV系统的算法中增长。我们证明了截断的-$\数学{q}$和整体$q$-函数之间的逼近误差的界。安全动作映射为基于控制屏障功能的训练和执行提供了可证明的安全保障。使用CALA模拟器进行了实验,结果表明,在不同的CAV比和不同的交通密度下,我们的方法在平均速度和舒适性方面都提高了CAV系统的效率。我们还表明,我们的方法避免执行不安全的动作,并始终保持与其他车辆的安全距离。我们构建了一个弯道障碍物场景,表明共同的愿景可以帮助骑手更早地观察到障碍物,并采取行动避免交通拥堵。实验视频在https://songyanghan.github.io/cavmarl/.上
The recent advancements in wireless technology enable connected autonomous vehicles (CAVs) to gather information about their environment by vehicle-to-vehicle (V2V) communication. In this work, we design an information-sharing-based multi-agent reinforcement learning (MARL) framework for CAVs, to take advantage of the extra information when making decisions to improve traffic efficiency and safety. The safe actor-critic algorithm we propose has two new techniques: the truncated $\mathcal {Q}$ -function and safe action mapping. The truncated $\mathcal {Q}$ -function utilizes the shared information from neighboring CAVs such that the joint state and action spaces of the $\mathcal {Q}$ -function do not grow in our algorithm for a large-scale CAV system. We prove the bound of the approximation error between the truncated- $\mathcal {Q}$ and global $Q$ -functions. The safe action mapping provides a provable safety guarantee for both the training and execution based on control barrier functions. Using the CARLA simulator for experiments, we show that our approach improves the CAV system’s efficiency in terms of average velocity and comfort under different CAV ratios and different traffic densities. We also show that our approach avoids the execution of unsafe actions and always maintains a safe distance from other vehicles. We construct an obstacle-at-corner scenario to show that the shared vision can help CAVs to observe obstacles earlier and take action to avoid traffic jams. The experiment video is on https://songyanghan.github.io/cavmarl/.