Spatial-Temporal-Aware Safe Multi-Agent Reinforcement Learning of Connected Autonomous Vehicles in Challenging Scenarios

Spatial-Temporal-Aware Safe Multi-Agent Reinforcement Learning of Connected Autonomous Vehicles in Challenging Scenarios
复制标题

DOI:
10.1109/icra48891.2023.10161216
复制
发表时间:
2022-10
期刊:
2023 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Zhili Zhang;Songyang Han;Jiangwei Wang;Fei Miao
Zhili Zhang;Songyang Han;Jiangwei Wang;Fei Miao
中科院分区:
其他
文献类型:
--
作者:
Zhili Zhang;Songyang Han;Jiangwei Wang;Fei Miao

文献摘要

被引文献

相似文献

通信技术使互联车辆和自动驾驶车辆(CAV)之间能够协调。然而,如何利用共享信息来提高CAV系统在动态和复杂驾驶场景中的安全性和效率仍然不清楚。在这项工作中,我们提出了一个框架的约束多智能体强化学习(MARL)与并行安全盾的CAV在具有挑战性的驾驶场景,包括未连接的危险车辆。MARL的协调机制包括信息共享和协作策略学习,采用图卷积网络(GCN)-Transformer作为时空编码器,增强了Agent的环境感知能力。安全盾模块具有基于控制屏障功能(CBF)的安全检查功能,可防止座席采取不安全的操作。我们设计了一个受约束的多智能体优势行动者-批评者(CMAA 2C)算法来训练CAV的安全和合作策略。部署在CARLA模拟器的实验,我们验证了安全检查,时空编码器,并在我们的方法设计的协调机制的性能,通过比较实验,在几个具有挑战性的情况下,与未连接的危险车辆。结果表明,我们提出的方法显着提高系统的安全性和效率在具有挑战性的情况下。
Communication technologies enable coordination among connected and autonomous vehicles (CAVs). However, it remains unclear how to utilize shared information to improve the safety and efficiency of the CAV system in dynamic and complicated driving scenarios. In this work, we propose a framework of constrained multi-agent reinforcement learning (MARL) with a parallel Safety Shield for CAVs in challenging driving scenarios that includes unconnected hazard vehicles. The coordination mechanisms of the proposed MARL include information sharing and cooperative policy learning, with Graph Convolutional Network (GCN)-Transformer as a spatial-temporal encoder that enhances the agent's environment awareness. The Safety Shield module with Control Barrier Functions (CBF)-based safety checking protects the agents from taking unsafe actions. We design a constrained multi-agent advantage actor-critic (CMAA2C) algorithm to train safe and cooperative policies for CAVs. With the experiment deployed in the CARLA simulator, we verify the performance of the safety checking, spatial-temporal encoder, and coordination mechanisms designed in our method by comparative experiments in several challenging scenarios with unconnected hazard vehicles. Results show that our proposed methodology significantly increases system safety and efficiency in challenging scenarios.