Scalable multi-region perimeter metering control for urban networks: A multi-agent deep reinforcement learning approach

Scalable multi-region perimeter metering control for urban networks: A multi-agent deep reinforcement learning approach
复制标题

DOI:
10.1016/j.trc.2023.104033
复制
发表时间:
2023-03
期刊:
Transportation Research Part C: Emerging Technologies
影响因子:
--
通讯作者:
Dongqin Zhou;V. Gayah
Dongqin Zhou;V. Gayah
中科院分区:
其他
文献类型:
--
作者:
Dongqin Zhou;V. Gayah

文献摘要

被引文献

相似文献

基于宏观基本图的周界计量控制在过去的十年中引起了越来越多的研究兴趣。该策略提供了一种方便的方法来缓解城市拥堵,通过操纵均匀区域的车辆运动,而无需建模与单个车辆存在相关的详细行为和相互作用。特别是,多区域周界计量控制为大规模城市网络中的高效交通管理提供了希望。然而,用于多区域控制的大多数现有方法需要环境交通动态或网络属性(即,临界累积),而这样的信息通常难以获得,并且会产生显著的估计误差。另一方面,最近开发的无模型技术尚未显示出可扩展性或适用于大型城市网络。为了填补这一空白,本文提出了一种基于多智能体深度强化学习的可扩展无模型方案。该方案的特点是在集中训练与分散执行的范式中进行价值函数分解,再加上单智能体深度强化学习和领域专业知识指导下的问题重构的关键进展。在七个城市网络上的综合实验结果表明,该方案是有效的,具有与模型预测控制方法相当的一致收敛性,(B)具有适应性,在环境输入信息不准确的情况下具有上级学习和控制效果;以及(c)可转移,具有充分的实施前景以及真实的时间适用性,以适应以增加不确定性为特征的未遇到的环境。
Perimeter metering control based on macroscopic fundamental diagrams has attracted increasing research interests over the past decade. This strategy provides a convenient way to mitigate urban congestion by manipulating vehicular movements across homogeneous regions without modeling the detailed behaviors and interactions involved with individual vehicle presence. In particular, multi-region perimeter metering control holds promise for efficient traffic management in large-scale urban networks. However, most existing methods for multi-region control require knowledge of either the environment traffic dynamics or network properties (i.e., the critical accumulations), whereas such information is generally difficult to obtain and subject to significant estimation errors. The recently developed model-free techniques, on the other hand, have not yet been shown scalable or applicable to large urban networks. To fill this gap, this paper proposes a scalable model-free scheme based on multi-agent deep reinforcement learning. The proposed scheme features value function decomposition in the paradigm of centralized training with decentralized execution, coupled with critical advances of single-agent deep reinforcement learning and problem reformulation guided by domain expertise. Comprehensive experiment results on a seven-region urban network suggest the scheme is: (a) effective, with consistent convergence to final control outcomes that are comparable to the model predictive control method; (b) resilient, with superior learning and control efficacy in the presence of inaccurate input information from the environment; and (c) transferable, with sufficient implementation prospect as well as real time applicability to unencountered environments featuring increased uncertainty.