Model-free perimeter metering control for two-region urban networks using deep reinforcement learning

Model-free perimeter metering control for two-region urban networks using deep reinforcement learning
复制标题

DOI:
10.1016/j.trc.2020.102949
复制
发表时间:
2021-03
影响因子:
8.3
通讯作者:
Dongqin Zhou;V. Gayah
Dongqin Zhou;V. Gayah
中科院分区:
工程技术1区
文献类型:
--
作者:
Dongqin Zhou;V. Gayah

文献摘要

被引文献

相似文献

针对城市交通网络提出了各种周界计量控制策略,这些策略依赖于网络生产力和累积之间存在明确的关系,通常称为网络宏观基本图(MFD)。大多数现有的周界计量控制策略需要对交通动态进行精确建模,并充分了解网络 MFD 和动态方程,以描述车辆如何跨网络区域移动。然而,此类信息通常难以获得并且容易出错。最近文献中提出了一些无模型周界计量控制方案。然而,这些现有方法需要估计控制器设计中的网络属性(例如,与最大网络生产率相关的临界累积)。本文提出了一种针对两区域城市网络的无模型深度强化学习边界控制(MFDRLPC)方案,该方案的特征是具有连续或离散动作空间的代理。所提出的代理通过强化学习过程学习选择控制动作,而无需假设任何有关环境动态的信息。大量数值实验的结果表明,所提出的智能体:(a)能够在各种环境配置下持续学习周界控制策略; (b) 性能与最先进的模型预测控制(MPC)相当; (c) 高度可移植到各种交通条件和环境动态。
Various perimeter metering control strategies have been proposed for urban traffic networks that rely on the existence of well-defined relationships between network productivity and accumulation, known more commonly as network Macroscopic Fundamental Diagrams (MFD). Most existing perimeter metering control strategies require accurate modeling of traffic dynamics with full knowledge of the network MFD and dynamic equations to describe how vehicles move across regions of the network. However, such information is generally difficult to obtain and subject to error. Some model free perimeter metering control schemes have been recently proposed in the literature. However, these existing approaches require estimates of network properties (e.g., the critical accumulation associated with maximum network productivity) in the controller designs. In this paper, a model free deep reinforcement learning perimeter control (MFDRLPC) scheme is proposed for two-region urban networks that features agents with either continuous or discrete action spaces. The proposed agents learn to select control actions through a reinforcement learning process without assuming any information about environment dynamics. Results from extensive numerical experiments demonstrate that the proposed agents: (a) can consistently learn perimeter control strategies under various environment configurations; (b) are comparable in performance to the state-of-the-art, model predictive control (MPC); and, (c) are highly transferable to a wide range of traffic conditions and dynamics in the environment.