Arterial traffic control using reinforcement learning agents and information from adjacent intersections in the state and reward structure

Arterial traffic control using reinforcement learning agents and information from adjacent intersections in the state and reward structure
复制标题

使用强化学习代理和状态和奖励结构中相邻交叉口的信息进行主干道交通控制

DOI:
--
复制
发表时间:
2010
期刊:
13th International IEEE Conference on Intelligent Transportation Systems
影响因子:
--
通讯作者:
R. Benekohal
R. Benekohal
中科院分区:
--
文献类型:
--
作者:
J. Medina;Ali Hajbabaie;R. Benekohal

文献摘要

被引文献

相似文献

应用强化学习(RL)代理的交通控制沿着动脉下的高交通量。RL代理使用Q学习和状态表示的修改版本进行训练,该状态表示包括来自相邻交叉口的链路占用信息。建议的结构还包括一个奖励,考虑潜在的阻塞从下游交叉口(由于饱和条件),以及压力,以协调信号响应与未来到达的交通从上游交叉口。采用微观仿真软件对一条5交叉口的主干道在高冲突流量下进行了实验,并与最佳的协调预定时相位设置进行了比较。数据显示,RL代理的延迟更低,停靠次数更少,系统中所有车辆的延迟分布也更均衡。协调行为的证据被发现,因为通过5个交叉口的停车次数平均低于1.5,而且所有交叉口的绿色时间分布非常相似。然而,随着交通量接近容量,提前定相的延误低于RL代理,但代理产生更低的最大延误时间和每辆车的最大停靠次数。未来的研究将分析系统的状态和奖励结构中的可变系数,以更好地科普各种各样的交通量,包括从过饱和到欠饱和的过渡,反之亦然。
An application that uses reinforcement learning (RL) agents for traffic control along an arterial under high traffic volumes is presented. RL agents were trained using Q learning and a modified version of the state representation that included information on the occupancy of the links from neighboring intersections. The proposed structure also includes a reward that considers potential blockage from downstream intersections (due to saturated conditions), as well as pressure to coordinate the signal response with the future arrival of traffic from upstream intersections. Experiments using microscopic simulation software were conducted for an arterial with 5 intersections under high conflicting volumes, and results were compared with the best settings of coordinated pre-timed phasing. Data showed lower delays and less number of stops with RL agents, as well as a more balanced distribution of the delay among all vehicles in the system. Evidence of coordinated-like behavior was found as the number of stops to traverse the 5 intersections was on average lower than 1.5, and also since the distribution of green times from all intersections was very similar. As traffic approached to capacity, however, delays with the pre-timed phasing were lower than with RL agents, but the agents produced lower maximum delay times and lower maximum number of stops per vehicle. Future research will analyze variable coefficients in the state and reward structures for the system to better cope with a wide variety of traffic volumes, including transitions from oversaturation to undersaturation and vice versa.