Deep Reinforcement Learning Algorithm for Dynamic Pricing of Express Lanes with Multiple Access Locations

Deep Reinforcement Learning Algorithm for Dynamic Pricing of Express Lanes with Multiple Access Locations
复制标题

DOI:
10.1016/j.trc.2020.102715
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Venktesh Pandey;Evana Wang;S. Boyles
Venktesh Pandey;Evana Wang;S. Boyles
中科院分区:
其他
文献类型:
--
作者:
Venktesh Pandey;Evana Wang;S. Boyles

文献摘要

被引文献

相似文献

本文开发了一个深度强化学习(Deep-RL)框架,用于管理车道上的动态定价,该车道具有多个访问位置和旅行者时间,起源和目的地价值的异质性。该框架放宽了文献中的假设,考虑多个起源和目的地,多个访问位置的管理车道,enroutedevoration的旅客,部分可观测性的传感器读数,随机需求和观察。该问题被制定为一个部分可观察的马尔可夫决策过程(POMDP)和政策梯度的方法来确定收费作为一个功能的实时观测。通行费被建模为连续和随机变量,并使用前馈神经网络确定。该方法进行了比较,对用于动态定价的反馈控制方法。我们表明,当在现实世界的交通网络上进行测试时,Deep-RL在学习最大化收入、最小化总系统旅行时间和其他联合加权目标的收费政策方面是有效的。Deep-RL收费策略通过产生比启发式算法高8.5%的收入来实现收入最大化目标,并且通过产生比启发式算法低8.4%的TSTT来实现最小化总系统行驶时间(TSTT)的目标,从而优于反馈控制启发式算法。我们还提出了POMDP的奖励成形方法,以克服不希望的行为收费政策,如收入最大化政策的干扰和收获行为。此外,我们测试了在一组输入上训练的算法对于新输入分布的可移植性,并提供了关于Deep-RL算法实时实现的建议。我们实验的源代码可以在https://github.com/venktesh22/ExpressLanes_Deep-RL上在线获得。
This article develops a deep reinforcement learning (Deep-RL) framework for dynamic pricing on managed lanes with multiple access locations and heterogeneity in travelers’ value of time, origin, and destination. This framework relaxes assumptions in the literature by considering multiple origins and destinations, multiple access locations to the managed lane,en routediversion of travelers, partial observability of the sensor readings, and stochastic demand and observations. The problem is formulated as a partially observable Markov decision process (POMDP) and policy gradient methods are used to determine tolls as a function of real-time observations. Tolls are modeled as continuous and stochastic variables and are determined using a feedforward neural network. The method is compared against a feedback control method used for dynamic pricing. We show that Deep-RL is effective in learning toll policies for maximizing revenue, minimizing total system travel time, and other joint weighted objectives, when tested on real-world transportation networks. The Deep-RL toll policies outperform the feedback control heuristic for the revenue maximization objective by generating revenues up to 8.5% higher than the heuristic and for the objective minimizing total system travel time (TSTT) by generating TSTT up to 8.4% lower than the heuristic. We also propose reward shaping methods for the POMDP to overcome the undesired behavior of toll policies, like thejam-and-harvestbehavior of revenue-maximizing policies. Additionally, we test transferability of the algorithm trained on one set of inputs for new input distributions and offer recommendations on real-time implementations of Deep-RL algorithms. The source code for our experiments is available online at https://github.com/venktesh22/ExpressLanes_Deep-RL.