Graph Meta-Reinforcement Learning for Transferable Autonomous Mobility-on-Demand

Graph Meta-Reinforcement Learning for Transferable Autonomous Mobility-on-Demand
复制标题

用于按需可转移自主移动的图元强化学习

DOI:
10.1145/3534678.3539180
复制
发表时间:
2022
期刊:
Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Pavone, Marco
Pavone, Marco
中科院分区:
--
文献类型:
--
作者:
Gammelli, Daniele;Yang, Kaidi;Harrison, James;Rodrigues, Filipe;Pereira, Francisco;Pavone, Marco

文献摘要

参考文献

被引文献

相似文献

自主移动按需(阿莫德)系统代表了现有交通模式的一种有吸引力的替代方案,目前受到城市化和日益增长的旅行需求的挑战。通过集中控制自动驾驶车队,这些系统为客户提供移动服务,目前已开始在全球多个城市部署。当前用于控制阿莫德系统的基于学习的方法仅限于单个城市场景,由此允许服务运营商在同一交通系统内采取无限量的运营决策。然而,现实世界的系统运营商几乎无法负担为他们运营的每个城市完全重新培训阿莫德控制器的费用,因为这可能导致培训期间大量的低质量决策,使单一城市策略成为一个潜在的不切实际的解决方案。为了解决这些局限性,我们建议通过元强化学习(meta-RL)的透镜来形式化多城市的阿莫德问题,并设计一个基于递归图神经网络的演员-评论家算法。在我们的方法中,阿莫德控制器被明确地训练,使得在新城市中的少量经验将产生良好的系统性能。从经验上讲,我们展示了通过元RL学习的控制策略如何能够通过学习快速适应的策略在看不见的城市中实现接近最优的性能,从而使它们不仅对新环境更具鲁棒性,而且对现实世界运营中常见的分布变化更具鲁棒性,例如特殊事件,意外拥堵和动态定价方案。
Autonomous Mobility-on-Demand (AMoD) systems represent an attractive alternative to existing transportation paradigms, currently challenged by urbanization and increasing travel needs. By centrally controlling a fleet of self-driving vehicles, these systems provide mobility service to customers and are currently starting to be deployed in a number of cities around the world. Current learning-based approaches for controlling AMoD systems are limited to the single-city scenario, whereby the service operator is allowed to take an unlimited amount of operational decisions within the same transportation system. However, real-world system operators can hardly afford to fully re-train AMoD controllers for every city they operate in, as this could result in a high number of poor-quality decisions during training, making the single-city strategy a potentially impractical solution. To address these limitations, we propose to formalize the multi-city AMoD problem through the lens of meta-reinforcement learning (meta-RL) and devise an actor-critic algorithm based on recurrent graph neural networks. In our approach, AMoD controllers are explicitly trained such that a small amount of experience within a new city will produce good system performance. Empirically, we show how control policies learned through meta-RL are able to achieve near-optimal performance on unseen cities by learning rapidly adaptable policies, thus making them more robust not only to novel environments, but also to distribution shifts common in real-world operations, such as special events, unexpected congestion, and dynamic pricing schemes.
当前矩阵乘法时间的确定性线性规划求解器
DOI: 10.1137/1.9781611975994.16
发表时间: 2019
期刊: ArXiv
影响因子: --
作者:
Jan van den Brand
通讯作者: Jan van den Brand
DOI: 10.1145/3474837
发表时间: 2022-01
期刊: ACM Transactions on Intelligent Systems and Technology (TIST)
影响因子: --
作者:
Yingxue Zhang;Yanhua Li;Xun Zhou;Jun Luo;Zhi-Li Zhang
通讯作者: Yingxue Zhang;Yanhua Li;Xun Zhou;Jun Luo;Zhi-Li Zhang
DOI: 10.1136/ebmh.11.4.102
发表时间: 2008-10
期刊: Evidence Based Mental Health
影响因子: --
作者:
P. Cochat;L. Vaucoret;J. Sarles
通讯作者: P. Cochat;L. Vaucoret;J. Sarles
按需移动系统中混合车队的实时控制
DOI: 10.1109/itsc48978.2021.9564770
发表时间: 2021
期刊: Intelligent Transportation Systems Conference
影响因子: --
作者:
Yang, Kaidi;Tsao, Matthew W.;Xu, Xin;Pavone, Marco
通讯作者: Pavone, Marco
具有动态网络加载和动态乘车共享应用程序的共享自动驾驶车辆建模通用框架
DOI: 10.1016/j.compenvurbsys.2017.04.006
发表时间: 2017
期刊: Comput. Environ. Urban Syst.
影响因子: --
作者:
Michael W. Levin;Kara Kockelman;S. Boyles;Tianxin Li
通讯作者: Tianxin Li