A Robust and Constrained Multi-Agent Reinforcement Learning Electric Vehicle Rebalancing Method in AMoD Systems

A Robust and Constrained Multi-Agent Reinforcement Learning Electric Vehicle Rebalancing Method in AMoD Systems
复制标题

DOI:
10.1109/iros55552.2023.10342342
复制
发表时间:
2022-09
期刊:
2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Sihong He;Yue Wang;Shuo Han;Shaofeng Zou;Fei Miao
Sihong He;Yue Wang;Shuo Han;Shaofeng Zou;Fei Miao
中科院分区:
其他
文献类型:
--
作者:
Sihong He;Yue Wang;Shuo Han;Shaofeng Zou;Fei Miao

文献摘要

相似文献

电动汽车 (EV) 在自主按需出行 (AMoD) 系统中发挥着关键作用,但其独特的充电模式增加了 AMoD 系统中的模型不确定性(例如状态转换概率)。由于训练和测试/真实环境之间通常存在不匹配,因此将模型不确定性纳入系统设计在实际应用中至关重要。然而,现有文献尚未在EV AMoD系统再平衡中明确考虑模型不确定性,模型不确定性与决策应满足的约束条件并存使得问题更具挑战性。在这项工作中,我们为 EV AMoD 系统设计了一个鲁棒且受限的多智能体强化学习 (MARL) 框架,该框架具有状态转换内核不确定性。然后,我们提出了一种具有鲁棒自然策略梯度(RNPG)的鲁棒约束MARL算法(ROCOMA),该算法训练鲁棒的电动汽车再平衡策略,以在模型不确定性下平衡整个城市的供需比和充电利用率。实验表明,ROCOMA 可以学习有效且稳健的再平衡策略。在存在模型不确定性的情况下,它的性能优于非鲁棒 MARL 方法。它将系统公平性提高了 19.6%,重新平衡成本降低了 75.8%。
Electric vehicles (EVs) play critical roles in autonomous mobility-on-demand (AMoD) systems, but their unique charging patterns increase the model uncertainties in AMoD systems (e.g. state transition probability). Since there usually exists a mismatch between the training and test/true environments, incorporating model uncertainty into system design is of critical importance in real-world applications. However, model uncertainties have not been considered explicitly in EV AMoD system rebalancing by existing literature yet, and the coexistence of model uncertainties and constraints that the decision should satisfy makes the problem even more challenging. In this work, we design a robust and constrained multi-agent reinforcement learning (MARL) framework with state transition kernel uncertainty for EV AMoD systems. We then propose a robust and constrained MARL algorithm (ROCOMA) with robust natural policy gradients (RNPG) that trains a robust EV rebalancing policy to balance the supply-demand ratio and the charging utilization rate across the city under model uncertainty. Experiments show that the ROCOMA can learn an effective and robust rebalancing policy. It outperforms non-robust MARL methods in the presence of model uncertainties. It increases the system fairness by 19.6% and decreases the rebalancing costs by 75.8%.