Joint Charging and Relocation Recommendation for E-Taxi Drivers via Multi-Agent Mean Field Hierarchical Reinforcement Learning

Joint Charging and Relocation Recommendation for E-Taxi Drivers via Multi-Agent Mean Field Hierarchical Reinforcement Learning
复制标题

通过多智能体平均场分层强化学习为电动出租车司机提供联合收费和搬迁建议

DOI:
10.1109/tmc.2020.3022173
复制
发表时间:
2022-04
影响因子:
7.9
通讯作者:
Xinbing Wang
Xinbing Wang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Enshu Wang;Rong Ding;Zhaoxing Yang;Haiming Jin;Chenglin Miao;Lu Su;Fan Zhang;Chunming Qiao;Xinbing Wang

文献摘要

参考文献

相似文献

如今,大多数出租车司机已经成为在线乘车平台提供的搬迁推荐服务的用户(例如,Uber和滴滴出行(Didi Chuxing),这通常可以将司机带到有利润订单的地方。与此同时,电动出租车(e-taxi)越来越多地被采用,并逐渐取代汽油出租车在今天的公共交通系统中,由于其环保的性质。虽然对传统的汽油出租车有效,但现有的搬迁推荐方案对于电动出租车司机的用户体验来说是次优的。一方面,现有的方案没有考虑到出租车的加油决策,因为汽油出租车的加油时间通常很短,可以忽略不计。然而,电动出租车在充电站的充电时间可能长达数小时。显然,电动出租车的电池很容易被现有计划所建议的连续搬迁所耗尽,因此必须在很长一段时间后充电,使电动出租车司机错过了许多订单服务机会。另一方面,充电桩通常稀疏且不均匀地分布在整个城市中。如果不考虑充电机会,现有的计划可能会将电动出租车送到一个没有充电桩的地区,即使它的电池电量不足。为了优化电动出租车司机的用户体验,本文设计了一个电动出租车司机联合充电和搬迁推荐系统(CARE)。我们从电动出租车司机的角度出发,将他们的决策制定为一个多智能体强化学习问题,每个电动出租车司机的目标是最大化自己的累积奖励。更具体地说,我们提出了一种新的多智能体平均场层次强化学习(MFHRL)框架。MFHRL的分层架构有助于拟议的CARE为电动出租车司机提供有远见的收费和搬迁建议。此外,我们整合每个层次的MFHRL分别与平均场近似,以考虑在决策中的电子出租车的相互影响。我们使用中国深圳最大的真实世界电动出租车数据集之一建立了一个模拟器,其中包含2017年6月1日至6月30日3848辆电动出租车的GPS轨迹数据和交易数据,以及165个充电站,包括317个快速充电桩和1421个慢速充电桩。我们采用该模拟器生成6个动态城市环境,反映了电动出租车司机面临的不同现实场景。在所有这些环境中,我们进行了大量的实验,以验证所提出的MFHRL框架大大优于所有基线显着增加的奖励获得的电动出租车司机。此外,我们还表明,MFHRL学习的收费政策可以有效地降低电动出租车司机的里程焦虑,这显着提高电动出租车司机的体验质量。
Nowadays, most of the taxi drivers have become users of the relocation recommendation service offered by online ride-hailing platforms (e.g., Uber and Didi Chuxing), which could oftentimes lead drivers to places with profitable orders. At the same time, electric taxis (e-taxis) are increasingly adopted and gradually replacing gasoline taxis in today’s public transportation systems due to their environmental-friendly nature. Though effective for traditional gasoline taxis, existing relocation recommendation schemes are rather suboptimal for e-taxi drivers’ user experience. On one hand, the existing schemes take no account of taxis’ refueling decisions, as the refueling durations of gasoline taxis are usually short enough to be ignored. However, the charging duration of the e-taxis spent at charging stations can be as long as hours. Obviously, an e-taxi’s battery could be easily depleted by the continuous relocations suggested by existing schemes, and thus will have to be charged for a long time afterwards, making the e-taxi driver miss numerous order-serving opportunities. On the other hand, charging posts are typically sparsely and unevenly distributed across a city. With no consideration of charging opportunities, existing schemes could probably send an e-taxi to an area with no charging post around, even though its battery is running low. To optimize e-taxi drivers’ user experience, in this paper, we design a joint charging and relocation recommendation system for e-taxi drivers (CARE). We take the perspective of e-taxi drivers and formulate their decision making as a multi-agent reinforcement learning problem where each e-taxi driver aims to maximize his own cumulative rewards. More specifically, we propose a novel multi-agent mean field hierarchical reinforcement learning (MFHRL) framework. The hierarchical architecture of MFHRL helps the proposed CARE provide far-sighted charging and relocation recommendations for e-taxi drivers. Besides, we integrate each hierarchical level of MFHRL separately with the mean field approximation to incorporate e-taxis’ mutual influences in decision making. We set up a simulator with one of the largest real-world e-taxi datasets in Shenzhen, China, which contains the GPS trajectory data and transaction data of 3848 e-taxis from June 1st to June 30th, 2017, coupled with 165 charging stations including 317 fast charging posts and 1421 slow charging posts. We adopt this simulator to generate 6 dynamic urban environments, which reflect the different real-world scenarios faced by e-taxi drivers. In all of these environments, we conduct extensive experiments to validate that the proposed MFHRL framework greatly outperforms all baselines by significantly increasing the rewards obtained by e-taxi drivers. Besides, we also show that the charging policy learned by MFHRL can effectively reduce the range anxiety of e-taxi drivers, which significantly boosts e-taxi drivers’ quality of experience.
DOI: 10.1109/infocom.2018.8486414
发表时间: 2018-04
期刊: IEEE INFOCOM 2018 - IEEE Conference on Computer Communications
影响因子: --
作者:
Yongmin Zhang;Jiayi Chen;Lin X. Cai;Jianping Pan
通讯作者: Yongmin Zhang;Jiayi Chen;Lin X. Cai;Jianping Pan
DOI: 10.1109/rtss.2016.016
发表时间: 2016-11
期刊: 2016 IEEE Real-Time Systems Symposium (RTSS)
影响因子: --
作者:
Fanxin Kong;Qiao Xiang;L. Kong;Xue Liu
通讯作者: Fanxin Kong;Qiao Xiang;L. Kong;Xue Liu
DOI: 10.1145/2971648.2971664
发表时间: 2016-09
期刊: Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing
影响因子: --
作者:
Mengwen Xu;Dong Wang-;Jian Li
通讯作者: Mengwen Xu;Dong Wang-;Jian Li
DOI: --
发表时间: 2017-03
期刊: ArXiv
影响因子: --
作者:
A. Vezhnevets;Simon Osindero;T. Schaul;N. Heess;Max Jaderberg;David Silver;K. Kavukcuoglu
通讯作者: A. Vezhnevets;Simon Osindero;T. Schaul;N. Heess;Max Jaderberg;David Silver;K. Kavukcuoglu
DOI: 10.1145/3267305.3274161
发表时间: 2018-10
期刊: Proceedings of the 2018 ACM International Joint Conference and 2018 International Symposium on Pervasive and Ubiquitous Computing and Wearable Computers
影响因子: --
作者:
Teerawat Kumsila;S. Phithakkitnukoon
通讯作者: Teerawat Kumsila;S. Phithakkitnukoon