Deep Reinforcement Learning for Crowdsourced Urban Delivery

Deep Reinforcement Learning for Crowdsourced Urban Delivery
复制标题

DOI:
10.1016/j.trb.2021.08.015
复制
发表时间:
2021-09-13
影响因子:
6.8
通讯作者:
Tulabandhula, Theja
Tulabandhula, Theja
中科院分区:
工程技术1区
文献类型:
--
作者:
Ahamed, Tanvir;Zou, Bo;Tulabandhula, Theja

文献摘要

被引文献

相似文献

本文研究了在城市众包配送的背景下,如何将配送请求分配给临时快递员的问题。发货请求在空间上分布,每个发货请求在最早的提货时间和最晚的发货时间之间有一个有限的时间窗口。临时快递员被称为众包,他们的时间可获得性和运载能力也有限。我们提出了一种新的基于深度强化学习(DRL)的方法来解决这个分配问题。训练了一种深度Q网络(DQN)算法,该算法具有经验回放和目标网络两个显著特征,提高了DRL训练的效率、收敛和稳定性。更重要的是,本文在方法上做出了三点贡献:1)提出了一种全面而新颖的众包系统状态刻画方法,它包含了众源和请求的时空和容量信息;2)嵌入了启发式算法,利用状态表示提供的信息,基于直观的推理来指导具体的行动,以保持易处理性,提高训练效率;3)结合规则插入,防止在路径改进过程中重复访问相同的路线和节点序列,从而通过加速学习进一步提高训练效率。研究了启发式算法和整体DQN训练的计算复杂性。通过大量的数值分析,验证了该方法的有效性。结果表明,在DRL训练中,启发式引导的动作选择、规则插入和在状态空间中具有与时间相关的信息所带来的好处,所获得的解的近最优性,以及该方法在解质量、计算时间和可伸缩性方面的优势。
This paper investigates the problem of assigning shipping requests to ad hoc couriers in the context of crowdsourced urban delivery. The shipping requests are spatially distributed each with a limited time window between the earliest time for pickup and latest time for delivery. The ad hoc couriers, termed crowdsourcees, also have limited time availability and carrying capacity. We propose a new deep reinforcement learning (DRL)-based approach to tackling this assignment problem. A deep Q network (DQN) algorithm is trained which entails two salient features of experience replay and target network that enhance the efficiency, convergence, and stability of DRL training. More importantly, this paper makes three methodological contributions: 1) presenting a comprehensive and novel characterization of crowdshipping system states that encompasses spatial-temporal and capacity information of crowdsourcees and requests; 2) embedding heuristics that leverage information offered by the state representation and are based on intuitive reasonings to guide specific actions to take, to preserve tractability and enhance efficiency of training; and 3) integrating rule-interposing to prevent repeated visiting of the same routes and node sequences during routing improvement, thereby further enhancing the training efficiency by accelerating learning. The computational complexities of the heuristics and the overall DQN training are investigated. The effectiveness of the proposed approach is demonstrated through extensive numerical analysis. The results show the benefits brought by the heuristics-guided action choice, rule-interposing, and having time-related information in the state space in DRL training, the near-optimality of the solutions obtained, and the superiority of the proposed approach over existing methods in terms of solution quality, computation time, and scalability.