Multi-Armed Bandit On-Time Arrival Algorithms for Sequential Reliable Route Selection under Uncertainty

Multi-Armed Bandit On-Time Arrival Algorithms for Sequential Reliable Route Selection under Uncertainty
复制标题

DOI:
10.1177/0361198119850457
复制
发表时间:
2019-06
影响因子:
1.7
通讯作者:
Jinkai Zhou;Xuebo Lai;Joseph Y. J. Chow
Jinkai Zhou;Xuebo Lai;Joseph Y. J. Chow
中科院分区:
工程技术4区
文献类型:
--
作者:
Jinkai Zhou;Xuebo Lai;Joseph Y. J. Chow

文献摘要

被引文献

相似文献

传统上,车辆在运送乘客和货物时只起到服务生的作用。随着包括自动车辆在内的车辆中传感器设备的增加,需要测试考虑车辆既是服务器又是传感器的双重角色的算法。本文将序列路径选择问题描述为一种强化学习模型--多臂强盗环境下的具有准时到达可靠性的最短路径问题。决策者必须按顺序对固定的始发地-目的地对之间的出发时间和路径做出有限集合的决策,使得准点可靠性最大化而旅行时间最小化。扩展了上置信限算法以处理该问题。进行了几次测试。首先,模拟数据成功地验证了该方法,然后构建了一个真实数据场景,即从纽约市曼哈顿市中心提供每小时一次的酒店班车服务,前往约翰·F·肯尼迪国际机场。结果表明,采用多臂强盗学习算法的路径选择是有效的,但忽略乘客调度约束对正点到达可靠性的负面影响高达4.8%,对综合可靠性和行程时间的负面影响高达66.1%。
Traditionally vehicles act only as servers in transporting passengers and goods. With increasing sensor equipment in vehicles, including automated vehicles, there is a need to test algorithms that consider the dual role of vehicles as both servers and sensors. The paper formulates a sequential route selection problem as a shortest path problem with on-time arrival reliability under a multi-armed bandit setting, a type of reinforcement learning model. A decision-maker has to make a finite set of decisions sequentially on departure time and path between a fixed origin-destination pair such that on-time reliability is maximized while travel time is minimized. The upper confidence bound algorithm is extended to handle this problem. Several tests are conducted. First, simulated data successfully verifies the method, then a real-data scenario is constructed of a hotel shuttle service from midtown Manhattan in New York City providing hourly access to John F. Kennedy International Airport. Results suggest that route selection with multi-armed bandit learning algorithms can be effective but neglecting passenger scheduling constraints can have negative effects on on-time arrival reliability by as much as 4.8% and combined reliability and travel time by 66.1%.