Adaptive railway traffic control using approximate dynamic programming

Adaptive railway traffic control using approximate dynamic programming
复制标题

DOI:
10.1016/j.trpro.2019.05.012
复制
发表时间:
2019-12
期刊:
Transportation Research Part C: Emerging Technologies
影响因子:
--
通讯作者:
T. Ghasempour;B. Heydecker
T. Ghasempour;B. Heydecker
中科院分区:
其他
文献类型:
--
作者:
T. Ghasempour;B. Heydecker

文献摘要

被引文献

相似文献

本研究提出了一种基于近似动态规划(ADP)的实时自适应铁路交通控制器。通过评估需求和机会,控制人员旨在通过及时在关键位置对列车进行排序来限制因晚于计划进入控制区域的列车而导致的连续延误,从而代表铁路运营的实际要求。这种方法依赖于从指定状态优化后动态规划价值函数的近似值,该函数是使用强化学习技术根据操作经验动态估计的。通过使用这种近似,ADP 避免了对性能的广泛显式评估,从而大大减少了计算负担。在这项研究中,我们探索了近似函数的公式以及用于估计它的学习技术的变体。在随机模拟环境中对 ADP 方法进行的评估表明,与当前先到先服务排序的行业实践相比,连续延迟方面有相当大的改进。我们还发现,在具有不同平均列车进入延误的一系列测试场景中,近似值函数的参数估计是相似的。
This study presents an adaptive railway traffic controller for real-time operations based on approximate dynamic programming (ADP). By assessing requirements and opportunities, the controller aims to limit consecutive delays resulting from trains that entered a control area behind schedule by sequencing them at critical locations in a timely manner, thus representing the practical requirements of railway operations. This approach depends on an approximation to the value function of dynamic programming after optimisation from a specified state, which is estimated dynamically from operational experience using reinforcement learning techniques. By using this approximation, the ADP avoids extensive explicit evaluation of performance and so reduces the computational burden substantially. In this investigation, we explore formulations of the approximation function and variants of the learning techniques used to estimate it. Evaluation of the ADP methods in a stochastic simulation environment shows considerable improvements in consecutive delays by comparison with the current industry practice of First-Come-First-Served sequencing. We also found that estimates of parameters of the approximate value function are similar across a range of test scenarios with different mean train entry delays.