Predicting Train Occupancies based on Query Logs and External Data Sources

Predicting Train Occupancies based on Query Logs and External Data Sources
复制标题

根据查询日志和外部数据源预测列车占用情况

DOI:
--
复制
发表时间:
2017
期刊:
The Web Conference
影响因子:
--
通讯作者:
F. Turck
F. Turck
中科院分区:
--
文献类型:
--
作者:
Gilles Vandewiele;Pieter Colpaert;Olivier Janssens;J. Herwegen;R. Verborgh;E. Mannens;F. Ongenae;F. Turck

文献摘要

被引文献

相似文献

在密集的铁路网络上,例如在比利时,火车旅客经常面临过度拥挤的火车,特别是在高峰时段。火车上的拥挤导致服务质量下降,对乘客的福祉产生负面影响。为了刺激旅客考虑不那么拥挤的列车,iRail项目希望通过预测建模的方式在他们的路线规划应用程序中显示占用指标。由于没有官方的占用数据,培训数据是通过使用iRail网络应用程序和iPhone的移动的Railer应用程序众包获得的。用户可以指出他们的出发和到达站,他们在什么时候乘坐火车,并将该列车的占用率分为:低,中或高。虽然对有限数据集的初步结果得出结论,这些模型的表现还不够好,但我们相信,通过进一步的研究和更大量的数据,我们的预测模型将能够实现更高的预测性能。为此,目前研究中使用的所有数据集都在iRail网站上以Kaggle竞赛的形式公开提供。此外,还建立了一个基础设施,可以自动处理用户提交的新日志,以便我们的模型不断学习。通过API可获得未来列车的占用预测。
On dense railway networks "such as in Belgium" train travelers are frequently confronted with overly occupied trains, especially during peak hours. Crowdedness on trains leads to a deterioration in the quality of service and has a negative impact on the well-being of the passenger. In order to stimulate travelers to consider less crowded trains, the iRail project wants to show an occupancy indicator in their route planning applications by the means of predictive modeling. As there is no official occupancy data available, training data is obtained by crowd-sourcing using the iRail web app and the mobile Railer application for iPhone. Users can indicate their departure & arrival station, at what time they took a train and classify the occupancy of that train into the classes: low, medium or high. While preliminary results on a limited dataset conclude that the models do not yet perform sufficiently well, we are convinced that with further research and a larger amount of data, our predictive model will be able to achieve higher predictive performances. All datasets used in the current research are, for that purpose, made publicly available under an open license on the iRail website and in the form of a Kaggle competition. Moreover, an infrastructure is set up that automatically processes new logs submitted by users in order for our model to continuously learn. Occupancy predictions for future trains are made available through an api.