A Data-Driven Method for Trip Ends Identification Using Large-Scale Smartphone-Based GPS Tracking Data

A Data-Driven Method for Trip Ends Identification Using Large-Scale Smartphone-Based GPS Tracking Data
复制标题

使用基于智能手机的大规模 GPS 跟踪数据进行行程终点识别的数据驱动方法

DOI:
10.1109/tits.2016.2630733
复制
发表时间:
2017-08
影响因子:
8.5
通讯作者:
Xiao Guangnian
Xiao Guangnian
中科院分区:
工程技术1区
文献类型:
--
作者:
Zhou Chaoran;Jia Hongfei;Juan Zhicai;Fu Xuemei;Xiao Guangnian

文献摘要

参考文献

被引文献

相似文献

利用从智能手机和互联网调查中获得的跟踪数据,提出了一种数据驱动的机器学习方法来识别旅行终点。在以前的文献中,这通常是基于一些预定义的规则,已被证实是有效的。然而,这些基于规则的方法在很大程度上依赖于研究人员自己的知识,这不可避免地是主观和任意的。此外,它们在处理大数据时代的海量数据方面不够有效。本文针对数百万基于智能手机的GPS跟踪数据。一组属性,如旅行速度,距离和航向,推导出表征智能手机持有人的旅行状态。换句话说,跟踪点可以被识别为处于行驶或非行驶状态,基于该状态可以容易地检测行程结束。与基于规则的方法不同,本文采用随机森林作为分类模型,没有预先定义的主观分类规则。这个数据驱动的模型是自动构建的。结果表明,利用随机森林对1393 d的GPS跟踪数据和PR调查数据进行训练后,对697 d的跟踪数据的行程终点识别准确率达到96.17%。目前的分析是免费的个人经验,这将是有用的,基于智能手机的调查数据在大数据时代。
Using tracking data obtained from the smartphone and Internet survey, a data-driven machine learning method is proposed to identify trip ends. In previous literature, this is usually done based on some predefined rules, which have been confirmed to be valid. Nonetheless, these rule-based methods largely depend on researchers’ own knowledge, which is inevitably subjective and arbitrary. Moreover, they are not effective enough to process the huge amount of data in the era of big data. In this paper, millions of smartphone-based GPS tracking data are targeted. A group of attributes, such as travel speed, distance, and heading, are derived to characterize the smartphone holders’ travel status. In other words, the tracking points could be identified as being at the state of traveling or non-traveling, based on which the trip ends are easily detected. In contrast to those rule-based methods, a random forest is utilized in this paper as the classification model, with no subjective rules predefined for classification. This data-driven model is automatically built. The results show that after training the GPS tracking data of 1393 days and the prompted recall (PR) survey data using the random forest, the accuracy of trip ends identification on tracking data of 697 days is 96.17%. The current analysis is free from personal experiences, which is expected to be useful for the smartphone-based survey data in the era of big data.
DOI: 10.1177/0361198105191700108
发表时间: 2005
影响因子: 1.7
作者:
T. Forrest;D. Pearson
通讯作者: T. Forrest;D. Pearson
DOI: 10.3141/1917-08
发表时间: 2005-01-01
期刊: DATA INITIATIVES
影响因子: --
作者:
Forrest, TL;Pearson, DF
通讯作者: Pearson, DF
DOI: 10.1016/j.compenvurbsys.2015.05.005
发表时间: 2015-11
期刊: Comput. Environ. Urban Syst.
影响因子: --
作者:
Guangnian Xiao;Z. Juan;Chunqin Zhang
通讯作者: Guangnian Xiao;Z. Juan;Chunqin Zhang
DOI: 10.1007/978-1-4419-9326-7_5
发表时间: 2012-01-01
期刊: ENSEMBLE MACHINE LEARNING: METHODS AND APPLICATIONS
影响因子: --
作者:
Cutler, Adele;Cutler, D. Richard;Stevens, John R.
通讯作者: Stevens, John R.
DOI: 10.3328/tl.2009.01.01.59-79
发表时间: 2009-01-01
影响因子: 2.8
作者:
Auld, Joshua;Williams, Chad;Nelson, Peter
通讯作者: Nelson, Peter