A novel variable selection method based on frequent pattern tree for real-time traffic accident risk prediction

A novel variable selection method based on frequent pattern tree for real-time traffic accident risk prediction
复制标题

DOI:
10.1016/j.trc.2015.03.015
复制
发表时间:
2015-06-01
影响因子:
8.3
通讯作者:
Sadek, Adel W.
Sadek, Adel W.
中科院分区:
工程技术1区
文献类型:
--
作者:
Lin, Lei;Wang, Qian;Sadek, Adel W.

文献摘要

被引文献

相似文献

随着大量的实时交通流数据沿着交通事故信息的出现,人们对交通事故风险实时预测模型的开发重新产生了兴趣。然而,一个挑战是,现有的数据通常是复杂的,嘈杂的,甚至误导。这就提出了一个问题,如何选择最重要的解释变量,以达到可接受的准确度水平的实时交通事故风险预测。针对这一问题,提出了一种基于频繁模式树(FP树)的变量选择方法。该方法的工作原理是首先识别交通事故数据集中的所有频繁模式。接下来,对于每个频繁模式,我们引入一个新的度量,在本文中称为相对对象纯度比(ROPR)。然后,使用ROPR来计算每个解释变量的重要性得分,该重要性得分又可以用于排名和选择对解释事故模式贡献最大的变量。为了证明所提出的变量选择方法的优点,该研究开发了两个交通事故风险预测模型,基于事故数据收集的州际公路1-64在弗吉尼亚州,即一个k-近邻模型和贝叶斯网络。在模型开发之前,使用两种变量选择方法:(1)本文提出的基于FP树的方法;以及(2)随机森林方法,一种广泛使用的变量选择方法,用作比较的基本情况。结果表明,无论预测模型的类型(k-最近邻或贝叶斯网络)、参数设置以及用于模型训练和测试的数据类型如何,基于FP树的事故风险预测模型的性能都优于基于随机森林的模型。最好的模型是基于FP树的贝叶斯网络模型,可以预测61.11%的事故,而误报率为38.16%。这些结果与文献中报道的其他事故预测模型相比非常有利。(C)2015爱思唯尔有限公司版权所有。
With the availability of large volumes of real-time traffic flow data along with traffic accident information, there is a renewed interest in the development of models for the real-time prediction of traffic accident risk. One challenge, however, is that the available data are usually complex, noisy, and even misleading. This raises the question of how to select the most important explanatory variables to achieve an acceptable level of accuracy for real-time traffic accident risk prediction. To address this, the present paper proposes a novel Frequent Pattern tree (FP tree) based variable selection method. The method works by first identifying all the frequent patterns in the traffic accident dataset Next, for each frequent pattern, we introduce a new metric, herein referred to as the Relative Object Purity Ratio (ROPR). The ROPR is then used to calculate the importance score of each explanatory variable which in turn can be used for ranking and selecting the variables that contribute most to explaining the accident patterns. To demonstrate the advantages of the proposed variable selection method, the study develops two traffic accident risk prediction models, based on accident data collected on interstate highway 1-64 in Virginia, namely a k-nearest neighbor model and a Bayesian network. Prior to model development, two variable selection methods are utilized: (1) the FP tree based method proposed in this paper; and (2) the random forest method, a widely used variable selection method, which is used as the base case for comparison. The results show that the FP tree based accident risk prediction models perform better than the random forest based models, regardless of the type of prediction models (i.e. k-nearest neighbor or Bayesian network), the settings of their parameters, and the types of datasets used for model training and testing. The best model found is a FP tree based Bayesian network model that can predict 61.11% of accidents while having a false alarm rate of 38.16%. These results compare very favorably with other accident prediction models reported in the literature. (C) 2015 Elsevier Ltd. All rights reserved.