A Method for Granular Traffic Data Imputation Based on PARATUCK2

A Method for Granular Traffic Data Imputation Based on PARATUCK2
复制标题

DOI:
10.1177/03611981221089298
复制
发表时间:
2022-10
影响因子:
1.7
通讯作者:
Mina Nouri;Mostafa Reisi-Gahrooei;Mohammad Ilbeigi
Mina Nouri;Mostafa Reisi-Gahrooei;Mohammad Ilbeigi
中科院分区:
工程技术4区
文献类型:
--
作者:
Mina Nouri;Mostafa Reisi-Gahrooei;Mohammad Ilbeigi

文献摘要

相似文献

缺失数据的录入是数据驱动的智能交通系统中的一项关键任务。近几十年来,在开发各种类型的传感器和智能系统方面进行了大量投资,包括固定设备(例如环路探测器)和配备全球定位系统(GPS)跟踪器的浮动车辆,以收集大规模交通数据。然而,由于不同的原因,收集的数据可能不包括交通网络中所有路段的观测,包括传感器故障、传输错误,以及因为配备GPS的车辆可能并不总是通过所有路段。开发实时交通监控和中断预测模型的第一步是通过系统的数据推算过程估计缺失值。许多现有的数据填充方法是基于利用交通数据固有的时空特性的矩阵补全技术。然而,这些方法可能不能完全捕获数据的集群结构。针对这一问题,提出了一种新的基于PARATUCK2分解的数据填充方法。该方法同时捕捉了交通数据的时空信息,构建了交通模式的低维聚类表示。所识别的时空簇被用来恢复网络流量简档并估计缺失值。所提出的方法是使用纽约市曼哈顿道路网的交通数据来实现的。通过与两种最先进的基准测试方法的比较,对该方法的性能进行了评估。结果表明,在复杂的大规模交通网络中,该方法的性能优于现有的最新方法。
Imputing missing data is a critical task in data-driven intelligent transportation systems. During recent decades there has been a considerable investment in developing various types of sensors and smart systems, including stationary devices (e.g., loop detectors) and floating vehicles equipped with global positioning system (GPS) trackers to collect large-scale traffic data. However, collected data may not include observations from all road segments in a traffic network for different reasons, including sensor failure, transmission error, and because GPS-equipped vehicles may not always travel through all road segments. The first step toward developing real-time traffic monitoring and disruption prediction models is to estimate missing values through a systematic data imputation process. Many of the existing data imputation methods are based on matrix completion techniques that utilize the inherent spatiotemporal characteristics of traffic data. However, these methods may not fully capture the clustered structure of the data. This paper addresses this issue by developing a novel data imputation method using PARATUCK2 decomposition. The proposed method captures both spatial and temporal information of traffic data and constructs a low-dimensional and clustered representation of traffic patterns. The identified spatiotemporal clusters are used to recover network traffic profiles and estimate missing values. The proposed method is implemented using traffic data in the road network of Manhattan in New York City. The performance of the proposed method is evaluated in comparison with two state-of-the-art benchmark methods. The outcomes indicate that the proposed method outperforms the existing state-of-the-art imputation methods in complex and large-scale traffic networks.