Traffic At-a-Glance: Time-Bounded Analytics on Large Visual Traffic Data

Traffic At-a-Glance: Time-Bounded Analytics on Large Visual Traffic Data
复制标题

DOI:
10.1109/tpds.2017.2684158
复制
发表时间:
2017-09-01
影响因子:
5.3
通讯作者:
Zhao, Wei
Zhao, Wei
中科院分区:
计算机科学2区
文献类型:
--
作者:
Li, Gang;Li, Xinfeng;Zhao, Wei

文献摘要

被引文献

相似文献

最近出现了大量的视觉交通数据。尽管它打开了智能流量分析的领域,但及时处理数据对于时间敏感的决策来说很困难,但至关重要,这对于交通相关管理来说是典型的。在本文中,我们研究了大型视觉交通数据(包括交通图像和视频)的时间限制聚合分析。我们首先发现当前的MapReduce框架由于两个挑战而不能很好地工作:首先,数据分布和处理时间上存在显着的双重多样性;其次,关于这些分布和时间成本的先验知识并不总是可用。然而,我们还观察数据值和处理时间的空间和时间局部性。根据检查,我们设计了 Traffic At-a-Glance (TaG),这是一个用于有时间限制的流量分析作业的增强型 MapReduce 框架。特别是,我们提出了一种新颖的采样算法,该算法利用流量数据局部性并根据数据分布和处理时间对样本进行分层。它以迭代、自适应的方式运行,无需先验知识。此外,我们提出了一种考虑批处理开销的启发式调度算法。此外,我们根据数据处理时间局部性改进了负载平衡机制,以尊重作业时间限制。此外,我们还扩展了 TaG,通过根据视频中编码的运动信息对视频数据进行采样来很好地处理交通视频。我们在 Hadoop 上实现 TaG,并在大型视觉流量数据集上进行了广泛的实验。对不同数据大小的评估表明,TaG 能够在时间范围内实现高精度。
Massive visual traffic data have become available recently. Though it opens the realm of intelligent traffic analysis, processing the data in a timely manner is difficult yet critical to time sensitive decisions, which are typical to traffic related management. In this paper, we study time-bounded aggregation analytics on large visual traffic data including traffic images and videos. We first find that current MapReduce framework can not work well due to two challenges: first, significant dual diversities exist on data distributions and processing time; second, apriori knowledge on these distributions and time costs are not always available. However, we also observe spatial and temporal locality on data values and processing time. Based on the examination, we design Traffic At-a-Glance (TaG), an augmented MapReduce framework for time-bounded traffic analytics jobs. Particularly, we propose a novel sampling algorithm that exploits traffic data localities and stratifies samples based on data distributions and processing time. It runs in an iterative, adaptive manner without apriori knowledge. Moreover, we propose a heuristic scheduling algorithm with considerations of batch processing overhead. Further, we refine the load balancing mechanism based on data processing time locality to respect job time bounds. In addition, we extend TaG to well handle traffic videos by sampling video data based on motion information encoded in the videos. We implement TaG on Hadoop and conduct extensive experiments on a large visual traffic dataset. The evaluations on different data sizes show TaG is able to achieve high accuracy within time bounds.