A Dynamic Spatial-Temporal Attention Network for Early Anticipation of Traffic Accidents

A Dynamic Spatial-Temporal Attention Network for Early Anticipation of Traffic Accidents
复制标题

DOI:
10.1109/tits.2022.3155613
复制
发表时间:
2022-03-09
影响因子:
8.5
通讯作者:
Yin, Zhaozheng
Yin, Zhaozheng
中科院分区:
工程技术1区
文献类型:
--
作者:
Karim, Muhammad Monjurul;Li, Yu;Yin, Zhaozheng

文献摘要

被引文献

相似文献

传感器技术和人工智能的快速发展为加强交通安全创造了新的机遇。仪表盘摄像头(dashcams)已广泛应用于人类驾驶车辆和自动驾驶车辆。一个计算智能模型可以准确、迅速地从行车记录仪的视频中预测事故,这将加强事故预防的准备工作。交通主体的时空交互作用是复杂的。预测未来事故的视觉线索深深嵌入在行车记录仪的视频数据中。因此,对交通事故的早期预测仍然是一个挑战。受人类视觉感知事故风险时的注意行为的启发,本文提出了一种基于行车记录仪视频的动态时空注意(DSTA)网络,用于事故的早期预测。dsta网络通过动态时间注意(Dynamic temporal Attention, DTA)模块学习选择视频序列的判别时间段。它还通过动态空间注意(DSA)模块学习关注帧的信息空间区域。门控循环单元(GRU)与关注模块联合训练,用于预测未来事故的概率。在两个基准数据集上对dsta -网络的评估证实,它已经超过了最先进的性能。一项全面的消融研究在组件级别评估dsta -网络,揭示了网络是如何实现这种性能的。此外,本文还提出了一种将两个互补模型的预测分数融合的方法,并验证了其在进一步提高事故早期预测性能方面的有效性。
The rapid advancement of sensor technologies and artificial intelligence are creating new opportunities for traffic safety enhancement. Dashboard cameras (dashcams) have been widely deployed on both human driving vehicles and automated driving vehicles. A computational intelligence model that can accurately and promptly predict accidents from the dashcam video will enhance the preparedness for accident prevention. The spatial-temporal interaction of traffic agents is complex. Visual cues for predicting a future accident are embedded deeply in dashcam video data. Therefore, the early anticipation of traffic accidents remains a challenge. Inspired by the attention behavior of humans in visually perceiving accident risks, this paper proposes a Dynamic Spatial-Temporal Attention (DSTA) network for the early accident anticipation from dashcam videos. The DSTA-network learns to select discriminative temporal segments of a video sequence with a Dynamic Temporal Attention (DTA) module. It also learns to focus on the informative spatial regions of frames with a Dynamic Spatial Attention (DSA) module. A Gated Recurrent Unit (GRU) is trained jointly with the attention modules to predict the probability of a future accident. The evaluation of the DSTA-network on two benchmark datasets confirms that it has exceeded the state-of-the-art performance. A thorough ablation study that assesses the DSTA-network at the component level reveals how the network achieves such performance. Furthermore, this paper proposes a method to fuse the prediction scores from two complementary models and verifies its effectiveness in further boosting the performance of early accident anticipation.