An Attention-Guided Multistream Feature Fusion Network for Early Localization of Risky Traffic Agents in Driving Videos

An Attention-Guided Multistream Feature Fusion Network for Early Localization of Risky Traffic Agents in Driving Videos
复制标题

DOI:
10.1109/tiv.2023.3275543
复制
发表时间:
2024-01
影响因子:
8.2
通讯作者:
Muhammad Monjurul Karim;Zhaozheng Yin;Ruwen Qin
Muhammad Monjurul Karim;Zhaozheng Yin;Ruwen Qin
中科院分区:
工程技术2区
文献类型:
--
作者:
Muhammad Monjurul Karim;Zhaozheng Yin;Ruwen Qin

文献摘要

相似文献

在车载仪表板摄像头(仪表盘摄像头)拍摄的视频中检测危险的交通代理对于确保复杂环境中的安全导航至关重要。事故相关视频只是驾驶相关大数据中的一小部分,事故前的瞬时过程是高度动态和复杂的。此外,危险和非危险的交通代理在外观上可能是相似的。这些都使得危险的交通代理在驾驶视频中的本地化特别具有挑战性。为此,本文提出了一种注意力引导的多数据流特征融合网络(AM-Net),用于在潜在事故发生前从仪表盘摄像头视频中定位危险交通代理。两个门控递归单元(GRU)网络使用从连续视频帧中提取的对象包围盒和光流特征来捕捉时空线索,以区分危险的交通代理。注意模块与GRU相结合,学习识别与事故相关的交通代理。AM-Net融合了全局特征和对象级特征两种特征流,预测视频中交通代理的风险得分。为了支持这项研究,本文还介绍了一个新的基准数据集,称为风险对象本地化(ROL)。数据集包含事故、对象和场景级别属性的空间、时间和分类注释。提出的AM-Net在ROL数据集上取得了85.59%的AUC性能。此外,AM-Net在公共DOTA数据集上的视频异常检测性能比当前最先进的视频异常检测技术高3.5%。一项彻底的烧蚀研究通过评估其成分的影响进一步揭示了AM-Net的优点。
Detecting dangerous traffic agents in videos captured by vehicle-mounted dashboard cameras (dashcams) is essential to ensure safe navigation in complex environments. Accident-related videos are just a minor portion of the driving-related Big Data, and the transient pre-accident process is highly dynamic and complex. Besides, risky and non-risky traffic agents can be similar in their appearance. These make risky traffic agent localization in the driving video particularly challenging. To this end, this article proposes an attention-guided multistream feature fusion network (AM-Net) to localize dangerous traffic agents from dashcam videos ahead of potential accidents. Two Gated Recurrent Unit (GRU) networks use object bounding box and optical flow features extracted from consecutive video frames to capture spatio-temporal cues for distinguishing risky traffic agents. An attention module, coupled with the GRUs, learns to identify traffic agents that are relevant to an accident. Fusing the two streams of global and object-level features, AM-Net predicts the riskiness scores of traffic agents in the video. In supporting this study, the article also introduces a new benchmark dataset called Risky Object Localization (ROL). The dataset contains spatial, temporal, and categorical annotations of the accident, object, and scene-level attributes. The proposed AM-Net achieves a promising performance of 85.59% AUC on the ROL dataset. Additionally, the AM-Net outperforms the current state-of-the-art for video anomaly detection by 3.5% AUC on the public DoTA dataset. A thorough ablation study further reveals AM-Net's merits by assessing the impact of its constituents.