Towards Interpretable Video Anomaly Detection

Towards Interpretable Video Anomaly Detection
复制标题

DOI:
10.1109/wacv56688.2023.00268
复制
发表时间:
2023-01
期刊:
2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Keval Doshi;Yasin Yılmaz
Keval Doshi;Yasin Yılmaz
中科院分区:
其他
文献类型:
--
作者:
Keval Doshi;Yasin Yılmaz

文献摘要

相似文献

大多数视频异常检测方法都是基于数据密集型的端到端训练神经网络,从视频中提取时空特征。这些方法提取的特征表示是不可解释的,这阻碍了异常原因的自动识别。为此,我们提出了一个新的框架来解释监控视频中检测到的异常事件。除了单独监视对象之外,我们还监视它们之间的交互,以检测异常事件并解释其根本原因。具体来说,我们证明了通过监测物体相互作用获得的场景图为异常的背景提供了解释,同时与最新的最先进的方法相比具有竞争力。此外,所提出的可解释方法实现了跨域适应性(即在另一个监控场景中的迁移学习),这对于大多数现有的端到端方法来说是不可行的,因为缺乏足够的每个监控场景的标记训练数据。通过理论(通过渐近最优性证明)和经验在流行的基准数据集上评估了所提出方法的快速可靠的检测性能。
Most video anomaly detection approaches are based on data-intensive end-to-end trained neural networks, which extract spatiotemporal features from videos. The extracted feature representations in such approaches are not interpretable, which prevents the automatic identification of anomaly cause. To this end, we propose a novel framework which can explain the detected anomalous event in a surveillance video. In addition to monitoring objects independently, we also monitor the interactions between them to detect anomalous events and explain their root causes. Specifically, we demonstrate that the scene graphs obtained by monitoring the object interactions provide an interpretation for the context of the anomaly while performing competitively with respect to the recent state-of-the-art approaches. Moreover, the proposed interpretable method enables cross-domain adaptability (i.e., transfer learning in another surveillance scene), which is not feasible for most existing end-to-end methods due to the lack of sufficient labeled training data for every surveillance scene. The quick and reliable detection performance of the proposed method is evaluated both theoretically (through an asymptotic optimality proof) and empirically on the popular benchmark datasets.