Video anomaly detection with spatio-temporal dissociation

Video anomaly detection with spatio-temporal dissociation
复制标题

DOI:
10.1016/j.patcog.2021.108213
复制
发表时间:
2021-08-22
影响因子:
8
通讯作者:
Yuan, Junsong
Yuan, Junsong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chang, Yunpeng;Tu, Zhigang;Yuan, Junsong

文献摘要

被引文献

相似文献

由于异常定义的模糊性和来自真实的视频数据的视觉场景的复杂性,视频中的异常检测仍然是一个具有挑战性的任务。不同于以前的工作,利用重建或预测作为辅助任务来学习的时间规律性,在这项工作中,我们探索了一种新的卷积自动编码器架构,可以分离的时空表示分别捕获的空间和时间的信息,因为异常事件通常是从外观和/或运动行为的正常不同。具体而言,空间自动编码器通过学习重建第一个单独帧(FIF)的输入来在外观特征空间上建模常态,而时间部分将前四个连续帧作为输入并将RGB差作为输出,以有效的方式模拟光流的运动。异常事件的出现或运动行为不规则,导致重建误差很大。为了提高快速移动离群点的检测性能,我们利用了基于方差的注意力模块,并将其插入到运动自动编码器中,以突出大的运动区域。此外,我们提出了一个深度K均值聚类策略,以迫使空间和运动编码器提取一个紧凑的表示。在一些公开数据集上的大量实验证明了该方法的有效性,达到了最先进的性能。代码在链接1上公开发布。(C)2021爱思唯尔有限公司保留所有权利。
Anomaly detection in videos remains a challenging task due to the ambiguous definition of anomaly and the complexity of visual scenes from real video data. Different from the previous work which utilizes reconstruction or prediction as an auxiliary task to learn the temporal regularity, in this work, we explore a novel convolution autoencoder architecture that can dissociate the spatio-temporal representation to separately capture the spatial and the temporal information, since abnormal events are usually different from the normality in appearance and/or motion behavior. Specifically, the spatial autoencoder models the normality on the appearance feature space by learning to reconstruct the input of the first individual frame (FIF), while the temporal part takes the first four consecutive frames as the input and the RGB difference as the output to simulate the motion of optical flow in an efficient way. The abnormal events, which are irregular in appearance or in motion behavior, lead to a large reconstruction error. To improve detection performance on fast moving outliers, we exploit a variance-based attention module and insert it into the motion autoencoder to highlight large movement areas. In addition, we propose a deep K means cluster strategy to force the spatial and the motion encoder to extract a compact representation. Extensive experiments on some publicly available datasets have demonstrated the effectiveness of our method which achieves the state-of-the-art performance. The code is publicly released at the link1. (C) 2021 Elsevier Ltd. All rights reserved.