CLIP-TSA: Clip-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection

CLIP-TSA: Clip-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection
复制标题

DOI:
10.1109/icip49359.2023.10222289
复制
发表时间:
2022-12
期刊:
2023 IEEE International Conference on Image Processing (ICIP)
影响因子:
--
通讯作者:
Hyekang Joo;Khoa T. Vo;Kashu Yamazaki;Ngan T. H. Le
Hyekang Joo;Khoa T. Vo;Kashu Yamazaki;Ngan T. H. Le
中科院分区:
其他
文献类型:
--
作者:
Hyekang Joo;Khoa T. Vo;Kashu Yamazaki;Ngan T. H. Le

文献摘要

被引文献

相似文献

视频异常检测(VAD)--由于其劳动密集性,通常被描述为弱监督方式下的多示例学习问题--是视频监控中的一个具有挑战性的问题,其中需要在未裁剪的视频中定位异常帧。在本文中,我们首先提出了利用VIT编码的视觉特征,而不是传统的C3D或I3D特征,来有效地提取区分表示。然后,我们对时间依赖进行建模,并通过利用我们提出的时间自我关注(TSA)来提名感兴趣的片段。消融研究证实了TSA和VIT特征的有效性。大量的实验表明,我们提出的CLIP-TSA方法在VAD问题中常用的三个基准数据集(UCF-犯罪、上海科技校园和XD-暴力)上的性能明显优于现有的最新方法(SOTA)。我们的源代码可以在https://github.com/joos2010kj/CLIP-TSA.上找到
Video anomaly detection (VAD) – commonly formulated as a multiple-instance learning problem in a weakly-supervised manner due to its labor-intensive nature – is a challenging problem in video surveillance where the frames of anomaly need to be localized in an untrimmed video. In this paper, we first propose to utilize the ViT-encoded visual features from CLIP, in contrast with the conventional C3D or I3D features in the domain, to efficiently extract discriminative representations in the novel technique. We then model temporal dependencies and nominate the snippets of interest by leveraging our proposed Temporal Self-Attention (TSA). The ablation study confirms the effectiveness of TSA and ViT feature. The extensive experiments show that our proposed CLIP-TSA outperforms the existing state-of-the-art (SOTA) methods by a large margin on three commonly-used benchmark datasets in the VAD problem (UCF-Crime, ShanghaiTech Campus and XD-Violence). Our source code is available at https://github.com/joos2010kj/CLIP-TSA.