Collaborative Foreground, Background, and Action Modeling Network for Weakly Supervised Temporal Action Localization

Collaborative Foreground, Background, and Action Modeling Network for Weakly Supervised Temporal Action Localization
复制标题

DOI:
10.1109/tcsvt.2023.3272891
复制
发表时间:
2023-11
影响因子:
8.4
通讯作者:
Md Moniruzzaman;Zhaozheng Yin
Md Moniruzzaman;Zhaozheng Yin
中科院分区:
工程技术1区
文献类型:
--
作者:
Md Moniruzzaman;Zhaozheng Yin

文献摘要

相似文献

在本文中,我们探讨了弱监督时间动作定位(W-TAL)问题,其任务是在只有视频级监督的情况下定位未修剪视频中所有动作实例的时间边界。现有的W-TAL方法通过分离判别动作和背景帧,实现了良好的动作定位性能。然而,弱监督方法和完全监督方法之间仍然存在很大的性能差距。其主要原因在于,除了判别性的动作和背景框架外,还存在大量的模糊动作和背景框架。由于W-TAL缺乏时间注释,模糊的背景帧可能被定位为前景,而模糊的动作帧可能被抑制为背景,分别导致假阳性和假阴性。在本文中,我们引入了一种新的协同前景、背景和动作建模网络(FBA-Net)来抑制背景(即判别性和模糊性背景)帧,并将实际动作相关(即判别性和模糊性动作)帧定位为前景,以实现精确的时间动作定位。我们设计的FBA-Net有三个分支:前景建模(FM)分支、背景建模(BM)分支和类特定的动作和背景建模(CM)分支。CM分支学习突出显示与$C$ action类相关的视频帧,并将$C$ action类的与动作相关的帧与$(C+1)$ th背景类分开。FM和CM之间的协作规范了FM和CM的$C$动作类之间的一致性,通过将视频中不同的实际动作相关(即判别性和模糊性动作)帧定位为前景来降低假阴性率。另一方面,BM和CM之间的协作使BM与CM的第$(C+1)$背景类之间的一致性得到了正则化,通过抑制判别性和模糊性背景帧来降低误报率。此外,FM和BM之间的协作加强了更有效的前景和背景分离。为了评估FBA-Net的有效性,我们在两个具有挑战性的数据集THUMOS14和ActivityNet1.3上进行了广泛的实验。实验表明,我们的FBA-Net取得了较好的效果。
In this paper, we explore the problem of Weakly-supervised Temporal Action Localization (W-TAL), where the task is to localize the temporal boundaries of all action instances in an untrimmed video with only video-level supervision. The existing W-TAL methods achieve a good action localization performance by separating the discriminative action and background frames. However, there is still a large performance gap between the weakly and fully supervised methods. The main reason comes from that there are plenty of ambiguous action and background frames in addition to the discriminative action and background frames. Due to the lack of temporal annotations in W-TAL, the ambiguous background frames may be localized as foreground and the ambiguous action frames may be suppressed as background, which result in false positives and false negatives, respectively. In this paper, we introduce a novel collaborative Foreground, Background, and Action Modeling Network (FBA-Net) to suppress the background (i.e., both the discriminative and ambiguous background) frames, and localize the actual-action-related (i.e., both the discriminative and ambiguous action) frames as foreground, for the precise temporal action localization. We design our FBA-Net with three branches: the foreground modeling (FM) branch, the background modeling (BM) branch, and the class-specific action and background modeling (CM) branch. The CM branch learns to highlight the video frames related to $C$ action classes, and separate the action-related frames of $C$ action classes from the $(C+1)$ th background class. The collaboration between FM and CM regularizes the consistency between the FM and the $C$ action classes of CM, which reduces the false negative rate by localizing different actual-action-related (i.e., both the discriminative and ambiguous action) frames in a video as foreground. On the other hand, the collaboration between BM and CM regularizes the consistency between the BM and the $(C+1)$ th background class of CM, which reduces the false positive rate by suppressing both the discriminative and ambiguous background frames. Furthermore, the collaboration between FM and BM enforces more effective foreground-background separation. To evaluate the effectiveness of our FBA-Net, we perform extensive experiments on two challenging datasets, THUMOS14 and ActivityNet1.3. The experiments show that our FBA-Net attains superior results.