Action Completeness Modeling with Background Aware Networks for Weakly-Supervised Temporal Action Localization

Action Completeness Modeling with Background Aware Networks for Weakly-Supervised Temporal Action Localization
复制标题

DOI:
10.1145/3394171.3413687
复制
发表时间:
2020-10
期刊:
Proceedings of the 28th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
M. Moniruzzaman;Zhaozheng Yin;Zhihai He;Ruwen Qin;M. Leu
M. Moniruzzaman;Zhaozheng Yin;Zhihai He;Ruwen Qin;M. Leu
中科院分区:
其他
文献类型:
--
作者:
M. Moniruzzaman;Zhaozheng Yin;Zhihai He;Ruwen Qin;M. Leu

文献摘要

被引文献

相似文献

最先进的完全监督的方法,从未修剪的视频时间动作定位取得了令人印象深刻的结果。然而,它仍然不令人满意的弱监督的时间动作定位,其中只有视频级的动作标签,没有时间戳注释的动作发生时。其主要原因是,弱监督网络只关注高区分度的帧,而背景和动作类中都存在一些模糊帧。背景类中的模糊帧与真实的动作非常相似,这可能被视为目标动作并导致误报。另一方面,动作类中可能包含动作实例的模糊帧容易被弱监督网络误判,导致定位粗糙。为了解决这些问题,我们介绍了一种新的弱监督的动作完整性建模与背景感知网络(ACM-BANets)。我们的背景感知网络(BANet)包含一个权重共享的两个分支架构,具有一个动作引导的背景感知时间注意模块(B-TAM)和一个不对称的训练策略,以抑制高度歧视性和模糊的背景帧,以消除误报。我们的动作完整性建模包含多个BANets,BANets被迫发现不同但互补的动作实例,以完全本地化的动作实例在高度歧视性和模糊的动作框架。在第i次迭代中,第i个BANet发现区分性特征,然后将其从特征图中删除。部分擦除的特征图被馈送到下一次迭代的第(i+1)个BANet中,以迫使该BANet发现与第i个BANet不同的区别性特征。在两个具有挑战性的未经修剪的视频数据集,THUMOS 14和ActivityNet1.3上进行评估,我们的方法优于目前所有的弱监督时间动作定位方法。
The state-of-the-art of fully-supervised methods for temporal action localization from untrimmed videos has achieved impressive results. Yet, it remains unsatisfactory for the weakly-supervised temporal action localization, where only video-level action labels are given without the timestamp annotation on when the actions occur. The main reason comes from that, the weakly-supervised networks only focus on the highly discriminative frames, but there are some ambiguous frames in both background and action classes. The ambiguous frames in background class are very similar to the real actions, which may be treated as target actions and result in false positives. On the other hand, the ambiguous frames in action class which possibly contain action instances, are prone to be false negatives by the weakly-supervised networks and result in a coarse localization. To solve these problems, we introduce a novel weakly-supervised Action Completeness Modeling with Background Aware Networks (ACM-BANets). Our Background Aware Network (BANet) contains a weight-sharing two-branch architecture, with an action guided Background aware Temporal Attention Module (B-TAM) and an asymmetrical training strategy, to suppress both highly discriminative and ambiguous background frames to remove the false positives. Our action completeness modeling contains multiple BANets, and the BANets are forced to discover different but complementary action instances to completely localize the action instances in both highly discriminative and ambiguous action frames. In the i-th iteration, the i-th BANet discovers the discriminative features, which are then erased from the feature map. The partially-erased feature map is fed into the (i+1)-th BANet of the next iteration to force this BANet to discover discriminative features different from the i-th BANet. Evaluated on two challenging untrimmed video datasets, THUMOS14 and ActivityNet1.3, our approach outperforms all the current weakly-supervised methods for temporal action localization.