Temporal Feature Enhancement Dilated Convolution Network for Weakly-supervised Temporal Action Localization

Temporal Feature Enhancement Dilated Convolution Network for Weakly-supervised Temporal Action Localization
复制标题

DOI:
10.1109/wacv56688.2023.00597
复制
发表时间:
2023-01
期刊:
2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Jianxiong Zhou;Ying Wu
Jianxiong Zhou;Ying Wu
中科院分区:
其他
文献类型:
--
作者:
Jianxiong Zhou;Ying Wu

文献摘要

相似文献

弱监督时间动作定位(WTAL)的目的是分类和定位动作实例在未经修剪的视频只有视频级的标签。现有方法通常使用直接从预训练提取器提取的片段级RGB和光流特征。由于两个限制:由于片段的时间跨度短和初始特征不合适,这些WTAL方法缺乏对时间信息的有效利用,性能有限。在本文中,我们提出了时间特征增强扩张卷积网络(TFE-DCN)来解决这两个限制。所提出的TFE-DCN具有扩大的感受野,其覆盖了长时间跨度以观察动作实例的完整动态,这使得它能够捕获片段之间的时间依赖性。此外,我们提出了模态增强模块,可以增强RGB功能的帮助下,增强光流功能,使整体功能适合WTAL任务。在THUMOS'14和ActivityNet v1.3数据集上进行的实验表明,我们提出的方法远远优于最先进的WTAL方法。
Weakly-supervised Temporal Action Localization (WTAL) aims to classify and localize action instances in untrimmed videos with only video-level labels. Existing methods typically use snippet-level RGB and optical flow features extracted from pre-trained extractors directly. Because of two limitations: the short temporal span of snippets and the inappropriate initial features, these WTAL methods suffer from the lack of effective use of temporal information and have limited performance. In this paper, we propose the Temporal Feature Enhancement Dilated Convolution Network (TFE-DCN) to address these two limitations. The proposed TFE-DCN has an enlarged receptive field that covers a long temporal span to observe the full dynamics of action instances, which makes it powerful to capture temporal dependencies between snippets. Furthermore, we propose the Modality Enhancement Module that can enhance RGB features with the help of enhanced optical flow features, making the overall features appropriate for the WTAL task. Experiments conducted on THUMOS’14 and ActivityNet v1.3 datasets show that our proposed approach far outperforms state-of-the-art WTAL methods.