Weakly-Supervised Action Detection Guided by Audio Narration

Weakly-Supervised Action Detection Guided by Audio Narration
复制标题

DOI:
10.1109/cvprw56347.2022.00159
复制
发表时间:
2022-05
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
Keren Ye;Adriana Kovashka
Keren Ye;Adriana Kovashka
中科院分区:
其他
文献类型:
--
作者:
Keren Ye;Adriana Kovashka

文献摘要

相似文献

视频是比图像更组织良好的视觉概念学习数据源。与仅涉及空间信息的二维图像不同,附加的时间维度桥接并同步多种模态。然而,在大多数视频检测基准中,这些附加模式并未得到充分利用。例如,EPIC Kitchens 是第一人称(以自我为中心)视觉中最大的数据集,但它仍然依赖众包信息来细化动作边界,以提供实例级动作注释。我们探索了如何消除视频检测数据中昂贵的注释,从而提供细化的边界。我们提出了一个从叙述监督中学习并利用多模态特征的模型,包括 RGB、运动流和环境声音。我们的模型学习关注与叙述标签相关的帧,同时抑制不相关的帧的使用。我们的实验表明,嘈杂的音频叙述足以学习良好的动作检测模型,从而减少注释费用。
Videos are more well-organized curated data sources for visual concept learning than images. Unlike the 2-dimensional images which only involve the spatial information, the additional temporal dimension bridges and synchronizes multiple modalities. However, in most video detection benchmarks, these additional modalities are not fully utilized. For example, EPIC Kitchens is the largest dataset in first-person (egocentric) vision, yet it still relies on crowdsourced information to refine the action boundaries to provide instance-level action annotations.We explored how to eliminate the expensive annotations in video detection data which provide refined boundaries. We propose a model to learn from the narration supervision and utilize multimodal features, including RGB, motion flow, and ambient sound. Our model learns to attend to the frames related to the narration label while suppressing the irrelevant frames from being used. Our experiments show that noisy audio narration suffices to learn a good action detection model, thus reducing annotation expenses.