Long-Term Feature Banks for Detailed Video Understanding

Long-Term Feature Banks for Detailed Video Understanding
复制标题

DOI:
10.1109/cvpr.2019.00037
复制
发表时间:
2018-12
期刊:
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Chao-Yuan Wu;Christoph Feichtenhofer;Haoqi Fan;Kaiming He;Philipp Krähenbühl;Ross B. Girshick
Chao-Yuan Wu;Christoph Feichtenhofer;Haoqi Fan;Kaiming He;Philipp Krähenbühl;Ross B. Girshick
中科院分区:
其他
文献类型:
--
作者:
Chao-Yuan Wu;Christoph Feichtenhofer;Haoqi Fan;Kaiming He;Philipp Krähenbühl;Ross B. Girshick

文献摘要

被引文献

相似文献

为了理解世界,我们人类需要不断地将现在与过去联系起来,并将事件放在背景中。在本文中,我们使现有的视频模型也能做到这一点。我们提出了一个长期特征库--在整个视频范围内提取支持信息--以增强最先进的视频模型,否则只能观看2-5秒的短片段。我们的实验表明,使用长期特征库来增强3D卷积网络可以在三个具有挑战性的视频数据集上产生最先进的结果:AVA、EPIC-Kitchens和Charade。代码可以在网上获得。
To understand the world, we humans constantly need to relate the present to the past, and put events in context. In this paper, we enable existing video models to do the same. We propose a long-term feature bank—supportive information extracted over the entire span of a video—to augment state-of-the-art video models that otherwise would only view short clips of 2-5 seconds. Our experiments demonstrate that augmenting 3D convolutional networks with a long-term feature bank yields state-of-the-art results on three challenging video datasets: AVA, EPIC-Kitchens, and Charades. Code is available online.