Multi-query Video Retrieval

Multi-query Video Retrieval
复制标题

DOI:
10.1007/978-3-031-19781-9_14
复制
发表时间:
2022-01
期刊:
--
影响因子:
--
通讯作者:
Zeyu Wang;Yu Wu;Karthik Narasimhan;Olga Russakovsky
Zeyu Wang;Yu Wu;Karthik Narasimhan;Olga Russakovsky
中科院分区:
其他
文献类型:
--
作者:
Zeyu Wang;Yu Wu;Karthik Narasimhan;Olga Russakovsky

文献摘要

相似文献

基于文本描述的目标视频检索是一个具有重要实用价值的课题,近年来受到越来越多的关注。尽管最近取得了进展,现有的视频检索数据集的不完善的注释模型的评估和发展带来了重大挑战。在本文中,我们解决了这个问题,重点是研究较少的设置多查询视频检索,其中多个描述提供给模型搜索的视频档案。我们首先表明,多查询检索任务有效地减轻了不完善的注释所引入的数据集噪声,并更好地与人类对当前模型检索能力的评估判断相关。然后,我们研究了几种在训练时利用多个查询的方法,并证明了多查询启发的训练可以带来上级性能和更好的泛化。我们希望在这个方向上的进一步研究可以为构建在现实世界的视频检索应用程序中表现更好的系统带来新的见解(代码可以在https://github.com/princetonvisualai/MQVR上获得)。
Retrieving target videos based on text descriptions is a task of great practical value and has received increasing attention over the past few years. Despite recent progress, imperfect annotations in existing video retrieval datasets have posed significant challenges on model evaluation and development. In this paper, we tackle this issue by focusing on the less-studied setting of multi-query video retrieval, where multiple descriptions are provided to the model for searching over the video archive. We first show that multi-query retrieval task effectively mitigates the dataset noise introduced by imperfect annotations and better correlates with human judgement on evaluating retrieval abilities of current models. We then investigate several methods which leverage multiple queries at training time, and demonstrate that the multi-query inspired training can lead to superior performance and better generalization. We hope further investigation in this direction can bring new insights on building systems that perform better in real-world video retrieval applications (Code is available at https://github.com/princetonvisualai/MQVR).