Learning a Multi-concept Video Retrieval Model with Multiple Latent Variables

Learning a Multi-concept Video Retrieval Model with Multiple Latent Variables
复制标题

DOI:
10.1145/3176647
复制
发表时间:
2016-12
期刊:
2016 IEEE International Symposium on Multimedia (ISM)
影响因子:
--
通讯作者:
Amir Mazaheri;Boqing Gong;M. Shah
Amir Mazaheri;Boqing Gong;M. Shah
中科院分区:
其他
文献类型:
--
作者:
Amir Mazaheri;Boqing Gong;M. Shah

文献摘要

被引文献

相似文献

在“大视频”时代,高效的视频检索已经成为一个迫切的需求,而如何处理多概念查询是其中的核心内容。这项工作的目的是提供一个原则性的模型,用于计算响应于多个概念的视频的排名分数。然而,它一直被忽视,并简单地实现了加权平均相应的概念检测器的分数。我们的方法可以被视为潜在排名支持向量机,它集成了最近关于文本和图像检索的各种工作的优点,例如选择排名而不是结构化预测,以及对查询概念与其他概念之间的相互依赖性进行建模。视频由镜头组成,我们使用潜变量来解释镜头内和镜头间的互补线索。我们介绍了一种简单有效的方法,使我们的模型鲁棒离群值和稀缺的数据。我们的方法产生了上级性能时,它不仅在训练中看到的查询,但也新的查询,其中一些包括更多的概念比看到的查询用于训练。
Effective and efficient video retrieval has become a pressing need in the "big video" era and how to deal with multi-concept queries is a central component. The objective of this work is to provide a principled model for calculating the ranking scores of video in response to multiple concepts. However, it has been long overlooked and simply implemented by weighted averaging the corresponding concept detectors' scores. Our approach, which can be considered as a latent ranking SVM, integrates the advantages of various recent works on text and image retrieval, such as choosing ranking over structured prediction and modeling inter-dependencies between querying concepts and the others. Videos consist of shots and we use latent variables to account for the mutually complementary cues within and across shots. We introduce a simple and effective way to make our model robust to outliers and scarce data. Our approach gives rise to superior performance when it is tested on not only the queries seen at training, but also novel queries, some of which consist of more concepts than the seen queries used for training.