Efficient multi-modal retrieval in conceptual space

Efficient multi-modal retrieval in conceptual space
复制标题

DOI:
10.1145/2072298.2071944
复制
发表时间:
2011-11
期刊:
Proceedings of the 19th ACM international conference on Multimedia
影响因子:
--
通讯作者:
Jun Imura;Teppei Fujisawa;T. Harada;Y. Kuniyoshi
Jun Imura;Teppei Fujisawa;T. Harada;Y. Kuniyoshi
中科院分区:
其他
文献类型:
--
作者:
Jun Imura;Teppei Fujisawa;T. Harada;Y. Kuniyoshi

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新的、高效的检索系统,适用于包括视频轨道在内的大规模多模态数据。大规模多模态数据,数据量巨大、内容多样,导致检索结果的效率和精度下降。最近关于图像注释和检索的研究表明,基于视觉词袋方法和 SIFT 等局部描述符的图像特征在处理大规模图像数据集时表现得令人惊讶。这些强大的描述符往往是高维的,这给原始特征空间中的近似最近邻搜索带来了很高的计算成本。我们的视频检索方法侧重于同时记录的图像、声音和位置信息之间的相关性,并学习描述数据内容的概念空间以实现高效搜索。实验表明我们的检索系统具有良好的性能,内存使用率低,时间复杂度低。
In this paper, we propose a new, efficient retrieval system for large-scale multi-modal data including video tracks. With large-scale multi-modal data, the huge data size and various contents cause degradation of efficiency and precision of retrieval results. Recent research on image annotation and retrieval shows that image features based on the Bag-of-Visual Words approach with local descriptors such as SIFT perform surprisingly well with large-scale image datasets. Those powerful descriptors tend to be high-dimensional, imposing a high computational cost for approximate nearest neighbor searching in raw feature space. Our video retrieval method is focused on the correlation between image, sound, and location information recorded simultaneously, and to learn conceptual space describing the contents of the data to realize efficient searching. Experiments show good performance of our retrieval system with low memory usage and temporal complexity.