Multiple Visual-Semantic Embedding for Video Retrieval from Query Sentence

Multiple Visual-Semantic Embedding for Video Retrieval from Query Sentence
复制标题

DOI:
10.3390/app11073214
复制
发表时间:
2020-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Huy Nguyen;Tomo Miyazaki;Yoshihiro Sugaya;S. Omachi
Huy Nguyen;Tomo Miyazaki;Yoshihiro Sugaya;S. Omachi
中科院分区:
其他
文献类型:
--
作者:
Huy Nguyen;Tomo Miyazaki;Yoshihiro Sugaya;S. Omachi

文献摘要

相似文献

视觉语义嵌入的目的是学习相关视频和句子实例相互靠近的联合嵌入空间。大多数现有方法将实例放在单个嵌入空间中。然而,他们很难嵌入实例,因为很难将视频中的视觉动态与句子中的文本特征相匹配。一个单独的空间不足以容纳各种视频和句子。在本文中,我们提出了一种新的框架,它将实例映射到多个独立的嵌入空间中,以便捕获实例之间的多个关系,从而实现引人注目的视频检索。我们建议通过使用加权和策略融合在每个嵌入空间中测量的相似度来产生实例之间的最终相似度。我们根据一句话来确定权重。因此,我们可以灵活地强调嵌入空间。我们在一个基准数据集上进行了句子到视频的检索实验。该方法取得了较好的性能,其结果与目前最先进的方法相比具有很强的竞争力。实验结果表明,与已有方法相比,本文提出的多重嵌入方法是有效的。
Visual-semantic embedding aims to learn a joint embedding space where related video and sentence instances are located close to each other. Most existing methods put instances in a single embedding space. However, they struggle to embed instances due to the difficulty of matching visual dynamics in videos to textual features in sentences. A single space is not enough to accommodate various videos and sentences. In this paper, we propose a novel framework that maps instances into multiple individual embedding spaces so that we can capture multiple relationships between instances, leading to compelling video retrieval. We propose to produce a final similarity between instances by fusing similarities measured in each embedding space using a weighted sum strategy. We determine the weights according to a sentence. Therefore, we can flexibly emphasize an embedding space. We conducted sentence-to-video retrieval experiments on a benchmark dataset. The proposed method achieved superior performance, and the results are competitive to state-of-the-art methods. These experimental results demonstrated the effectiveness of the proposed multiple embedding approach compared to existing methods.