Learning Joint Representations of Videos and Sentences with Web Image Search

Learning Joint Representations of Videos and Sentences with Web Image Search
复制标题

DOI:
10.1007/978-3-319-46604-0_46
复制
发表时间:
2016-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Mayu Otani;Yuta Nakashima;Esa Rahtu;J. Heikkilä;N. Yokoya
Mayu Otani;Yuta Nakashima;Esa Rahtu;J. Heikkilä;N. Yokoya
中科院分区:
其他
文献类型:
--
作者:
Mayu Otani;Yuta Nakashima;Esa Rahtu;J. Heikkilä;N. Yokoya

文献摘要

被引文献

相似文献

我们的目标是基于自然语言查询的视频检索。此外,我们还考虑了给定输入视频检索句子或生成描述的类似问题。最近的工作通过将视觉和文本输入嵌入到语义相似度与距离相关的公共空间中来解决这个问题。我们还采用了嵌入方法,并做出了以下贡献:首先,我们在句子嵌入过程中利用网络图像搜索来消除细粒度视觉概念的歧义。其次,我们提出了同时学习参数的句子、图像和视频输入的嵌入模型。最后,我们展示了如何将所提出的模型应用于描述生成。总的来说,我们观察到在视频和句子检索任务中,比最先进的方法有明显的改进。在描述生成方面,虽然我们的嵌入是为检索任务训练的,但性能水平与当前的最先进水平相当。
Our objective is video retrieval based on natural language queries. In addition, we consider the analogous problem of retrieving sentences or generating descriptions given an input video. Recent work has addressed the problem by embedding visual and textual inputs into a common space where semantic similarities correlate to distances. We also adopt the embedding approach, and make the following contributions: First, we utilize web image search in sentence embedding process to disambiguate fine-grained visual concepts. Second, we propose embedding models for sentence, image, and video inputs whose parameters are learned simultaneously. Finally, we show how the proposed model can be applied to description generation. Overall, we observe a clear improvement over the state-of-the-art methods in the video and sentence retrieval tasks. In description generation, the performance level is comparable to the current state-of-the-art, although our embeddings were trained for the retrieval tasks.