Discriminative Unsupervised Alignment of Natural Language Instructions with Corresponding Video Segments

Discriminative Unsupervised Alignment of Natural Language Instructions with Corresponding Video Segments
复制标题

自然语言指令与相应视频片段的有区别的无监督对齐

DOI:
--
复制
发表时间:
2015
期刊:
North American Chapter of the Association for Computational Linguistics
影响因子:
--
通讯作者:
D. Gildea
D. Gildea
中科院分区:
--
文献类型:
--
作者:
Iftekhar Naim;Y. Song;Qiguang Liu;Liang Huang;Henry A. Kautz;Jiebo Luo;D. Gildea

文献摘要

被引文献

相似文献

我们解决了在没有任何直接监督的情况下自动将自然语言句子与相应的视频片段对齐的问题。大多数现有的将语言和视频结合在一起的算法依赖于手动对齐的并行数据,其中每个自然语言句子都与其对应的图像或视频片段手动对齐。最近,使用分层生成模型,文本与视频的完全无监督对齐已被证明是可行的。与以往的产生式模型相比,我们提出了三种用于无监督比对任务的潜变量判别模型。通过添加不同的和重叠的特征,所提出的判别模型能够融合领域知识。结果表明,在对齐精度方面,判别模型优于生成模型。
We address the problem of automatically aligning natural language sentences with corresponding video segments without any direct supervision. Most existing algorithms for integrating language with videos rely on handaligned parallel data, where each natural language sentence is manually aligned with its corresponding image or video segment. Recently, fully unsupervised alignment of text with video has been shown to be feasible using hierarchical generative models. In contrast to the previous generative models, we propose three latent-variable discriminative models for the unsupervised alignment task. The proposed discriminative models are capable of incorporating domain knowledge, by adding diverse and overlapping features. The results show that discriminative models outperform the generative models in terms of alignment accuracy.