Discriminative Unsupervised Alignment of Natural Language Instructions with Corresponding Video Segments
Discriminative Unsupervised Alignment of Natural Language Instructions with Corresponding Video Segments
复制标题
自然语言指令与相应视频片段的有区别的无监督对齐
DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
D. Gildea
中科院分区:
文献类型:
--
作者:
Iftekhar Naim;Y. Song;Qiguang Liu;Liang Huang;Henry A. Kautz;Jiebo Luo;D. Gildea
We address the problem of automatically aligning natural language sentences with corresponding video segments without any direct supervision. Most existing algorithms for integrating language with videos rely on handaligned parallel data, where each natural language sentence is manually aligned with its corresponding image or video segment. Recently, fully unsupervised alignment of text with video has been shown to be feasible using hierarchical generative models. In contrast to the previous generative models, we propose three latent-variable discriminative models for the unsupervised alignment task. The proposed discriminative models are capable of incorporating domain knowledge, by adding diverse and overlapping features. The results show that discriminative models outperform the generative models in terms of alignment accuracy.