Video-aided Unsupervised Grammar Induction

Video-aided Unsupervised Grammar Induction
复制标题

DOI:
10.18653/v1/2021.naacl-main.119
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Songyang Zhang;Linfeng Song;Lifeng Jin;Kun Xu;Dong Yu;Jiebo Luo
Songyang Zhang;Linfeng Song;Lifeng Jin;Kun Xu;Dong Yu;Jiebo Luo
中科院分区:
其他
文献类型:
--
作者:
Songyang Zhang;Linfeng Song;Lifeng Jin;Kun Xu;Dong Yu;Jiebo Luo

文献摘要

被引文献

相似文献

我们研究了视频辅助语法归纳,它从未标记的文本和相应的视频中学习一个选区解析器。现有的多模态语法归纳方法主要集中在文本-图像对的语法归纳上,结果表明静态图像信息在归纳中是有用的。然而,视频提供了更丰富的信息,不仅包括静态对象,还包括有助于诱导动词短语的动作和状态变化。在本文中,我们以最近的复合PCFG模型为基线,从视频中探索丰富的特征(如动作、对象、场景、音频、人脸、OCR和语音)。我们进一步提出了一个多模态复合PCFG模型(MMC-PCFG)来有效地聚合这些来自不同模态的丰富特征。我们提出的MMC-PCFG是端到端训练的,在三个基准(即DiDeMo, YouCook2和MSRVTT)上优于每个单独的模态和以前最先进的系统,证实了利用视频信息进行无监督语法归纳的有效性。
We investigate video-aided grammar induction, which learns a constituency parser from both unlabeled text and its corresponding video. Existing methods of multi-modal grammar induction focus on grammar induction from text-image pairs, with promising results showing that the information from static images is useful in induction. However, videos provide even richer information, including not only static objects but also actions and state changes useful for inducing verb phrases. In this paper, we explore rich features (e.g. action, object, scene, audio, face, OCR and speech) from videos, taking the recent Compound PCFG model as the baseline. We further propose a Multi-Modal Compound PCFG model (MMC-PCFG) to effectively aggregate these rich features from different modalities. Our proposed MMC-PCFG is trained end-to-end and outperforms each individual modality and previous state-of-the-art systems on three benchmarks, i.e. DiDeMo, YouCook2 and MSRVTT, confirming the effectiveness of leveraging video information for unsupervised grammar induction.