What Is More Likely to Happen Next? Video-and-Language Future Event Prediction

What Is More Likely to Happen Next? Video-and-Language Future Event Prediction
复制标题

DOI:
10.18653/v1/2020.emnlp-main.706
复制
发表时间:
2020-10
期刊:
--
影响因子:
--
通讯作者:
Jie Lei;Licheng Yu;Tamara L. Berg;Mohit Bansal
Jie Lei;Licheng Yu;Tamara L. Berg;Mohit Bansal
中科院分区:
其他
文献类型:
--
作者:
Jie Lei;Licheng Yu;Tamara L. Berg;Mohit Bansal

文献摘要

被引文献

相似文献

给定一个对话一致的视频,人们通常可以推断出接下来更可能发生的事情。做出这样的预测不仅需要深刻理解视频和对话背后的丰富动态,还需要大量的常识性知识。在这项工作中,我们探索人工智能模型是否能够学习做出这种多模态常识性的下一个事件预测。为了支持这一方向的研究,我们收集了一个新的数据集,名为视频和语言事件预测(VLEP),其中包括来自10,234个不同电视节目和YouTube生活方式视频片段的28,726个未来事件预测示例(以及它们的基本原理)。为了促进非平凡挑战性示例的收集,我们采用了对抗性的人与模型在循环中的数据收集过程。我们还提出了一个强有力的基线,包括来自视频、对话和常识的信息。实验表明,每种类型的信息对这一具有挑战性的任务都是有用的,与人类在VLEP上的高表现相比,我们的模型提供了一个很好的起点,但为未来的工作留下了很大的空间。我们的数据集和代码可在:这个https URL
Given a video with aligned dialogue, people can often infer what is more likely to happen next. Making such predictions requires not only a deep understanding of the rich dynamics underlying the video and dialogue, but also a significant amount of commonsense knowledge. In this work, we explore whether AI models are able to learn to make such multimodal commonsense next-event predictions. To support research in this direction, we collect a new dataset, named Video-and-Language Event Prediction (VLEP), with 28,726 future event prediction examples (along with their rationales) from 10,234 diverse TV Show and YouTube Lifestyle Vlog video clips. In order to promote the collection of non-trivial challenging examples, we employ an adversarial human-and-model-in-the-loop data collection procedure. We also present a strong baseline incorporating information from video, dialogue, and commonsense knowledge. Experiments show that each type of information is useful for this challenging task, and that compared to the high human performance on VLEP, our model provides a good starting point but leaves large room for future work. Our dataset and code are available at: this https URL