Text2Video: An End-to-end Learning Framework for Expressing Text With Videos

Text2Video: An End-to-end Learning Framework for Expressing Text With Videos
复制标题

DOI:
10.1109/tmm.2018.2807588
复制
发表时间:
2018-02
影响因子:
7.3
通讯作者:
Xiaoshan Yang;Tianzhu Zhang;Changsheng Xu
Xiaoshan Yang;Tianzhu Zhang;Changsheng Xu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Xiaoshan Yang;Tianzhu Zhang;Changsheng Xu

文献摘要

被引文献

相似文献

视频创建是一项具有挑战性和高度专业性的任务,通常涉及大量的手动工作。为了减轻这种负担,更好的方法是根据任意文本从大量现有视频中自动生成新视频。在本文中,我们制定视频创建为一个问题,检索一个句子流的视频序列。为了实现这一目标,我们提出了一种新的多模态经常性架构的自动视频制作。与现有的方法相比,该模型具有三大优点。首先,据我们所知,它是第一个完全集成的端到端深度学习系统,用于现实世界的生产。我们是第一批解决检索句子流的视频序列的问题。其次,通过语义一致性建模,可以有效地利用句子和视频片段之间的对应关系。第三,它可以通过要求所产生的视频应该在视觉外观方面连贯地组织来很好地建模视觉连贯性。我们已经进行了广泛的实验上的两个应用程序,包括视频检索和视频合成。在2016年大规模电影描述挑战赛中使用的两个公共数据集上获得的定性和定量结果都证明了所提出的模型与其他最先进的算法相比的有效性。
Video creation is a challenging and highly profession-al task that generally involves substantial manual efforts. To ease this burden, a better approach is to automatically produce new videos based on clips from the massive amount of existing videos according to arbitrary text. In this paper, we formulate video creation as a problem of retrieving a sequence of videos for a sentence stream. To achieve this goal, we propose a novel multimodal recurrent architecture for automatic video production. Compared with existing methods, the proposed model has three major advantages. First, it is the first completely integrated end-to-end deep learning system for real-world production to the best of our knowledge. We are among the first to address the problem of retrieving a sequence of videos for a sentence stream. Second, it can effectively exploit the correspondence between sentences and video clips through semantic consistency modeling. Third, it can model the visual coherence well by requiring that the produced videos should be organized coherently in terms of visual appearance. We have conducted extensive experiments on two applications, including video retrieval and video composition. The qualitative and quantitative results obtained on two public datasets used in the Large Scale Movie Description Challenge 2016 both demonstrate the effectiveness of the proposed model compared with other state-of-the-art algorithms.