ERA: Entity Relationship Aware Video Summarization with Wasserstein GAN

ERA: Entity Relationship Aware Video Summarization with Wasserstein GAN
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Guande Wu;Jianzhe Lin;Cláudio T. Silva
Guande Wu;Jianzhe Lin;Cláudio T. Silva
中科院分区:
其他
文献类型:
--
作者:
Guande Wu;Jianzhe Lin;Cláudio T. Silva

文献摘要

相似文献

视频摘要旨在通过生成简洁的简短摘要来简化大规模视频浏览,这些简洁摘要从中偏离但很好地代表了原始视频。由于视频注释的稀缺性,视频摘要的最新进展集中在无监督的方法上,其中基于GAN的方法最为普遍。这种类型的方法包括摘要和歧视器。摘要中的摘要视频将被假定为最终输出,只有当从本摘要中重建的视频不能由歧视者歧视原始摘要。这种基于GAN的方法的主要问题是两个折叠。首先,以这种方式进行的摘要视频是原始视频的子集,其冗余性低,并包含高优先事件/实体。此摘要标准还不够。其次,GAN框架的培训不稳定。本文提出了一种新颖的实体关系意识到视频摘要方法(ERA),以解决上述问题。更具体地说,我们引入了一个对抗性时空时间网络来构建实体之间的关系,我们认为在摘要中也应高度优先。通过引入Wasserstein Gan和两个新提出的视频补丁/分数损失来解决GAN训练问题。此外,分数损失还可以减轻对不同视频长度的模型敏感性,这对于大多数当前的视频分析任务来说是一个固有的问题。我们的方法显着提高了目标基准数据集上的性能,并超过了当前排行榜1的最新状态CSNET状态(TVSUM的F1得分增加了2.1%,Summe的F1得分提高了3.1%)。我们希望我们直接而有效的方法能够阐明未来无监督视频摘要的研究。
Video summarization aims to simplify large scale video browsing by generating concise, short summaries that diver from but well represent the original video. Due to the scarcity of video annotations, recent progress for video summarization concentrates on unsupervised methods, among which the GAN based methods are most prevalent. This type of methods includes a summarizer and a discriminator. The summarized video from the summarizer will be assumed as the final output, only if the video reconstructed from this summary cannot be discriminated from the original one by the discriminator. The primary problems of this GAN based methods are two folds. First, the summarized video in this way is a subset of original video with low redundancy and contains high priority events/entities. This summarization criterion is not enough. Second, the training of the GAN framework is not stable. This paper proposes a novel Entity relationship Aware video summarization method (ERA) to address the above problems. To be more specific, we introduce an Adversarial Spatio Temporal network to construct the relationship among entities, which we think should also be given high priority in the summarization. The GAN training problem is solved by introducing the Wasserstein GAN and two newly proposed video patch/score sum losses. In addition, the score sum loss can also relieve the model sensitivity to the varying video lengths, which is an inherent problem for most current video analysis tasks. Our method substantially lifts the performance on the target benchmark datasets and exceeds the current leaderboard Rank 1 state of the art CSNet (2.1% F1 score increase on TVSum and 3.1% F1 score increase on SumMe). We hope our straightforward yet effective approach will shed some light on the future research of unsupervised video summarization.