Closing-the-Loop: A Data-Driven Framework for Effective Video Summarization

Closing-the-Loop: A Data-Driven Framework for Effective Video Summarization
复制标题

DOI:
10.1109/ism.2020.00042
复制
发表时间:
2020-12
期刊:
2020 IEEE International Symposium on Multimedia (ISM)
影响因子:
--
通讯作者:
Ran Xu;Haoliang Wang;Stefano Petrangeli;Viswanathan Swaminathan;S. Bagchi
Ran Xu;Haoliang Wang;Stefano Petrangeli;Viswanathan Swaminathan;S. Bagchi
中科院分区:
其他
文献类型:
--
作者:
Ran Xu;Haoliang Wang;Stefano Petrangeli;Viswanathan Swaminathan;S. Bagchi

文献摘要

相似文献

今天,视频是通过互联网共享信息的主要方式。鉴于视频共享平台的巨大普及,使视频吸引最终用户势在必行。内容创作者依靠自己的经验从原始内容开始创作引人入胜的短视频。过去已经提出了几种方法来帮助创作者进行摘要处理。然而,很难量化这些编辑对最终用户参与度的影响。此外,视频观看量数据的可用性已经开启了在视频发布之前预测其有效性的可能性。在本文中,我们提出了一种新的框架,以关闭自动视频摘要和其数据驱动的评估之间的反馈回路。我们的闭环框架由迭代重复的两个主要步骤组成。给定一个输入视频,我们首先生成一组初始视频摘要。第二,我们预测的有效性的基础上生成的变种的数据驱动模型训练用户的视频观看量数据。我们采用遗传算法来搜索可能的摘要的空间(即,向视频添加/移除镜头),其中仅允许具有最高预测性能的那些变体存活并在其位置生成新变体。我们的研究结果表明,所提出的框架可以提高所生成的摘要的有效性与最小的计算开销相比,基线解决方案- 28.3%以上的视频摘要是在最高的有效性类比那些在基线。
Today, videos are the primary way in which information is shared over the Internet. Given the huge popularity of video sharing platforms, it is imperative to make videos engaging for the end-users. Content creators rely on their own experience to create engaging short videos starting from the raw content. Several approaches have been proposed in the past to assist creators in the summarization process. However, it is hard to quantify the effect of these edits on the end-user engagement. Moreover, the availability of video consumption data has opened the possibility to predict the effectiveness of a video before it is published. In this paper, we propose a novel framework to close the feedback loop between automatic video summarization and its data-driven evaluation. Our Closing-The-Loop framework is composed of two main steps that are repeated iteratively. Given an input video, we first generate a set of initial video summaries. Second, we predict the effectiveness of the generated variants based on a data-driven model trained on users' video consumption data. We employ a genetic algorithm to search the space of possible summaries (i.e., adding/removing shots to the video) in an efficient way, where only those variants with the highest predicted performance are allowed to survive and generate new variants in their place. Our results show that the proposed framework can improve the effectiveness of the generated summaries with minimal computation overhead compared to a baseline solution - 28.3% more video summaries are in the highest effectiveness class than those in the baseline.