Human Attention Based Movie Summarization: Dataset and Baseline Model

Human Attention Based Movie Summarization: Dataset and Baseline Model
复制标题

DOI:
10.1109/icme52920.2022.9859796
复制
发表时间:
2022-07
期刊:
2022 IEEE International Conference on Multimedia and Expo (ICME)
影响因子:
--
通讯作者:
Defang Zhao;Dandan Zhu;Xiongkuo Min;Jiaomin Yue;Kaiwei Zhang;Qiangqiang Zhou;Guangtao Zhai;Xiaokang Yang
Defang Zhao;Dandan Zhu;Xiongkuo Min;Jiaomin Yue;Kaiwei Zhang;Qiangqiang Zhou;Guangtao Zhai;Xiaokang Yang
中科院分区:
其他
文献类型:
--
作者:
Defang Zhao;Dandan Zhu;Xiongkuo Min;Jiaomin Yue;Kaiwei Zhang;Qiangqiang Zhou;Guangtao Zhai;Xiaokang Yang

文献摘要

相似文献

电影摘要模型可以通过选择关键帧来自动编辑电影的精简版本。以前的作品主要依靠手工制作,大多数都是无人监督的。有监督的电影摘要是一个新的研究领域,目前还没有公开的合适的数据集。此外,现有的作品只关注电影本身,而忽视了观众,他们最有发言权的电影的哪一部分更有吸引力。为了解决上述问题,我们建立了一个基于人类注意力的电影摘要数据集Movie 50。具体而言,我们探索了人类在观看视频时的注意力变化,并有以下发现:(1)人类的注意力在观看关键帧时集中。(2)人类的注意力在观看非关键帧时会分散。受这些发现的启发,我们收集了20名参与者在观看50部电影时的眼睛注视,并提出了一种新的基于人类注意力的注释管道。此外,我们介绍了A/V-MSNet,视听神经网络,利用时空视觉和听觉信息,以更好地模拟人类的注意力,以及利用更丰富的信息。大量的实验证明了该方法的优越性。
The movie summarization model can automatically edit a condensed and succinct version of the movie by selecting the keyframes. Previous works mainly resort to hand-crafted heuristics and most of them are unsupervised. Supervised movie summarization is a new research field and, there is currently no publicly suitable dataset available. Moreover, existing works only focus on the movies themselves while neglecting the audiences, who have the most say in which part of the movie is more attractive. To deal with the aforementioned limitations, we establish a human attention based movie summarization dataset Movie50. Specifically, we explore the human attention variations when watching videos and have the following findings: (1) The attention of humans is concentrated when watching keyframes. (2) The attention of humans is distracted when watching non-keyframes. Inspired by these findings, we collect the eye fixations of 20 participants when watching 50 movies and propose a novel human attention based annotation pipeline. In addition, we introduce A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better model human attention as well as exploit more plentiful information. Extensive experiments demonstrate the superiority of the proposed method.