Human Attention Based Movie Summarization: Dataset and Baseline Model
Human Attention Based Movie Summarization: Dataset and Baseline Model
复制标题
DOI:
10.1109/icme52920.2022.9859796
复制
发表时间:
2022-07
期刊:
影响因子:
--
通讯作者:
Defang Zhao;Dandan Zhu;Xiongkuo Min;Jiaomin Yue;Kaiwei Zhang;Qiangqiang Zhou;Guangtao Zhai;Xiaokang Yang
中科院分区:
文献类型:
--
作者:
Defang Zhao;Dandan Zhu;Xiongkuo Min;Jiaomin Yue;Kaiwei Zhang;Qiangqiang Zhou;Guangtao Zhai;Xiaokang Yang
The movie summarization model can automatically edit a condensed and succinct version of the movie by selecting the keyframes. Previous works mainly resort to hand-crafted heuristics and most of them are unsupervised. Supervised movie summarization is a new research field and, there is currently no publicly suitable dataset available. Moreover, existing works only focus on the movies themselves while neglecting the audiences, who have the most say in which part of the movie is more attractive. To deal with the aforementioned limitations, we establish a human attention based movie summarization dataset Movie50. Specifically, we explore the human attention variations when watching videos and have the following findings: (1) The attention of humans is concentrated when watching keyframes. (2) The attention of humans is distracted when watching non-keyframes. Inspired by these findings, we collect the eye fixations of 20 participants when watching 50 movies and propose a novel human attention based annotation pipeline. In addition, we introduce A/V-MSNet, an audiovisual neural network that takes advantage of spatio-temporal visual and auditory information to better model human attention as well as exploit more plentiful information. Extensive experiments demonstrate the superiority of the proposed method.