Emotion Knowledge Driven Video Highlight Detection

Emotion Knowledge Driven Video Highlight Detection
复制标题

情感知识驱动的视频亮点检测

DOI:
10.1109/tmm.2020.3035285
复制
发表时间:
2021
影响因子:
7.3
通讯作者:
Changsheng Xu
Changsheng Xu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Fan Qi;Xiaoshan Yang;Changsheng Xu

文献摘要

被引文献

相似文献

本文研究的是视频亮点检测,其目的是根据用户的主要兴趣或特殊兴趣选择较小的帧子集。传统方法的性能很大程度上依赖于大规模的人工标注训练数据,收集起来费时费力。为了解决这个问题,我们追溯到最初的问题定义,发现用户是否对特定的视频片段感兴趣在很大程度上取决于人的主观情感。利用这一洞察力,我们提出了一种情感知识驱动的视频检测框架,用于对人类的一般情感进行建模和推断重点强度。首先,通过前端网络获取视频片段的概念级表示。以概念为节点构建情感相关知识图,并通过外部公共知识图对图中概念间的关系进行建模。然后采用暹罗GCNS对图中节点间的依赖关系进行建模,并沿边传播消息。最后,我们计算了基于GCN层的视频片段的情感感知表示,并进一步利用它来预测精彩程度。我们的框架,包括前端网络、图卷积层和高亮映射网络,可以端到端的方式进行训练,并具有排序损失的约束。在两个基准数据集上的实验表明,我们提出的方法比最先进的方法具有更好的性能。
This paper addresses video highlight detection which aims to select a small subset of frames according to user's major or special interest. The performances of conventional methods highly depend on large-scale manually labeled training data which are time-consuming and labor-intensive to collect. To deal with this problem, we trace back to the original problem definition and find that whether a user is interested in a specific video segment heavily depends on human's subjective emotions. Leveraging this insight, we introduce an emotion knowledge driven video detection framework for modeling human's general emotion and inferencing highlight strength. Firstly, we obtain the concept-level representation of the video clip with a front-end network. The concepts are used as nodes to build an emotion-related knowledge graph, and their relationships in the graph are modeled via external public knowledge graphs. Then we adopt Siamese GCNs to model the dependencies between nodes in the graph and propagate messages along the edges. Finally, we compute the emotion-aware representation of the video clip based on the GCN layers and further use it to predict the highlight score. Our framework, including the front-end network, graph convolution layers and the highlight mapping network, can be trained in an end-to-end manner with the constraint of a ranking loss. Experiments on two benchmark datasets show that our proposed method performs favorably against the state-of-the-art methods.