An Entanglement-driven Fusion Neural Network for Video Sentiment Analysis

An Entanglement-driven Fusion Neural Network for Video Sentiment Analysis
复制标题

DOI:
10.24963/ijcai.2021/239
复制
发表时间:
2021-08
期刊:
Computación y Sistemas
影响因子:
--
通讯作者:
Dimitris Gkoumas;Qiuchi Li;Yijun Yu;Dawei Song
Dimitris Gkoumas;Qiuchi Li;Yijun Yu;Dawei Song
中科院分区:
其他
文献类型:
--
作者:
Dimitris Gkoumas;Qiuchi Li;Yijun Yu;Dawei Song

文献摘要

相似文献

视频数据本质上是多模态的,其中话语可以涉及语言,视觉和声学信息。因此,视频情感分析的一个关键挑战是如何有效地将联合收割机不同的模态结合起来进行情感识别。最新的神经网络方法实现了最先进的性能,但它们在很大程度上忽视了人类如何理解和推理情感状态。相比之下,量子概率神经模型的最新进展已经实现了与最先进技术相当的性能,但具有更好的透明度和更高的可解释性。然而,现有的量子启发模型将量子态视为经典混合物或跨模态的可分离张量积,而不会以它们是相关的或不可分离的方式触发它们的相互作用(即,纠缠)。这意味着目前的模型还没有充分利用量子概率的表达能力。为了填补这一空白,我们提出了一个透明的量子概率神经模型。该模型诱导不同的模态以一种可能不可分离的方式相互作用,以非经典相关性的形式编码跨模态信息。在两个视频情感分析基准数据集上的综合评价表明,该模型取得了显着的性能改善。我们还表明,模态之间的不可分离程度优化了事后可解释性。
Video data is multimodal in its nature, where an utterance can involve linguistic, visual and acoustic information. Therefore, a key challenge for video sentiment analysis is how to combine different modalities for sentiment recognition effectively. The latest neural network approaches achieve state-of-the-art performance, but they neglect to a large degree of how humans understand and reason about sentiment states. By contrast, recent advances in quantum probabilistic neural models have achieved comparable performance to the state-of-the-art, yet with better transparency and increased level of interpretability. However, the existing quantum-inspired models treat quantum states as either a classical mixture or as a separable tensor product across modalities, without triggering their interactions in a way that they are correlated or non-separable (i.e., entangled). This means that the current models have not fully exploited the expressive power of quantum probabilities. To fill this gap, we propose a transparent quantum probabilistic neural model. The model induces different modalities to interact in such a way that they may not be separable, encoding crossmodal information in the form of non-classical correlations. Comprehensive evaluation on two benchmarking datasets for video sentiment analysis shows that the model achieves significant performance improvement. We also show that the degree of non-separability between modalities optimizes the post-hoc interpretability.