Attention-Based Modality-Gated Networks for Image-Text Sentiment Analysis

Attention-Based Modality-Gated Networks for Image-Text Sentiment Analysis
复制标题

用于图像文本情感分析的基于注意力的模态门控网络

DOI:
10.1145/3388861
复制
发表时间:
2020-09-01
影响因子:
5.1
通讯作者:
Li, Zhoujun
Li, Zhoujun
中科院分区:
计算机科学3区
文献类型:
--
作者:
Huang, Feiran;Wei, Kaimin;Li, Zhoujun

文献摘要

被引文献

相似文献

社会多媒体数据的情感分析已经引起了广泛的研究兴趣,并已被应用于许多任务,如选举预测和产品评估。一种模态的情感分析(例如,文本或图像)已经被广泛地研究。然而,多模态数据的情感分析并没有受到太多的关注。不同的模式通常具有互补的信息。因此,有必要通过将视觉内容与文本描述相结合来学习整体情感。在这篇文章中,我们提出了一种新的方法-基于注意力的模态门控网络(AMGN)-利用图像和文本的模态之间的相关性,并提取多模态情感分析的判别特征。具体而言,视觉语义注意力模型,提出了学习每个单词的关注视觉特征。为了有效地将图像和文本两种模态上的情感信息联合收割机结合起来,提出了一种模态门控LSTM,通过自适应地选择呈现更强情感信息的模态来学习多模态特征。然后提出了一种语义自注意模型,用于自动关注情感分类的判别特征。在人工标注和机器弱标记数据集上进行了大量的实验。结果表明,我们的方法的优越性,通过比较与国家的最先进的模型。
Sentiment analysis of social multimedia data has attracted extensive research interest and has been applied to many tasks, such as election prediction and products evaluation. Sentiment analysis of one modality (e.g., text or image) has been broadly studied. However, not much attention has been paid to the sentiment analysis of multimodal data. Different modalities usually have information that is complementary. Thus, it is necessary to learn the overall sentiment by combining the visual content with text description. In this article, we propose a novel method-Attention-Based Modality-Gated Networks (AMGN)-to exploit the correlation between the modalities of images and texts and extract the discriminative features for multimodal sentiment analysis. Specifically, a visual-semantic attention model is proposed to learn attended visual features for each word. To effectively combine the sentiment information on the two modalities of image and text, a modalitygated LSTM is proposed to learn the multimodal features by adaptively selecting the modality that presents stronger sentiment information. Then a semantic self-attention model is proposed to automatically focus on the discriminative features for sentiment classification. Extensive experiments have been conducted on both manually annotated and machine weakly labeled datasets. The results demonstrate the superiority of our approach through comparison with state-of-the-art models.