Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph
复制标题

DOI:
10.18653/v1/p18-1208
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
Amir Zadeh;P. Liang;Soujanya Poria;E. Cambria;Louis-Philippe Morency
Amir Zadeh;P. Liang;Soujanya Poria;E. Cambria;Louis-Philippe Morency
中科院分区:
其他
文献类型:
--
作者:
Amir Zadeh;P. Liang;Soujanya Poria;E. Cambria;Louis-Philippe Morency

文献摘要

被引文献

相似文献

分析人类多模态语言是自然语言处理(NLP)中一个新兴的研究领域。从本质上讲,这种语言是多模态(异构的)、顺序的和异步的;它由语言(词汇)、视觉(表情)和声学(副语言)模态组成,所有这些模态都以异步协调序列的形式存在。从资源角度来看,确实需要大规模的数据集,以便对这种形式的语言进行深入研究。在本文中,我们介绍了卡内基梅隆大学多模态观点情感和情绪强度(CMU - MOSEI)数据集,这是迄今为止最大的情感分析和情绪识别数据集。利用来自CMU - MOSEI的数据以及一种名为动态融合图(DFG)的新型多模态融合技术,我们进行了实验,以探究在人类多模态语言中各种模态是如何相互作用的。与先前提出的融合技术不同,DFG具有很高的可解释性,并且与先前的先进技术相比,具有相当的性能。
Analyzing human multimodal language is an emerging area of research in NLP. Intrinsically this language is multimodal (heterogeneous), sequential and asynchronous; it consists of the language (words), visual (expressions) and acoustic (paralinguistic) modalities all in the form of asynchronous coordinated sequences. From a resource perspective, there is a genuine need for large scale datasets that allow for in-depth studies of this form of language. In this paper we introduce CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI), the largest dataset of sentiment analysis and emotion recognition to date. Using data from CMU-MOSEI and a novel multimodal fusion technique called the Dynamic Fusion Graph (DFG), we conduct experimentation to exploit how modalities interact with each other in human multimodal language. Unlike previously proposed fusion techniques, DFG is highly interpretable and achieves competative performance when compared to the previous state of the art.