Multimodal Routing: Improving Local and Global Interpretability of Multimodal Language Analysis.

Multimodal Routing: Improving Local and Global Interpretability of Multimodal Language Analysis.
复制标题

DOI:
10.18653/v1/2020.emnlp-main.143
复制
发表时间:
2020-11
期刊:
Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
Morency LP
Morency LP
中科院分区:
其他
文献类型:
--
作者:
Tsai YH;Ma MQ;Yang M;Salakhutdinov R;Morency LP

文献摘要

被引文献

相似文献

人类语言可以通过多种信息源来表达,这些信息源被称为模态,包括声调、面部手势和口语。最近在情感分析和情感识别等以人为中心的任务上表现出色的多模态学习通常是黑箱式的,可解释性非常有限。在本文中,我们提出了多模态路由,它动态地调整每个输入样本的输入模态和输出表示之间的权重。多模态路由可以识别单个模态和跨模态特征的相对重要性。此外,路由的权重分配使我们不仅可以在全局(即整个数据集的一般趋势)上解释模态预测关系,而且可以在局部解释每个单一输入样本,同时与最先进的方法相比保持竞争性能。
The human language can be expressed through multiple sources of information known as modalities, including tones of voice, facial gestures, and spoken language. Recent multimodal learning with strong performances on human-centric tasks such as sentiment analysis and emotion recognition are often black-box, with very limited interpretability. In this paper we propose Multimodal Routing, which dynamically adjusts weights between input modalities and output representations differently for each input sample. Multimodal routing can identify relative importance of both individual modalities and cross-modality features. Moreover, the weight assignment by routing allows us to interpret modality-prediction relationships not only globally (i.e. general trends over the whole dataset), but also locally for each single input sample, mean-while keeping competitive performance compared to state-of-the-art methods.