Integrating Multimodal Information in Large Pretrained Transformers.

Integrating Multimodal Information in Large Pretrained Transformers.
复制标题

DOI:
10.18653/v1/2020.acl-main.214
复制
发表时间:
2020-07
期刊:
Proceedings of the conference. Association for Computational Linguistics. Meeting
影响因子:
--
通讯作者:
Hoque E
Hoque E
中科院分区:
其他
文献类型:
--
作者:
Rahman W;Hasan MK;Lee S;Zadeh A;Mao C;Morency LP;Hoque E

文献摘要

被引文献

相似文献

最近基于Transformer的上下文单词表示,包括BERT和XLNet,在NLP的多个学科中显示了最先进的性能。在特定于任务的数据集上微调训练的上下文模型一直是实现下游卓越性能的关键。虽然微调这些预先训练的模型对于词汇应用程序(只有语言情态的应用程序)是直接的,但对于多模式语言(NLP中一个日益增长的领域,专注于建模面对面的交流)来说,这并不是微不足道的。预先训练的模型没有必要的组件来接受视觉和声学这两种额外的模式。在本文中,我们提出了一种附加到ERT和XLNet上的多模式适配门(MAG)。MAG允许Bert和XLNet在微调期间接受多模式非语言数据。它通过产生向BERT和XLNet的内部表示的转变来做到这一点;这种转变取决于视觉和听觉形式。在我们的实验中,我们研究了用于多通道情感分析的常用的CMU-MOSI和CMU-MOSEI数据集。MAG-BERT和MAG-XLNet的微调显著提高了之前基线的情感分析性能,以及BERT和XLNet的仅语言微调。在CMU-MOSI数据集上,MAG-XLNet在自然语言处理领域首次实现了人类级别的多通道情感分析。
Recent Transformer-based contextual word representations, including BERT and XLNet, have shown state-of-the-art performance in multiple disciplines within NLP. Fine-tuning the trained contextual models on task-specific datasets has been the key to achieving superior performance downstream. While fine-tuning these pre-trained models is straight-forward for lexical applications (applications with only language modality), it is not trivial for multimodal language (a growing area in NLP focused on modeling face-to-face communication). Pre-trained models don’t have the necessary components to accept two extra modalities of vision and acoustic. In this paper, we proposed an attachment to BERT and XLNet called Multimodal Adaptation Gate (MAG). MAG allows BERT and XLNet to accept multimodal nonverbal data during fine-tuning. It does so by generating a shift to internal representation of BERT and XLNet; a shift that is conditioned on the visual and acoustic modalities. In our experiments, we study the commonly used CMU-MOSI and CMU-MOSEI datasets for multimodal sentiment analysis. Fine-tuning MAG-BERT and MAG-XLNet significantly boosts the sentiment analysis performance over previous baselines as well as language-only fine-tuning of BERT and XLNet. On the CMU-MOSI dataset, MAG-XLNet achieves human-level multimodal sentiment analysis performance for the first time in the NLP community.