Found in Translation: Learning Robust Joint Representations by Cyclic Translations Between Modalities

Found in Translation: Learning Robust Joint Representations by Cyclic Translations Between Modalities
复制标题

DOI:
10.1609/aaai.v33i01.33016892
复制
发表时间:
2018-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Hai Pham;P. Liang;Thomas Manzini;Louis-Philippe Morency;B. Póczos
Hai Pham;P. Liang;Thomas Manzini;Louis-Philippe Morency;B. Póczos
中科院分区:
其他
文献类型:
--
作者:
Hai Pham;P. Liang;Thomas Manzini;Louis-Philippe Morency;B. Póczos

文献摘要

被引文献

相似文献

多模态情感分析是研究说话人从语言、视觉和听觉三方面表达情感的一个核心研究领域。多模态学习的核心挑战涉及推断可以处理和关联来自这些模态的信息的联合表征。然而,现有的工作通过要求所有模态作为输入来学习联合表征,因此,学习到的表征可能对测试时的噪声或缺失模态敏感。随着最近序列到序列(Seq2Seq)模型在机器翻译中的成功,有机会探索学习联合表示的新方法,这些方法在测试时可能不需要所有输入模态。在本文中,我们提出了一种通过模态之间的转换来学习鲁棒联合表示的方法。我们的方法基于一个关键的见解,即从源模态到目标模态的转换提供了一种仅使用源模态作为输入来学习联合表征的方法。我们通过循环一致性损失来增强模态翻译,以确保我们的联合表示保留了所有模态的最大信息。一旦我们的翻译模型使用配对的多模态数据进行训练,我们只需要在测试时来自源模态的数据来进行最终的情感预测。这确保了我们的模型在其他模态的扰动或缺失信息中保持鲁棒性。我们用一个耦合的翻译预测目标来训练我们的模型,它在多模态情感分析数据集上获得了最新的结果:CMU-MOSI, ICTMMMO和YouTube。另外的实验表明,我们的模型学习越来越判别联合表示与更多的输入模态,同时保持鲁棒性缺失或扰动模态。
Multimodal sentiment analysis is a core research area that studies speaker sentiment expressed from the language, visual, and acoustic modalities. The central challenge in multimodal learning involves inferring joint representations that can process and relate information from these modalities. However, existing work learns joint representations by requiring all modalities as input and as a result, the learned representations may be sensitive to noisy or missing modalities at test time. With the recent success of sequence to sequence (Seq2Seq) models in machine translation, there is an opportunity to explore new ways of learning joint representations that may not require all input modalities at test time. In this paper, we propose a method to learn robust joint representations by translating between modalities. Our method is based on the key insight that translation from a source to a target modality provides a method of learning joint representations using only the source modality as input. We augment modality translations with a cycle consistency loss to ensure that our joint representations retain maximal information from all modalities. Once our translation model is trained with paired multimodal data, we only need data from the source modality at test time for final sentiment prediction. This ensures that our model remains robust from perturbations or missing information in the other modalities. We train our model with a coupled translationprediction objective and it achieves new state-of-the-art results on multimodal sentiment analysis datasets: CMU-MOSI, ICTMMMO, and YouTube. Additional experiments show that our model learns increasingly discriminative joint representations with more input modalities while maintaining robustness to missing or perturbed modalities.