Strong and Simple Baselines for Multimodal Utterance Embeddings

Strong and Simple Baselines for Multimodal Utterance Embeddings
复制标题

DOI:
10.18653/v1/n19-1267
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
P. Liang;Y. Lim;Yao-Hung Hubert Tsai;R. Salakhutdinov;Louis-Philippe Morency
P. Liang;Y. Lim;Yao-Hung Hubert Tsai;R. Salakhutdinov;Louis-Philippe Morency
中科院分区:
其他
文献类型:
--
作者:
P. Liang;Y. Lim;Yao-Hung Hubert Tsai;R. Salakhutdinov;Louis-Philippe Morency

文献摘要

被引文献

相似文献

人类语言是一种丰富的多模态信号,包括口语、面部表情、肢体动作和语音语调。由于存在多种异构信息源,学习这些口语话语的表示是一个复杂的研究问题。多模态学习的最新进展遵循了建立更复杂的模型的大趋势,这些模型利用了各种注意力、记忆和循环成分。在本文中,我们提出了两个简单但强大的基线来学习多模态话语的嵌入。第一个基线假设有条件地将话语分解为单峰因素。每个单峰因子使用通过嵌入的线性变换得到的简单似然函数来建模。我们证明了最优嵌入可以通过单峰特征的加权平均以封闭形式导出。为了捕获更丰富的表示,我们的第二个基线通过分解为单峰、双峰和三峰因素来扩展第一个基线,同时在学习和推理过程中保持简单性和效率。从跨两个任务的一组实验中,我们在监督和半监督多模态预测上都表现出很强的性能,并且在推理过程中比神经模型显著(10倍)加速。总的来说,我们相信我们强大的基线模型为未来的多模态学习研究提供了新的基准选择。
Human language is a rich multimodal signal consisting of spoken words, facial expressions, body gestures, and vocal intonations. Learning representations for these spoken utterances is a complex research problem due to the presence of multiple heterogeneous sources of information. Recent advances in multimodal learning have followed the general trend of building more complex models that utilize various attention, memory and recurrent components. In this paper, we propose two simple but strong baselines to learn embeddings of multimodal utterances. The first baseline assumes a conditional factorization of the utterance into unimodal factors. Each unimodal factor is modeled using the simple form of a likelihood function obtained via a linear transformation of the embedding. We show that the optimal embedding can be derived in closed form by taking a weighted average of the unimodal features. In order to capture richer representations, our second baseline extends the first by factorizing into unimodal, bimodal, and trimodal factors, while retaining simplicity and efficiency during learning and inference. From a set of experiments across two tasks, we show strong performance on both supervised and semi-supervised multimodal prediction, as well as significant (10 times) speedups over neural models during inference. Overall, we believe that our strong baseline models offer new benchmarking options for future research in multimodal learning.