VALHALLA: Visual Hallucination for Machine Translation

VALHALLA: Visual Hallucination for Machine Translation
复制标题

DOI:
10.1109/cvpr52688.2022.00515
复制
发表时间:
2022-05
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yi Li-;Rameswar Panda;Yoon Kim;Chun-Fu Chen;R. Feris;David Cox;N. Vasconcelos
Yi Li-;Rameswar Panda;Yoon Kim;Chun-Fu Chen;R. Feris;David Cox;N. Vasconcelos
中科院分区:
其他
文献类型:
--
作者:
Yi Li-;Rameswar Panda;Yoon Kim;Chun-Fu Chen;R. Feris;David Cox;N. Vasconcelos

文献摘要

被引文献

相似文献

近年来,通过考虑图像等辅助输入来设计更好的机器翻译系统引起了人们的广泛关注。虽然现有的方法显示出比传统的纯文本翻译系统更有前途的性能,但它们通常需要成对的文本和图像作为推理过程中的输入,这限制了它们对现实世界场景的适用性。在本文中,我们介绍了一个视觉幻觉框架,称为VALHALLA,它只需要在推理时的源语句,而不是使用幻觉的视觉表示多模态机器翻译。特别地,给定源句子,自回归幻觉Transformer用于从输入文本预测离散视觉表示,并且组合的文本和幻觉表示用于获得目标翻译。我们使用具有交叉熵损失的标准反向传播与翻译Transformer联合训练幻觉Transformer,同时由额外的损失指导,该损失鼓励使用地面实况或幻觉视觉表示的预测之间的一致性。在三个标准翻译数据集上进行的大量实验表明,我们的方法在纯文本基线和最先进的方法上都是有效的。项目页面:http://www.svcl.ucsd.jects/valhalla.edu/pro。
Designing better machine translation systems by considering auxiliary inputs such as images has attracted much attention in recent years. While existing methods show promising performance over the conventional text-only translation systems, they typically require paired text and image as input during inference, which limits their applicability to real-world scenarios. In this paper, we introduce a visual hallucination framework, called VALHALLA, which requires only source sentences at inference time and instead uses hallucinated visual representations for multi-modal machine translation. In particular, given a source sentence an autoregressive hallucination transformer is used to predict a discrete visual representation from the input text, and the combined text and hallucinated representations are utilized to obtain the target translation. We train the hallucination transformer jointly with the translation transformer using standard backpropagation with crossentropy losses while being guided by an additional loss that encourages consistency between predictions using either groundtruth or hallucinated visual representations. Extensive experiments on three standard translation datasets with a diverse set of language pairs demonstrate the effectiveness of our approach over both text-only baselines and state-of-the-art methods. Project page: http://www.svcl.ucsd.jects/valhalla.edu/pro.