Double Attention-based Multimodal Neural Machine Translation with Semantic Image Regions

Double Attention-based Multimodal Neural Machine Translation with Semantic Image Regions
复制标题

DOI:
--
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Yuting Zhao;Mamoru Komachi;Tomoyuki Kajiwara;Chenhui Chu
Yuting Zhao;Mamoru Komachi;Tomoyuki Kajiwara;Chenhui Chu
中科院分区:
其他
文献类型:
--
作者:
Yuting Zhao;Mamoru Komachi;Tomoyuki Kajiwara;Chenhui Chu

文献摘要

相似文献

现有的多模态神经机器翻译(MNMT)研究主要集中在结合视觉和文本模态来改善翻译的效果。然而,有人认为视觉方式的益处有限。传统的视觉注意机制已用于从卷积神经网络(CNN)生成的同等大小的网格中选择视觉特征,并且对于对齐与文本对象相关的视觉概念可能具有适度的影响,因为网格视觉特征不捕获语义信息。相比之下,我们提出通过使用两种单独的注意机制(双重注意)集成视觉和文本特征来将语义图像区域应用于 MNMT。我们在 Multi30k 数据集上进行了实验,与具有网格视觉特征的 MNMT 相比,英德和英法翻译任务的 BLEU 分数分别提高了 0.5 和 0.9。我们还展示了语义图像区域对翻译性能的具体改进。
Existing studies on multimodal neural machine translation (MNMT) have mainly focused on the effect of combining visual and textual modalities to improve translations. However, it has been suggested that the visual modality is only marginally beneficial. Conventional visual attention mechanisms have been used to select the visual features from equally-sized grids generated by convolutional neural networks (CNNs), and may have had modest effects on aligning the visual concepts associated with textual objects, because the grid visual features do not capture semantic information. In contrast, we propose the application of semantic image regions for MNMT by integrating visual and textual features using two individual attention mechanisms (double attention). We conducted experiments on the Multi30k dataset and achieved an improvement of 0.5 and 0.9 BLEU points for English-German and English-French translation tasks, compared with the MNMT with grid visual features. We also demonstrated concrete improvements on translation performance benefited from semantic image regions.