Word-Region Alignment-Guided Multimodal Neural Machine Translation

Word-Region Alignment-Guided Multimodal Neural Machine Translation
复制标题

DOI:
10.1109/taslp.2021.3138719
复制
发表时间:
2022
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Yuting Zhao;Mamoru Komachi;Tomoyuki Kajiwara;Chenhui Chu
Yuting Zhao;Mamoru Komachi;Tomoyuki Kajiwara;Chenhui Chu
中科院分区:
其他
文献类型:
--
作者:
Yuting Zhao;Mamoru Komachi;Tomoyuki Kajiwara;Chenhui Chu

文献摘要

相似文献

我们提出了词区域对齐引导的多模式神经机器翻译(MNMT),该模型通过词区域对齐(WRA)将文本和视觉通道之间的语义关联联系起来。已有的关于MNMT的研究主要集中在视觉和文本通道的整合效应上。然而,它们没有利用这两个情态之间的语义相关性。通过引入WRA作为一座桥梁,我们提出了MNMT中文本和视觉通道之间的语义关联。该方案已在两种主流神经机器翻译体系结构上实现:递归神经网络(RNN)和转换器。在使用Multi30k数据集的英语-德语和英语-法语翻译任务和使用Flickr30kEnt-JP数据集的英日翻译任务上的实验证明,我们的模型在不同评估指标的竞争基线方面有显著的改善,并优于大多数现有的MNMT模型。例如,在Multi30k Test2016测试集上,英语-德语任务的BLEU分数提高了1.0%,英语-法语任务的BLEU分数提高了1.1%;在Flickr30kEnt-JP测试集上,英语-日语任务的BLEU分数提高了0.7%。进一步的分析表明,我们的模型可以通过整合WRA来获得更好的翻译性能,从而更好地利用视觉信息。
We propose word-region alignment-guided multimodal neural machine translation (MNMT), a novel model for MNMT that links the semantic correlation between textual and visual modalities using word-region alignment (WRA). Existing studies on MNMT have mainly focused on the effect of integrating visual and textual modalities. However, they do not leverage the semantic relevance between the two modalities. We advance the semantic correlation between textual and visual modalities in MNMT by incorporating WRA as a bridge. This proposal has been implemented on two mainstream architectures of neural machine translation (NMT): the recurrent neural network (RNN) and the transformer. Experiments on two public benchmarks, English–German and English–French translation tasks using the Multi30k dataset and English–Japanese translation tasks using the Flickr30kEnt-JP dataset prove that our model has a significant improvement with respect to the competitive baselines across different evaluation metrics and outperforms most of the existing MNMT models. For example, 1.0 BLEU scores are improved for the English–German task and 1.1 BLEU scores are improved for the English–French task on the Multi30k test2016 set; and 0.7 BLEU scores are improved for the English–Japanese task on the Flickr30kEnt-JP test set. Further analysis demonstrates that our model can achieve better translation performance by integrating WRA, leading to better visual information use.