A Corpus for English-Japanese Multimodal Neural Machine Translation with Comparable Sentences

A Corpus for English-Japanese Multimodal Neural Machine Translation with Comparable Sentences
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Andrew C. Merritt;Chenhui Chu;Yuki Arase
Andrew C. Merritt;Chenhui Chu;Yuki Arase
中科院分区:
其他
文献类型:
--
作者:
Andrew C. Merritt;Chenhui Chu;Yuki Arase

文献摘要

被引文献

相似文献

多年来,多模态神经机器翻译 (NMT) 已成为一个日益重要的研究领域,因为图像数据等附加模态可以为文本数据提供更多上下文。此外,由于带有图像的并行句子的可用性较低,特别是对于英语-日语数据,在没有大型并行语料库的情况下训练多模态 NMT 模型的可行性仍在研究中。然而,这个空白可以用包含双语术语和平行短语的类似句子来填补,这些句子是通过社交网络帖子和电子商务产品描述等媒体自然创建的。在本文中,我们提出了一个新的多模态英语-日语语料库,其中包含从现有图像字幕数据集编译而来的可比句子。此外,我们还使用较小的平行语料库来补充可比较的句子,以进行验证和测试。为了测试这种可比句子翻译场景的性能,我们使用可比语料库训练了几个基线 NMT 模型,并评估了它们的英日翻译性能。由于我们的基线实验中翻译得分较低,我们认为当前的多模态 NMT 模型并未设计为有效利用可比较的句子数据。尽管如此,我们希望我们的语料库能够用于进一步研究具有可比句子的多模态 NMT。
Multimodal neural machine translation (NMT) has become an increasingly important area of research over the years because additional modalities, such as image data, can provide more context to textual data. Furthermore, the viability of training multimodal NMT models without a large parallel corpus continues to be investigated due to low availability of parallel sentences with images, particularly for English-Japanese data. However, this void can be filled with comparable sentences that contain bilingual terms and parallel phrases, which are naturally created through media such as social network posts and e-commerce product descriptions. In this paper, we propose a new multimodal English-Japanese corpus with comparable sentences that are compiled from existing image captioning datasets. In addition, we supplement our comparable sentences with a smaller parallel corpus for validation and test purposes. To test the performance of this comparable sentence translation scenario, we train several baseline NMT models with our comparable corpus and evaluate their English-Japanese translation performance. Due to low translation scores in our baseline experiments, we believe that current multimodal NMT models are not designed to effectively utilize comparable sentence data. Despite this, we hope for our corpus to be used to further research into multimodal NMT with comparable sentences.