On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation Evaluation

On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation Evaluation
复制标题

DOI:
10.18653/v1/2020.acl-main.151
复制
发表时间:
2020-05
期刊:
--
影响因子:
--
通讯作者:
Wei Zhao;Goran Glavavs;Maxime Peyrard;Yang Gao-;Robert West;Steffen Eger
Wei Zhao;Goran Glavavs;Maxime Peyrard;Yang Gao-;Robert West;Steffen Eger
中科院分区:
其他
文献类型:
--
作者:
Wei Zhao;Goran Glavavs;Maxime Peyrard;Yang Gao-;Robert West;Steffen Eger

文献摘要

被引文献

相似文献

跨语言编码器的评估通常通过有监督的下游任务中的零镜头跨语言传输或通过无监督的跨语言文本相似性来执行。在本文中,我们关注的是无参考机器翻译(MT)评估,其中我们直接将源文本与(有时是低质量的)系统翻译进行比较,这代表了多语言编码器的自然对抗设置。无参考评估有望对机器翻译系统进行网络规模的比较。我们系统地研究了一系列基于最先进的跨语言语义表示的指标,这些语义表示是通过预训练的M-BERT和LASER获得的。我们发现,它们作为无参考机器翻译评估的语义编码器表现不佳,并确定了它们的两个关键限制,即,(a)相互翻译表示之间的语义不匹配,以及更突出的是,(B)无法惩罚“翻译”,即,低质量的直译。我们提出了两个部分补救措施:(1)事后重新对齐的向量空间和(2)耦合的语义相似性为基础的指标与目标端语言建模。在段级MT评估中,我们的最佳度量超过基于参考的BLEU 5.7个相关点。
Evaluation of cross-lingual encoders is usually performed either via zero-shot cross-lingual transfer in supervised downstream tasks or via unsupervised cross-lingual textual similarity. In this paper, we concern ourselves with reference-free machine translation (MT) evaluation where we directly compare source texts to (sometimes low-quality) system translations, which represents a natural adversarial setup for multilingual encoders. Reference-free evaluation holds the promise of web-scale comparison of MT systems. We systematically investigate a range of metrics based on state-of-the-art cross-lingual semantic representations obtained with pretrained M-BERT and LASER. We find that they perform poorly as semantic encoders for reference-free MT evaluation and identify their two key limitations, namely, (a) a semantic mismatch between representations of mutual translations and, more prominently, (b) the inability to punish “translationese”, i.e., low-quality literal translations. We propose two partial remedies: (1) post-hoc re-alignment of the vector spaces and (2) coupling of semantic-similarity based metrics with target-side language modeling. In segment-level MT evaluation, our best metric surpasses reference-based BLEU by 5.7 correlation points.