Discourse-Related Language Contrasts in English-Croatian Human and Machine Translation

Discourse-Related Language Contrasts in English-Croatian Human and Machine Translation
复制标题

英语-克罗地亚语人机翻译中与语篇相关的语言对比

DOI:
--
复制
发表时间:
2018
期刊:
Conference on Machine Translation
影响因子:
--
通讯作者:
Sara Stymne
Sara Stymne
中科院分区:
--
文献类型:
--
作者:
Margita Sostaric;Christian Hardmeier;Sara Stymne

文献摘要

被引文献

相似文献

我们提出了一个共指现象在英语-克罗地亚人和机器翻译的分析。其目的是阐明这些结构不同的语言利用话语信息的方式的差异,并为话语感知机器翻译系统的开发提供见解。使用解析器和单词对齐工具产生的注释,在并行数据中自动识别这些现象,使我们能够在两种语言中精确定位感兴趣的模式。我们使分析更细粒度的,包括三个语料库属于三个不同的寄存器。在第二步中,我们创建了一个测试集,具有挑战性的语言结构,并用它来评估三个MT系统的性能。我们发现,SMT和NMT系统的斗争与处理这些话语现象,即使NMT往往比SMT表现得更好。通过对实际语言使用中经常出现的模式的概述,以及指出当前机器翻译系统的弱点,通常误译他们,我们希望有助于解决的问题,话语现象在机器翻译应用的努力。
We present an analysis of a number of coreference phenomena in English-Croatian human and machine translations. The aim is to shed light on the differences in the way these structurally different languages make use of discourse information and provide insights for discourse-aware machine translation system development. The phenomena are automatically identified in parallel data using annotation produced by parsers and word alignment tools, enabling us to pinpoint patterns of interest in both languages. We make the analysis more fine-grained by including three corpora pertaining to three different registers. In a second step, we create a test set with the challenging linguistic constructions and use it to evaluate the performance of three MT systems. We show that both SMT and NMT systems struggle with handling these discourse phenomena, even though NMT tends to perform somewhat better than SMT. By providing an overview of patterns frequently occurring in actual language use, as well as by pointing out the weaknesses of current MT systems that commonly mistranslate them, we hope to contribute to the effort of resolving the issue of discourse phenomena in MT applications.