Findings of the WMT 2019 Shared Task on Parallel Corpus Filtering for Low-Resource Conditions

Findings of the WMT 2019 Shared Task on Parallel Corpus Filtering for Low-Resource Conditions
复制标题

WMT 2019 共享任务针对低资源条件的并行语料库过滤的结果

DOI:
10.18653/v1/w19-5404
复制
发表时间:
2019
期刊:
POLIBITS
影响因子:
--
通讯作者:
J. Pino
J. Pino
中科院分区:
--
文献类型:
--
作者:
Philipp Koehn;Francisco Guzmán;Vishrav Chaudhary;J. Pino

文献摘要

被引文献

相似文献

继 WMT 2018 并行语料库过滤共享任务之后,我们提出了为从网络爬取的非常嘈杂的句子对语料库分配句子级质量分数的挑战,目标是分选 2% 和 10% 的最高质量数据用于训练机器翻译系统。今年,该任务解决了尼泊尔语-英语和僧伽罗语-英语资源匮乏的问题。来自公司、国家研究实验室和大学的 11 名参与者参与了这项任务。
Following the WMT 2018 Shared Task on Parallel Corpus Filtering, we posed the challenge of assigning sentence-level quality scores for very noisy corpora of sentence pairs crawled from the web, with the goal of sub-selecting 2% and 10% of the highest-quality data to be used to train machine translation systems. This year, the task tackled the low resource condition of Nepali-English and Sinhala-English. Eleven participants from companies, national research labs, and universities participated in this task.