Findings of the WMT 2019 Shared Task on Parallel Corpus Filtering for Low-Resource Conditions
Findings of the WMT 2019 Shared Task on Parallel Corpus Filtering for Low-Resource Conditions
复制标题
WMT 2019 共享任务针对低资源条件的并行语料库过滤的结果
DOI:
10.18653/v1/w19-5404
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
J. Pino
中科院分区:
文献类型:
--
作者:
Philipp Koehn;Francisco Guzmán;Vishrav Chaudhary;J. Pino
Following the WMT 2018 Shared Task on Parallel Corpus Filtering, we posed the challenge of assigning sentence-level quality scores for very noisy corpora of sentence pairs crawled from the web, with the goal of sub-selecting 2% and 10% of the highest-quality data to be used to train machine translation systems. This year, the task tackled the low resource condition of Nepali-English and Sinhala-English. Eleven participants from companies, national research labs, and universities participated in this task.