OpusFilter: A Configurable Parallel Corpus Filtering Toolbox

OpusFilter: A Configurable Parallel Corpus Filtering Toolbox
复制标题

OpusFilter:可配置的并行语料库过滤工具箱

DOI:
10.18653/v1/2020.acl-demos.20
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
J. Tiedemann
J. Tiedemann
中科院分区:
--
文献类型:
--
作者:
Mikko Aulamo;Sami Virpioja;J. Tiedemann

文献摘要

被引文献

相似文献

本文介绍了一个灵活的、模块化的平行语料库过滤工具箱OpusFilter。它实现了许多基于启发式过滤器、语言识别库、基于字符的语言模型和单词对齐工具的组件,并且可以很容易地使用自定义过滤器进行扩展。可以使用单个特征或Logistic回归模型根据它们的质量或领域匹配来对比特段进行排名,所述Logistic回归模型可以在不手动标记训练数据的情况下进行训练。我们以一个基于噪声网络爬行的训练数据的芬兰语-英语新闻翻译任务为例,验证了OpusFilter的有效性。应用我们的工具提高了翻译质量,同时显著减少了训练数据的大小,也明显优于爬行数据集中给出的替代排名。此外,我们还展示了OpusFilter执行领域适配的数据选择的能力。
This paper introduces OpusFilter, a flexible and modular toolbox for filtering parallel corpora. It implements a number of components based on heuristic filters, language identification libraries, character-based language models, and word alignment tools, and it can easily be extended with custom filters. Bitext segments can be ranked according to their quality or domain match using single features or a logistic regression model that can be trained without manually labeled training data. We demonstrate the effectiveness of OpusFilter on the example of a Finnish-English news translation task based on noisy web-crawled training data. Applying our tool leads to improved translation quality while significantly reducing the size of the training data, also clearly outperforming an alternative ranking given in the crawled data set. Furthermore, we show the ability of OpusFilter to perform data selection for domain adaptation.