OpusTools and Parallel Corpus Diagnostics

OpusTools and Parallel Corpus Diagnostics
复制标题

OpusTools 和并行语料库诊断

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
J. Tiedemann
J. Tiedemann
中科院分区:
--
文献类型:
--
作者:
Mikko Aulamo;U. Sulubacak;Sami Virpioja;J. Tiedemann

文献摘要

被引文献

相似文献

本文介绍了OpusTools,一个用于下载和处理OPUS语料库中包含的平行语料库的软件包。该软件包实现了用于访问压缩数据的工具,这些压缩数据是以其归档发布格式发布的,并且可以轻松地在常见格式之间进行转换。OpusTools还包括用于语言识别和数据过滤的工具,以及用于将各种来源的数据导入OPUS格式的工具。我们展示了这些工具在并行语料库创建和数据诊断中的使用。后者对于在大量数据集中查明潜在问题和错误特别有用。使用这些工具,我们现在可以监测数据集的有效性,提高数据收集的整体质量和一致性。
This paper introduces OpusTools, a package for downloading and processing parallel corpora included in the OPUS corpus collection. The package implements tools for accessing compressed data in their archived release format and make it possible to easily convert between common formats. OpusTools also includes tools for language identification and data filtering as well as tools for importing data from various sources into the OPUS format. We show the use of these tools in parallel corpus creation and data diagnostics. The latter is especially useful for the identification of potential problems and errors in the extensive data set. Using these tools, we can now monitor the validity of data sets and improve the overall quality and consitency of the data collection.