Building Large Corpora from the Web Using a New Efficient Tool Chain

Building Large Corpora from the Web Using a New Efficient Tool Chain
复制标题

使用新的高效工具链从网络构建大型语料库

DOI:
--
复制
发表时间:
2012
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Felix Bildhauer
Felix Bildhauer
中科院分区:
--
文献类型:
--
作者:
R. Schäfer;Felix Bildhauer

文献摘要

被引文献

相似文献

近十年来,网络语料库构建方法和网络语料库评价方法得到了积极的研究。值得注意的是,WaCky 计划为选定的欧洲语言提供了理论结果和一组网络语料库。我们提供了一个用于网络语料库构建的软件工具包以及使用该软件构建的一组更大的语料库(多达超过 90 亿个令牌)。首先,我们讨论如何收集数据以确保数据不偏向某些主机。然后,我们描述我们的软件工具包,它执行基本的清理以及样板删除、简单的连接文本检测以及从语料库中删除重复项的重叠。我们最终报告了迄今为止构建的语料库的评估结果,例如 w.r.t.包含的重复量和文本类型/流派分布。在适用的情况下,我们将我们的语料库与 WaCky 语料库进行比较,因为我们认为将网络语料库与传统或平衡语料库进行比较是不合适的。虽然我们使用 WaCky 计划应用的一些方法,但我们可以表明我们已经引入了渐进式改进。
Over the last decade, methods of web corpus construction and the evaluation of web corpora have been actively researched. Prominently, the WaCky initiative has provided both theoretical results and a set of web corpora for selected European languages. We present a software toolkit for web corpus construction and a set of siginificantly larger corpora (up to over 9 billion tokens) built using this software. First, we discuss how the data should be collected to ensure that it is not biased towards certain hosts. Then, we describe our software toolkit which performs basic cleanups as well as boilerplate removal, simple connected text detection as well as shingling to remove duplicates from the corpora. We finally report evaluation results of the corpora built so far, for example w.r.t. the amount of duplication contained and the text type/genre distribution. Where applicable, we compare our corpora to the WaCky corpora, since it is inappropriate, in our view, to compare web corpora to traditional or balanced corpora. While we use some methods applied by the WaCky initiative, we can show that we have introduced incremental improvements.