The Good, the Bad, and the Hazy: Design Decisions in Web Corpus Construction

The Good, the Bad, and the Hazy: Design Decisions in Web Corpus Construction
复制标题

好的、坏的和模糊的:网络语料库构建中的设计决策

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Felix Bildhauer
Felix Bildhauer
中科院分区:
--
文献类型:
--
作者:
R. Schäfer;A. Barbaresi;Felix Bildhauer

文献摘要

被引文献

相似文献

在本文中,我们在网络语料库建设的背景下考察了文本质量的概念。Web文档通常包含使其无法被纳入语料库的材料(标签云、姓名或名词列表等)。首先,我们看看编码者(特别是语料库设计者)之间的协议,因为他们的任务是对文本质量进行评级。然后,我们对一种基于短词和高频词的简单、完全无监督的文本质量评估方法进行了评估。最后,我们描述了我们构建仔细清理和非破坏性规范化网络语料库的一般方法。在这种方法下,我们使用质量度量来注释文档,而不是实际删除那些被归类为低质量的文档。
In this paper, we examine notions of text quality in the context of web corpus construction. Web documents often contain material which disqualifies them from inclusion in a corpus (tag clouds, lists of names or nouns, etc.). First, we look at the agreement between coders (especially corpus designers) given the task of rating text quality. Then, we evaluate a simple and fully unsupervised method of text quality assessment based on short and very frequent words. Finally, we describe our general approach to the construction of carefully cleansed and non-destructively normalized web corpora. Under this approach, we annotate documents with quality metrics instead of actually removing those documents classified as being of low quality.