Efficient Web Crawling for Large Text Corpora

Efficient Web Crawling for Large Text Corpora
复制标题

大型文本语料库的高效网络爬行

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Jan Pomikálek
Jan Pomikálek
中科院分区:
--
文献类型:
--
作者:
Vít Suchomel;Jan Pomikálek

文献摘要

被引文献

相似文献

许多研究人员使用来自网络的文本,这是一个以各种语言提供的语言数据的简单来源。构建既大又高质量的文本语料库是我们目前面临的挑战。我们描述了如何处理低效的数据下载,以及如何专注于对文本丰富的Web域进行爬行。我们给出了用美国西班牙语、捷克语、日语、俄语、塔吉克语、波斯语、土耳其语爬行文本的效率数字以及由此产生的语料库的大小。这一思想已成功应用于六种语言数十亿字规模的语料库建设。俄语语料库中的文本,由202亿个令牌组成,在短短13天内被下载。
Many researchers use texts from the web, an easy source of linguistic data in a great variety of languages. Building both large and good quality text corpora is the challenge we face nowadays. We describe how to deal with inefficient data downloading and how to focus crawling on text rich web domains. We present efficiency figures from crawling texts in American Spanish, Czech, Japanese, Russian, Tajik Persian, Turkish and the sizes of the resulting corpora. The idea has been successfully applied for building billions of words scale corpora in six languages. Texts in the Russian corpus, consisting of 20.2 billions tokens, were downloaded in just 13 days.