Efficient Web Crawling for Large Text Corpora
Efficient Web Crawling for Large Text Corpora
复制标题
大型文本语料库的高效网络爬行
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Jan Pomikálek
中科院分区:
文献类型:
--
作者:
Vít Suchomel;Jan Pomikálek
Many researchers use texts from the web, an easy source of linguistic data in a great variety of languages. Building both large and good quality text corpora is the challenge we face nowadays. We describe how to deal with inefficient data downloading and how to focus crawling on text rich web domains. We present efficiency figures from crawling texts in American Spanish, Czech, Japanese, Russian, Tajik Persian, Turkish and the sizes of the resulting corpora. The idea has been successfully applied for building billions of words scale corpora in six languages. Texts in the Russian corpus, consisting of 20.2 billions tokens, were downloaded in just 13 days.