A Corpus Factory for Many Languages

A Corpus Factory for Many Languages
复制标题

多种语言的语料库工厂

DOI:
--
复制
发表时间:
2010
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Avinesh P.V.S
Avinesh P.V.S
中科院分区:
--
文献类型:
--
作者:
A. Kilgarriff;Siva Reddy;Jan Pomikálek;Avinesh P.V.S

文献摘要

被引文献

相似文献

对于许多语言,没有可用的大型通用语料库。在互联网出现之前,除了这些机构之外,所有机构都无能为力,只能沮丧地摇头,因为语料库的建设既漫长又缓慢,而且成本高昂。但随着网络的出现,它可以高度自动化,因此速度快,成本低。我们已经开发了一个语料库工厂,在那里我们建立了大型语料库。在这篇文章中,我们描述了我们使用的方法,以及它是如何工作的,以及如何解决八种语言的各种问题:荷兰语、印地语、印尼语、挪威语、瑞典语、泰卢固语、泰语和越南语。我们使用BootCaT方法:我们从维基百科中获取该语言的一组“种子词”。然后,数百次,我们*随机选择三到四个种子词*作为查询发送到谷歌或雅虎或必应,后者返回一个‘搜索命中’页面*收集谷歌或雅虎指向的页面并保存文本。这形成了语料库,然后我们*‘清理’(以移除导航栏、广告等)*移除重复项*标记并(如果工具可用)词条和词性标记*加载到我们的语料库查询工具中,我们开发的素描引擎语料库可用于素描引擎语料库查询工具。
For many languages there are no large, general-language corpora available. Until the web, all but the institutions could do little but shake their heads in dismay as corpus-building was long, slow and expensive. But with the advent of the Web it can be highly automated and thereby fast and inexpensive. We have developed a corpus factory where we build large corpora. In this paper we describe the method we use, and how it has worked, and how various problems were solved, for eight languages: Dutch, Hindi, Indonesian, Norwegian, Swedish, Telugu, Thai and Vietnamese. We use the BootCaT method: we take a set of 'seed words' for the language from Wikipedia. Then, several hundred times over, we * randomly select three or four of the seed words * send as a query to Google or Yahoo or Bing, which returns a 'search hits' page * gather the pages that Google or Yahoo point to and save the text. This forms the corpus, which we then * 'clean' (to remove navigation bars, advertisements etc) * remove duplicates * tokenise and (if tools are available) lemmatise and part-of-speech tag * load into our corpus query tool, the Sketch Engine The corpora we have developed are available for use in the Sketch Engine corpus query tool.