Automated construction and evaluation of Japanese Web-based reference corpora

Automated construction and evaluation of Japanese Web-based reference corpora
复制标题

日语网络参考语料库的自动化构建与评价

DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
Marco Baroni
Marco Baroni
中科院分区:
--
文献类型:
--
作者:
M. Ueyama;Marco Baroni

文献摘要

被引文献

相似文献

利用网络进行语言学研究的一个特别有前途的方法是通过自动查询搜索引擎来建立语料库,检索和后处理以这种方式找到的页面(Ghani et al. 2003,Baroni and Bernardini 2004,Sharoff)。这种方法不同于传统的语料库建设方法,需要花费大量的时间来寻找和选择要包含的文本,但对内容有很好的控制和意识。使用基于Web的自动语料库构建,情况正好相反:人们可以在很短的时间内构建语料库,但无法很好地控制语料库中的文本类型。
A particularly promising approach to the use of the Web for linguistic research is to build corpora via automated queries to search engines, retrieving and post-processing the pages found in this way (Ghani et al. 2003, Baroni and Bernardini 2004, Sharoff to appear). This approach differs from the traditional method of corpus construction, where one needs to spend considerable time finding and selecting the texts to be included, but has perfect control and awareness over contents. With automated Webbased corpus construction, the situation is reversed: one can build a corpus in very little time, but without a good control over what kinds of texts are in the corpus.