A Web Corpus and Word Sketches for Japanese

A Web Corpus and Word Sketches for Japanese
复制标题

日语网络语料库和单词草图

DOI:
10.5715/jnlp.15.2_137
复制
发表时间:
2008
影响因子:
--
通讯作者:
A. Kilgarriff
A. Kilgarriff
中科院分区:
--
文献类型:
--
作者:
Irena Srdanovic Erjavec;T. Erjavec;A. Kilgarriff

文献摘要

被引文献

相似文献

在世界上所有主要语言中,日语在公开访问和搜索语料库方面落后。在本文中,我们描述了JpWaC(日语网络语料库),一个大型语料库的4亿字的日本网络文本,其编码的草图引擎的发展。Sketch Engine是一个基于Web的语料库查询工具,支持快速索引,语法处理,“单词草图”(单词语法和搭配行为的一页摘要),分布式词库和机器人使用。我们描述了收集和处理语料库,并建立其有效性的步骤,在它包含的语言种类。然后,我们描述了一个浅的语法日语,使单词素描的发展。我们相信,加载到Sketch Engine中的日语网络语料库将成为大量日本研究人员,学习者和NLP开发人员的有用资源。
Of all the major world languages, Japanese is lagging behind in terms of publicly accessible and searchable corpora. In this paper we describe the development of JpWaC (Japanese Web as Corpus), a large corpus of 400 million words of Japanese web text, and its encoding for the Sketch Engine. The Sketch Engine is a web-based corpus query tool that supports fast concordancing, grammatical processing, ‘word sketching’ (one-page summaries of a word’s grammatical and collocational behaviour), a distributional thesaurus, and robot use. We describe the steps taken to gather and process the corpus and to establish its validity, in terms of the kinds of language it contains. We then describe the development of a shallow grammar for Japanese to enable word sketching. We believe that the Japanese web corpus as loaded into the Sketch Engine will be a useful resource for a wide number of Japanese researchers, learners, and NLP developers.