Language-Independent Methods for Compiling Monolingual Lexical Data
Language-Independent Methods for Compiling Monolingual Lexical Data
复制标题
用于编译单语词汇数据的与语言无关的方法
DOI:
--
复制
发表时间:
2004
期刊:
影响因子:
--
通讯作者:
Christian Wolff
中科院分区:
文献类型:
--
作者:
Chris Biemann;Stefan Bordag;Gerhard Heyer;U. Quasthoff;Christian Wolff
In this paper we describe a flexible, portable and language-independent infrastructure for setting up large monolingual language corpora. The approach is based on collecting a large amount of monolingual text from various sources. The input data is processed on the basis of a sentence-based text segmentation algorithm. We describe the entry structure of the corpus database as well as various query types and tools for information extraction. Among them, the extraction and usage of sentence-based word collocations is discussed in detail. Finally we give an overview of different applications for this language resource. A WWW interface allows for public access to most of the data and information extraction tools (http://wortschatz.uni-leipzig.de).