Language-Independent Methods for Compiling Monolingual Lexical Data

Language-Independent Methods for Compiling Monolingual Lexical Data
复制标题

用于编译单语词汇数据的与语言无关的方法

DOI:
--
复制
发表时间:
2004
期刊:
Conference on Intelligent Text Processing and Computational Linguistics
影响因子:
--
通讯作者:
Christian Wolff
Christian Wolff
中科院分区:
--
文献类型:
--
作者:
Chris Biemann;Stefan Bordag;Gerhard Heyer;U. Quasthoff;Christian Wolff

文献摘要

被引文献

相似文献

在本文中,我们描述了一个灵活的,可移植的和语言无关的基础设施,建立大型单语语料库。该方法是基于从各种来源收集大量的单语文本。输入数据的处理基于一个基于文本分割算法。我们描述了语料库的条目结构以及各种查询类型和信息提取工具。其中,对基于语义的词语搭配的提取和使用进行了详细的讨论。最后,我们给出了这个语言资源的不同应用的概述。万维网界面允许公众访问大多数数据和信息提取工具(wortschatz.uni-leipzig.de)。
In this paper we describe a flexible, portable and language-independent infrastructure for setting up large monolingual language corpora. The approach is based on collecting a large amount of monolingual text from various sources. The input data is processed on the basis of a sentence-based text segmentation algorithm. We describe the entry structure of the corpus database as well as various query types and tools for information extraction. Among them, the extraction and usage of sentence-based word collocations is discussed in detail. Finally we give an overview of different applications for this language resource. A WWW interface allows for public access to most of the data and information extraction tools (http://wortschatz.uni-leipzig.de).