A Dictionary- and Corpus-Independent Statistical Lemmatizer for Information Retrieval in Low Resource Languages

A Dictionary- and Corpus-Independent Statistical Lemmatizer for Information Retrieval in Low Resource Languages
复制标题

用于低资源语言信息检索的独立于词典和语料库的统计词形还原器

DOI:
--
复制
发表时间:
2010
期刊:
Conference and Labs of the Evaluation Forum
影响因子:
--
通讯作者:
K. Järvelin
K. Järvelin
中科院分区:
--
文献类型:
--
作者:
Aki Loponen;K. Järvelin

文献摘要

被引文献

相似文献

我们提出了一个独立于字典和语料库的统计词法分析器StaLe,它通过为任何屈变词形生成候选词来处理基于字典的词法结构化的词汇外问题。对于缺乏语言资源的语言,可以毫不费力地应用StaLe。我们使用四种高资源语言的多个数据集和查询类型,展示了StaLe在单独的词源化任务中的性能,以及作为IR系统中的组件的性能。在红外实验中,StaLe具有较强的竞争力,达到了商用浸染剂金标准性能的88- 108%。尽管具有竞争力的性能,但它紧凑、高效、快速地应用于新语言。
We present a dictionary- and corpus-independent statistical lemmatizer StaLe that deals with the out-of-vocabulary (OOV) problem of dictionary-based lemmatization by generating candidate lemmas for any inflected word forms. StaLe can be applied with little effort to languages lacking linguistic resources. We show the performance of StaLe both in lemmatization tasks alone and as a component in an IR system using several datasets and query types in four high resource languages. StaLe is competitive, reaching 88-108 % of gold standard performance of a commercial lemmatizer in IR experiments. Despite competitive performance, it is compact, efficient and fast to apply to new languages.