Acquiring a Poor Man’s Inflectional Lexicon for German

Acquiring a Poor Man’s Inflectional Lexicon for German
复制标题

获取穷人的德语屈折词典

DOI:
--
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
P. Adolphs
P. Adolphs
中科院分区:
--
文献类型:
--
作者:
P. Adolphs

文献摘要

被引文献

相似文献

许多 NLP 模块和应用程序需要一个模块来进行广泛覆盖的词形变化分析。获得此类分析的一种方法是结合使用形态分析器和词形变化词典。由于当今大型文本语料库很容易获得,并且词形变化系统通常很好理解,因此在我们的词形变化知识的指导下从原始文本获取词汇数据似乎是可行的。我按照这些思路为德语提供了一种获取方法。总体思路可以大致概括为:首先,为语料库中的每个词形生成一组词条假设;然后,选择解释语料库“best”中发现的词形的假设。为此,我将现有的形态语法(采用有限状态技术(Schmid et al. 2004))转变为词汇条目的假设器。简单地列出不规则形式,以便它们不会干扰假设者使用的规则规则。在文本语料库上运行假设器会产生大量的词汇输入假设。然后借助统计模型根据其有效性对这些模型进行排名,该统计模型基于每个假设的经过验证和预测的单词形式的数量。
Many NLP modules and applications require the availability of a module for wide-coverage inflectional analysis. One way to obtain such analyses is to use a morphological analyser in combination with an inflectional lexicon. Since large text corpora nowadays are easily available and inflectional systems are in general well understood, it seems feasible to acquire lexical data from raw texts, guided by our knowledge of inflection. I present an acquisition method along these lines for German. The general idea can be roughly summarised as follows: first, generate a set of lexical entry hypotheses for each word-form in the corpus; then, select hypotheses that explain the word-forms found in the corpus “best”. To this end, I have turned an existing morphological grammar, cast in finite-state technology (Schmid et al. 2004), into a hypothesiser for lexical entries. Irregular forms are simply listed so that they do not interfere with the regular rules used in the hypothesiser. Running the hypothesiser on a text corpus yields a large number of lexical entry hypotheses. These are then ranked according to their validity with the help of a statistical model that is based on the number of attested and predicted word forms for each hypothesis.