Estimating Lexical Priors for Low-Frequency Syncretic Forms

Estimating Lexical Priors for Low-Frequency Syncretic Forms
复制标题

估计低频融合形式的词汇先验

DOI:
--
复制
发表时间:
1995
期刊:
arXiv.org
影响因子:
--
通讯作者:
R. Sproat
R. Sproat
中科院分区:
--
文献类型:
--
作者:
Harald Baayen;R. Sproat

文献摘要

被引文献

相似文献

给定一个以前看不见的形式,它在形态学上是n路模糊的,对于该形式的各种函数,词汇先验概率的最佳估计是什么?我们认为,最好的估计是通过计算hapax legomena之间的各种功能的相对频率-在语料库中只出现一次的形式。这一结果具有重要意义的发展随机形态标签,特别是当一些初始的手标记的语料库是必需的:预测词汇先验非常低的频率形态模糊的类型(其中大部分不会发生在任何给定的语料库),应该集中在标记一个很好的代表性样本的hapax legomena,而不是广泛的标记的所有频率范围内的话。
Given a previously unseen form that is morphologically n-ways ambiguous, what is the best estimator for the lexical prior probabilities for the various functions of the form? We argue that the best estimator is provided by computing the relative frequencies of the various functions among the hapax legomena --- the forms that occur exactly once in a corpus. This result has important implications for the development of stochastic morphological taggers, especially when some initial hand-tagging of a corpus is required: For predicting lexical priors for very low-frequency morphologically ambiguous types (most of which would not occur in any given corpus) one should concentrate on tagging a good representative sample of the hapax legomena, rather than extensively tagging words of all frequency ranges.