Unsupervised Disambiguation of Syncretism in Inflected Lexicons

Unsupervised Disambiguation of Syncretism in Inflected Lexicons
复制标题

DOI:
10.18653/v1/n18-2087
复制
发表时间:
2018-06
期刊:
--
影响因子:
--
通讯作者:
Ryan Cotterell;Christo Kirov;Sabrina J. Mielke;Jason Eisner
Ryan Cotterell;Christo Kirov;Sabrina J. Mielke;Jason Eisner
中科院分区:
其他
文献类型:
--
作者:
Ryan Cotterell;Christo Kirov;Sabrina J. Mielke;Jason Eisner

文献摘要

相似文献

词汇歧义使得计算语料库的有用统计数据变得困难。一个给定的词形可能代表几个形态特征束中的任何一个。然而,人们可以使用无监督学习(如EM)来拟合一个概率消除单词形式歧义的模型。我们提出了这样一种方法,它采用神经网络来平滑地对特征束(即使是罕见的)的先验分布进行建模。虽然这个基本模型没有考虑令牌的上下文,但是这个属性允许它对一个简单的unigram类型计数列表进行操作,将每个计数划分为该unigram的不同分析。我们讨论了这个新任务的评估指标,并报告了5种语言的结果。
Lexical ambiguity makes it difficult to compute useful statistics of a corpus. A given word form might represent any of several morphological feature bundles. One can, however, use unsupervised learning (as in EM) to fit a model that probabilistically disambiguates word forms. We present such an approach, which employs a neural network to smoothly model a prior distribution over feature bundles (even rare ones). Although this basic model does not consider a token’s context, that very property allows it to operate on a simple list of unigram type counts, partitioning each count among different analyses of that unigram. We discuss evaluation metrics for this novel task and report results on 5 languages.