Weakly Supervised Part-of-Speech Tagging for Morphologically-Rich, Resource-Scarce Languages

Weakly Supervised Part-of-Speech Tagging for Morphologically-Rich, Resource-Scarce Languages
复制标题

针对形态丰富、资源稀缺的语言的弱监督词性标记

DOI:
10.3115/1609067.1609107
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
Vincent Ng
Vincent Ng
中科院分区:
--
文献类型:
--
作者:
K. Hasan;Vincent Ng

文献摘要

参考文献

被引文献

相似文献

本文研究了词性标注的无监督方法,用于词性标注丰富的,资源稀缺的语言,重点介绍了Goldwater和Griffiths(2007)最初为英语词性标注开发的全贝叶斯方法。我们认为现有的无监督词性标注器不切实际地假设一个完美的词性词汇作为输入,因此,我们提出了一种弱监督的全贝叶斯词性标注方法,该方法通过从少量的词性标注数据中自动获取词性词汇来缓解这种不切实际的假设。由于这种放松是以标记准确性下降为代价的,我们提出了贝叶斯框架的两个扩展,并证明它们有效地改进了孟加拉语的全贝叶斯POS标记器,孟加拉语是我们的代表性的形态学丰富,资源稀缺的语言。
This paper examines unsupervised approaches to part-of-speech (POS) tagging for morphologically-rich, resource-scarce languages, with an emphasis on Goldwater and Griffiths's (2007) fully-Bayesian approach originally developed for English POS tagging. We argue that existing unsupervised POS taggers unrealistically assume as input a perfect POS lexicon, and consequently, we propose a weakly supervised fully-Bayesian approach to POS tagging, which relaxes the unrealistic assumption by automatically acquiring the lexicon from a small amount of POS-tagged data. Since such relaxation comes at the expense of a drop in tagging accuracy, we propose two extensions to the Bayesian framework and demonstrate that they are effective in improving a fully-Bayesian POS tagger for Bengali, our representative morphologically-rich, resource-scarce language.
DOI: 10.1109/tpami.1984.4767596
发表时间: 1984-01-01
影响因子: 23.6
作者:
GEMAN, S;GEMAN, D
通讯作者: GEMAN, D