Contextual Dependencies in Unsupervised Word Segmentation

Contextual Dependencies in Unsupervised Word Segmentation
复制标题

DOI:
10.3115/1220175.1220260
复制
发表时间:
2006-07
期刊:
--
影响因子:
--
通讯作者:
S. Goldwater;T. Griffiths;Mark Johnson
S. Goldwater;T. Griffiths;Mark Johnson
中科院分区:
其他
文献类型:
--
作者:
S. Goldwater;T. Griffiths;Mark Johnson

文献摘要

被引文献

相似文献

开发更好的方法将连续文本分割成单词对于改善亚洲语言的处理非常重要,并可能揭示人类如何学习分割语音。我们提出了两个新的贝叶斯分词方法,假设单字和二元模型的词依赖分别。bigram模型大大优于unigram模型(和以前的概率模型),证明了这种依赖关系对分词的重要性。我们还表明,以前的概率模型依赖于次优搜索程序至关重要。
Developing better methods for segmenting continuous text into words is important for improving the processing of Asian languages, and may shed light on how humans learn to segment speech. We propose two new Bayesian word segmentation methods that assume unigram and bigram models of word dependencies respectively. The bigram model greatly outperforms the unigram model (and previous probabilistic models), demonstrating the importance of such dependencies for word segmentation. We also show that previous probabilistic models rely crucially on sub-optimal search procedures.