Interpolating between types and tokens by estimating power-law generators

Interpolating between types and tokens by estimating power-law generators
复制标题

DOI:
--
复制
发表时间:
2005-12
期刊:
--
影响因子:
--
通讯作者:
S. Goldwater;T. Griffiths;Mark Johnson
S. Goldwater;T. Griffiths;Mark Johnson
中科院分区:
其他
文献类型:
--
作者:
S. Goldwater;T. Griffiths;Mark Johnson

文献摘要

被引文献

相似文献

语言的标准统计模型无法捕捉自然语言最显著的特性之一:单词标记频率的幂律分布。我们提出了一个框架,用于开发一般产生幂律的统计模型,增强标准生成模型与适配器,产生适当的模式的令牌频率。我们表明,以一个特定的随机过程-Pitman-Yor过程-作为适配器证明了自然语言的形式分析中出现的类型频率,并提高了无监督学习的形态学模型的性能。
Standard statistical models of language fail to capture one of the most striking properties of natural languages: the power-law distribution in the frequencies of word tokens. We present a framework for developing statistical models that generically produce power-laws, augmenting standard generative models with an adaptor that produces the appropriate pattern of token frequencies. We show that taking a particular stochastic process - the Pitman-Yor process - as an adaptor justifies the appearance of type frequencies in formal analyses of natural language, and improves the performance of a model for unsupervised learning of morphology.