Collocation Map for Overcoming Data Sparseness

Collocation Map for Overcoming Data Sparseness
复制标题

克服数据稀疏性的搭配图

DOI:
--
复制
发表时间:
1995
期刊:
Conference of the European Chapter of the Association for Computational Linguistics
影响因子:
--
通讯作者:
Key
Key
中科院分区:
--
文献类型:
--
作者:
Moonjoo Kim;Young S. Han;Key

文献摘要

被引文献

相似文献

统计语言模型之所以有用,是因为它们可以在不确定的决策中提供概率信息。最常见的统计量是n-grams测量文本中的单词共发生。但是,该方法遇到了数据短缺问题。在本文中,我们建议将贝叶斯网络用于近似发生不足的统计数据,以及在示例文本中没有出现优雅降解的统计数据。搭配图是一个可以从Bigram构建的Sigmoid信念网络。我们比较了从Bigrams和搭配图计算出的条件概率和共同信息。结果表明,搭配图的值的差异小于不频繁对的频率度量的差异48%。还证明了搭配图的预测能力未从样本文本中观察到的任意关联的预测能力。
Statistical language models are useful because they can provide probabilistic information upon uncertain decision making. The most common statistic is n-grams measuring word cooccurrences in texts. The method suffers from data shortage problem, however. In this paper, we suggest Bayesian networks be used in approximating the statistics of insufficient occurrences and of those that do not occur in the sample texts with graceful degradation. Collocation map is a sigmoid belief network that can be constructed from bigrams. We compared the conditional probabilities and mutual information computed from bigrams and Collocation map. The results show that the variance of the values from Collocation map is smaller than that from frequency measure for the infrequent pairs by 48%. The predictive power of Collocation map for arbitrary associations not observed from sample texts is also demonstrated.