MODELING DOCUMENTS WITH MULTIPLE POISSON-DISTRIBUTIONS

MODELING DOCUMENTS WITH MULTIPLE POISSON-DISTRIBUTIONS
复制标题

DOI:
10.1016/0306-4573(93)90005-x
复制
发表时间:
1993-03-01
影响因子:
8.6
通讯作者:
MARGULIS, EL
MARGULIS, EL
中科院分区:
计算机科学1区
文献类型:
--
作者:
MARGULIS, EL

文献摘要

被引文献

相似文献

本文是一项研究的初步报告,该研究试图寻找大量全文文档的有用统计模型。我们研究了文档集合中单词分布的多重泊松 (nP) 模型的有效性。 nP 分布是具有不同均值的 n 个泊松分布的混合。我们描述了一种实用的算法,用于确定某个单词是否根据 nP 分布进行分布。该算法应用于三个不同文档集合中的每个术语(精简词)。研究发现,超过 70% 的频繁出现项确实遵循 nP 分布。结果表明,对于文档长度相似的集合,nP 术语的比例甚至更高(超过 80%)。大多数已识别的 nP 项是根据相对较少的单一泊松分布(两个、三个或四个)的混合分布的。有迹象表明,混合物中单个泊松分量的数量取决于项的收集频率。
This paper is the initial report of a study attempting to find a useful statistical model of large collections of full text documents. We investigate the validity of the Multiple Poisson (nP) model of word distribution in document collections. An nP distribution is a mixture of n Poisson distributions with different means. We describe a practical algorithm for determining if a certain word is distributed according to an nP distribution. The algorithm is applied to every term (reduced word) in three different document collections. It was found that over 70% of frequently occurring terms indeed behave according to the nP distributions. The results indicate that the proportion of nP terms is even higher (over 80%) for the collections in which documents have similar length. Most of the nP terms recognised are distributed according to the mixture of relatively few single Poisson distributions (two, three, or four). There is an indication that the number of single Poisson components in the mixture depends on the collection frequency of terms.