Pólya urn model and its application to text categorization

Pólya urn model and its application to text categorization
复制标题

DOI:
10.4310/sii.2019.v12.n2.a4
复制
发表时间:
2019
影响因子:
0.8
通讯作者:
Haibin Zhang;Xianyi Wu;Xueqin Zhou
Haibin Zhang;Xianyi Wu;Xueqin Zhou
中科院分区:
数学4区
文献类型:
--
作者:
Haibin Zhang;Xianyi Wu;Xueqin Zhou

文献摘要

相似文献

P´olya urn 模型是广泛应用于统计和文本挖掘的基本模型。大多数训练模型的算法都非常缓慢且复杂,因此通常很难将 P´olya urn 模型拟合到大数据集。本文提出了一种新的最小化最大化(MM)算法,用于 P´olya urn 模型的最大似然估计(MLE),其中代理函数是通过简单的凸函数构造的。分析了MM算法的收敛性,并推导了非同分布观测值对应的MLE的渐近正态性。这种新的MM算法的性能也与牛顿法和其他MM算法进行了比较。 P´olya urn 模型应用于文本分类。真实的新闻组数据集证明了它相对于朴素贝叶斯 (NB) 分类器、k-近邻 (k-NN) 和支持向量机 (SVM) 的优越性。
P´olya urn model is a basic model widely applied in statistics and text mining. Most algorithms to training the model are very slow and complicated so that it generally difficult to fit a P´olya urn model to big data sets. This paper proposes a new minorization-maximization (MM) algorithm for the maximum likelihood estimation (MLE) of the P´olya urn model in which the surrogate function is constructed by means of a simple convex function. The convergence of the MM algorithm is analyzed and the asymptotic normality of the corresponding MLE for non-identically distributed observations is also derived. The performance of this new MM algorithm is also compared with Newton method and other MM algorithms. The P´olya urn model is applied to text categorization. Its superiority to naive Bayes (NB) classi-fier, k-Nearest Neighbor (k-NN) and support vector machine (SVM) are demonstrated by a real newsgroup dataset.