Classification of binary vectors by stochastic complexity

Classification of binary vectors by stochastic complexity
复制标题

DOI:
10.1006/jmva.1997.1687
复制
发表时间:
1997-10-01
影响因子:
1.6
通讯作者:
Verlaan, M
Verlaan, M
中科院分区:
数学2区
文献类型:
--
作者:
Gyllenberg, M;Koski, T;Verlaan, M

文献摘要

被引文献

相似文献

随机复杂性被视为分类的工具,即,对于给定的二进制向量数据集,推断类的数量、类描述和类成员关系。根据最大熵原理得到的多元贝努利分布的有限混合所定义的统计模型族,对随机复杂性进行了评估。结果表明,随机复杂度与分类最大似然估计是渐近相关的。随机复杂性公式的解释是某些通用源代码的最小代码长度,用于将二进制数据向量及其赋值存储到分类中的类别。还可以将分类不确定性分解为类内不确定性、类间不确定性和特殊简约项的总和。结果表明,最小化随机复杂度相当于最大化分类的信息量。给出了一种交替最小化随机复杂度的算法。讨论了该方法与贝叶斯自动分类系统的关系。描述了随机复杂性分类方法在肠杆菌科菌株数据库中的应用。(C)1997年学术出版社。
Stochastic complexity is treated as a tool of classification, i.e., of inferring the number of classes, the class descriptions, and the class memberships for a given data set of binary vectors. The stochastic complexity is evaluated with respect to the family of statistical models defined by finite mixtures of multivariate Bernoulli distributions obtained by the principle of maximum entropy. It is shown that stochastic complexity is asymptotically related to the classification maximum likelihood estimate. The Formulae for stochastic complexity have an interpretation as minimum code lengths for certain universal source codes for storing the binary data Vectors and their assignments into the classes in a classification. There is also a decomposition of the classification uncertainty in a sum of an intraclass uncertainty, an interclass uncertainty, and a special parsimony term. It is shown that minimizing the stochastic complexity amounts to maximizing the information content of the classification. An algorithm of alternating minimization of stochastic complexity is given. We discuss the relation of the method to the AUTOCLASS system of Bayesian classification. The application of classification by stochastic complexity to an extensive data base of strains of Enterobacteriaceae is described. (C) 1997 Academic Press.