Bayesian model-based clustering procedures

Bayesian model-based clustering procedures
复制标题

DOI:
10.1198/106186007x238855
复制
发表时间:
2007-09-01
影响因子:
2.4
通讯作者:
Green, Peter J.
Green, Peter J.
中科院分区:
数学2区
文献类型:
--
作者:
Lau, John W.;Green, Peter J.

文献摘要

被引文献

相似文献

本文建立了基于贝叶斯模型的聚类的一般公式,其中子集标签是可交换的,项目也是可交换的,可能达到协变量效应。的符号框架是足够丰富的,以涵盖各种现有的程序,包括一些最近讨论的方法,涉及随机搜索或分层聚类,但更重要的是,允许制定的聚类程序是最佳的相对于一个指定的损失函数。我们的重点是基于成对重合的损失函数,也就是说,是否成对的项目被聚集到相同的subset.Optimization的后验预期损失函数可以制定为一个二进制整数规划问题,这可以很容易地解决标准的软件时,聚类适度数量的项目,但很快就变得不切实际的问题规模的增加。为了解决这个问题,一个新的启发式项目交换算法。这表现良好,在我们的数值实验,模拟和真实的数据的例子。本文包括(近似)最优聚类与早期方法的统计性能的比较,这些方法是基于节点的,但在其详细定义中是特别的。
This article establishes a general formulation for Bayesian model-based clustering, in which subset labels are exchangeable, and items are also exchangeable, possibly up to covariate effects. The notational framework is rich enough to encompass a variety of existing procedures, including some recently discussed methods involving stochastic search or hierarchical clustering, but more importantly allows the formulation of clustering procedures that are optimal with respect to a specified loss function. Our focus is on loss functions based on pairwise coincidences, that is, whether pairs of items are clustered into the same subset or not.Optimization of the posterior expected loss function can be formulated as a binary integer programming problem, which can be readily solved by standard software when clustering a modest number of items, but quickly becomes impractical as problem scale increases. To combat this, a new heuristic item-swapping algorithm is introduced. This performs well in our numerical experiments, on both simulated and real data examples. The article includes a comparison of the statistical performance of the (approximate) optimal clustering with earlier methods that are rnodel-based but ad hoc in their detailed definition.