A "Density-Based" Algorithm for Cluster Analysis Using Species Sampling Gaussian Mixture Models

A "Density-Based" Algorithm for Cluster Analysis Using Species Sampling Gaussian Mixture Models
复制标题

DOI:
10.1080/10618600.2013.856796
复制
发表时间:
2014-10-02
影响因子:
2.4
通讯作者:
Guglielmi, Alessandra
Guglielmi, Alessandra
中科院分区:
数学2区
文献类型:
--
作者:
Argiento, Raffaele;Cremaschi, Andrea;Guglielmi, Alessandra

文献摘要

被引文献

相似文献

我们提出了一个新的模型在贝叶斯非参数框架的聚类分析。我们的模型结合了两种成分,一方面是高斯分布的物种采样混合模型,另一方面是确定性聚类过程(DBSCAN)。在这里,如果对应于其潜在参数的密度之间的距离小于阈值,则来自底层物种采样混合物模型的两个观测值共享相同的聚类;这产生了比由物种采样混合物引起的随机分区更粗糙的随机分区。由于这个过程依赖于阈值的值,我们提出了一个策略来修复它。此外,我们讨论了该模型的实现和应用;比较更标准的聚类算法也将被给出。这篇文章的补充材料可在网上查阅。
We propose a new model for cluster analysis in a Bayesian nonparametric framework. Our model combines two ingredients, species sampling mixture models of Gaussian distributions on one hand, and a deterministic clustering procedure (DBSCAN) on the other. Here, two observations from the underlying species sampling mixture model share the same cluster if the distance between the densities corresponding to their latent parameters is smaller than a threshold; this yields a random partition which is coarser than the one induced by the species sampling mixture. Since this procedure depends on the value of the threshold, we suggest a strategy to fix it. In addition, we discuss implementation and applications of the model; comparison with more standard clustering algorithms will be given as well. Supplementary materials for the article are available online.