Cross-Clustering: A Partial Clustering Algorithm with Automatic Estimation of the Number of Clusters.

Cross-Clustering: A Partial Clustering Algorithm with Automatic Estimation of the Number of Clusters.
复制标题

DOI:
10.1371/journal.pone.0152333
复制
发表时间:
2016
期刊:
影响因子:
3.7
通讯作者:
Drăghici S
Drăghici S
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Tellaroli P;Bazzi M;Donato M;Brazzale AR;Drăghici S

文献摘要

被引文献

相似文献

许多可用的聚类方法的四个最常见的限制是:i)缺乏处理异常值的适当策略; ii) 需要对聚类数量进行良好的先验估计以获得合理的结果; iii) 缺乏能够检测特定数据集的划分何时不合适的方法; iv) 结果对初始化的依赖性。在这里,我们提出了交叉聚类(CC),这是一种部分聚类算法,它通过结合两种成熟的层次聚类算法的原理克服了这四个限制:Ward 最小方差和完全链接。我们通过将 CC 与许多现有的聚类方法(包括 Ward 聚类方法和 Complete-linkage 聚类方法)进行比较来对其进行验证。我们在模拟和真实数据集上都表明,CC 在以下方面比其他方法表现更好:识别正确的聚类数量、识别异常值以及确定真实的聚类成员资格。我们使用 CC 对样本进行聚类,以识别疾病亚型,并使用基因图谱来确定具有相同行为的基因组。在非生物数据集上获得的结果表明该方法足够通用,可以成功地用于如此多样化的应用。该算法已使用统计语言 R 实现,并且可以从 CRAN 贡献包存储库免费获取。
Four of the most common limitations of the many available clustering methods are: i) the lack of a proper strategy to deal with outliers; ii) the need for a good a priori estimate of the number of clusters to obtain reasonable results; iii) the lack of a method able to detect when partitioning of a specific data set is not appropriate; and iv) the dependence of the result on the initialization. Here we propose Cross-clustering (CC), a partial clustering algorithm that overcomes these four limitations by combining the principles of two well established hierarchical clustering algorithms: Ward’s minimum variance and Complete-linkage. We validated CC by comparing it with a number of existing clustering methods, including Ward’s and Complete-linkage. We show on both simulated and real datasets, that CC performs better than the other methods in terms of: the identification of the correct number of clusters, the identification of outliers, and the determination of real cluster memberships. We used CC to cluster samples in order to identify disease subtypes, and on gene profiles, in order to determine groups of genes with the same behavior. Results obtained on a non-biological dataset show that the method is general enough to be successfully used in such diverse applications. The algorithm has been implemented in the statistical language R and is freely available from the CRAN contributed packages repository.