MULTI-K: accurate classification of microarray subtypes using ensemble k-means clustering.

MULTI-K: accurate classification of microarray subtypes using ensemble k-means clustering.
复制标题

Multi-K:使用集合K均值聚类对微阵列亚型进行准确分类。

DOI:
10.1186/1471-2105-10-260
复制
发表时间:
2009-08-22
期刊:
影响因子:
3
通讯作者:
Nam D
Nam D
中科院分区:
生物学4区
文献类型:
--
作者:
Kim EY;Kim SY;Ashlock D;Nam D

文献摘要

参考文献

被引文献

相似文献

从微阵列样本中揭示疾病的亚型具有重要的临床意义,例如存活时间和个体患者对特定疗法的敏感性。无监督聚类方法已被用于对这类数据进行分类。然而,大多数现有的方法集中于具有紧凑形状的簇,并且没有反映高维微阵列簇的几何复杂性,这限制了它们的性能。我们提出了一个基于聚类数的集成聚类算法,称为MULTI-K,用于微阵列样本分类,它表现出显着的准确性。该方法通过改变聚类的数量来合并多个k-means运行,并识别出表现出最强大的元素共同成员关系的聚类。除了原来的算法,我们新设计的熵图来控制单态或小集群的分离。与简单的k-means或其他广泛使用的方法不同,MULTI-K能够准确捕获具有复杂和高维结构的聚类。MULTI-K优于其他方法,包括最近开发的集成聚类算法在测试中与五个模拟和八个真实的基因表达数据集。为了对微阵列数据进行准确的分类,需要考虑聚类的几何复杂性,而将集合聚类应用于聚类数目可以很好地解决这个问题。C++代码和测试的数据集可从作者那里获得。
Uncovering subtypes of disease from microarray samples has important clinical implications such as survival time and sensitivity of individual patients to specific therapies. Unsupervised clustering methods have been used to classify this type of data. However, most existing methods focus on clusters with compact shapes and do not reflect the geometric complexity of the high dimensional microarray clusters, which limits their performance. We present a cluster-number-based ensemble clustering algorithm, called MULTI-K, for microarray sample classification, which demonstrates remarkable accuracy. The method amalgamates multiple k-means runs by varying the number of clusters and identifies clusters that manifest the most robust co-memberships of elements. In addition to the original algorithm, we newly devised the entropy-plot to control the separation of singletons or small clusters. MULTI-K, unlike the simple k-means or other widely used methods, was able to capture clusters with complex and high-dimensional structures accurately. MULTI-K outperformed other methods including a recently developed ensemble clustering algorithm in tests with five simulated and eight real gene-expression data sets. The geometric complexity of clusters should be taken into account for accurate classification of microarray data, and ensemble clustering applied to the number of clusters tackles the problem very well. The C++ code and the data sets tested are available from the authors.
DOI: 10.1093/bioinformatics/bti483
发表时间: 2005-07-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Qiu, P;Wang, ZJ;Liu, KJR
通讯作者: Liu, KJR
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1002/j.1538-7305.1948.tb01338.x
发表时间: 1948-01-01
影响因子: --
作者:
SHANNON, CE
通讯作者: SHANNON, CE
DOI: 10.1093/bioinformatics/18.9.1194
发表时间: 2002-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Medvedovic, M;Sivaganesan, S
通讯作者: Sivaganesan, S
DOI: 10.1007/bf02294245
发表时间: 1985-01-01
期刊: PSYCHOMETRIKA
影响因子: 3
作者:
MILLIGAN, GW;COOPER, MC
通讯作者: COOPER, MC