Computational cluster validation for microarray data analysis: experimental assessment of Clest, Consensus Clustering, Figure of Merit, Gap Statistics and Model Explorer.

Computational cluster validation for microarray data analysis: experimental assessment of Clest, Consensus Clustering, Figure of Merit, Gap Statistics and Model Explorer.
复制标题

DOI:
10.1186/1471-2105-9-462
复制
发表时间:
2008-10-29
期刊:
影响因子:
3
通讯作者:
Utro F
Utro F
中科院分区:
生物学4区
文献类型:
--
作者:
Giancarlo R;Scaturro D;Utro F

文献摘要

参考文献

被引文献

相似文献

从微阵列数据中推断簇结构是所谓的组学科学的一项基本任务。这也是统计、数据分析和分类中的一个基本问题,特别是在预测数据集中的聚类数量方面,通常通过内部验证措施建立。尽管在文献中提供了丰富的内部措施,最近提出了新的,其中一些专门用于微阵列数据。我们考虑五个这样的措施:Clest,共识(共识聚类),FOM(品质指数),差距(差距统计)和ME(模型浏览器),除了经典的WCSS(内集群平方和)和KL(Krzanowski和Lai指数)。我们进行了广泛的实验,六个基准微阵列数据集,使用分层和K-均值聚类算法,我们提供了一个分析评估的内在能力的措施,以预测正确的聚类数在一个数据集和它的优点相对于其他措施。我们特别注重精度和速度。此外,我们还提供了各种快速近似算法的计算间隙,FOM和WCSS。主要的结果是这些措施的精度和速度方面的层次结构,突出了一些文献中没有报道的优点和局限性。基于我们的分析,我们得出了几个结论,使用这些内部措施的微阵列数据。我们报告主要的。就预测能力和显著的算法独立性而言,共识是迄今为止表现最好的。不幸的是,在大型数据集上,它可能没有用处,因为它的计算机时间需求很大(在最先进的PC上需要数周)。FOM是表现第二好的,尽管令人惊讶的是,它在这种情况下可能没有竞争力:它基本上具有与WCSS相同的预测能力,但根据数据集的不同,它在时间上慢了6到100倍。用于计算FOM、Gap和WCSS的近似算法执行得非常好,即,它们更快,同时仍然非常接近FOM和WCSS。用于计算Gap的近似算法值得被挑选出来,因为它的预测能力远远好于Gap,它与其他度量具有竞争力,但它在时间上比Gap快至少两个数量级。从我们的分析中可以得出的另一个重要的新结论是,我们所考虑的所有度量在大型数据集上都表现出严重的局限性,这要么是由于计算需求(共识,如前所述,Clest和Gap),要么是由于缺乏精度(所有其他度量,包括它们的近似值)。软件和数据集可在补充材料网页上的GNU GPL下获得。
Inferring cluster structure in microarray datasets is a fundamental task for the so-called -omic sciences. It is also a fundamental question in Statistics, Data Analysis and Classification, in particular with regard to the prediction of the number of clusters in a dataset, usually established via internal validation measures. Despite the wealth of internal measures available in the literature, new ones have been recently proposed, some of them specifically for microarray data. We consider five such measures: Clest, Consensus (Consensus Clustering), FOM (Figure of Merit), Gap (Gap Statistics) and ME (Model Explorer), in addition to the classic WCSS (Within Cluster Sum-of-Squares) and KL (Krzanowski and Lai index). We perform extensive experiments on six benchmark microarray datasets, using both Hierarchical and K-means clustering algorithms, and we provide an analysis assessing both the intrinsic ability of a measure to predict the correct number of clusters in a dataset and its merit relative to the other measures. We pay particular attention both to precision and speed. Moreover, we also provide various fast approximation algorithms for the computation of Gap, FOM and WCSS. The main result is a hierarchy of those measures in terms of precision and speed, highlighting some of their merits and limitations not reported before in the literature. Based on our analysis, we draw several conclusions for the use of those internal measures on microarray data. We report the main ones. Consensus is by far the best performer in terms of predictive power and remarkably algorithm-independent. Unfortunately, on large datasets, it may be of no use because of its non-trivial computer time demand (weeks on a state of the art PC). FOM is the second best performer although, quite surprisingly, it may not be competitive in this scenario: it has essentially the same predictive power of WCSS but it is from 6 to 100 times slower in time, depending on the dataset. The approximation algorithms for the computation of FOM, Gap and WCSS perform very well, i.e., they are faster while still granting a very close approximation of FOM and WCSS. The approximation algorithm for the computation of Gap deserves to be singled-out since it has a predictive power far better than Gap, it is competitive with the other measures, but it is at least two order of magnitude faster in time with respect to Gap. Another important novel conclusion that can be drawn from our analysis is that all the measures we have considered show severe limitations on large datasets, either due to computational demand (Consensus, as already mentioned, Clest and Gap) or to lack of precision (all of the other measures, including their approximations). The software and datasets are available under the GNU GPL on the supplementary material web page.
DOI: 10.1007/bf02294245
发表时间: 1985-01-01
期刊: PSYCHOMETRIKA
影响因子: 3
作者:
MILLIGAN, GW;COOPER, MC
通讯作者: COOPER, MC
DOI: 10.1038/35000501
发表时间: 2000-02-03
期刊: NATURE
影响因子: 64.8
作者:
Alizadeh, AA;Eisen, MB;Staudt, LM
通讯作者: Staudt, LM
DOI: 10.1091/mbc.9.12.3273
发表时间: 1998-12-01
影响因子: 3.3
作者:
Spellman, PT;Sherlock, G;Futcher, B
通讯作者: Futcher, B
Genclust:用于聚类基因表达数据的遗传算法。
DOI: 10.1186/1471-2105-6-289
发表时间: 2005-12-07
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Di Gesú, V;Giancarlo, R;Lo Bosco, G;Raimondi, A;Scaturro, D
通讯作者: Scaturro, D
DOI: 10.1016/j.jmva.2004.02.002
发表时间: 2004-07-01
影响因子: 1.6
作者:
McLachlan, GJ;Khan, N
通讯作者: Khan, N