Challenges in microarray class discovery: a comprehensive examination of normalization, gene selection and clustering.

Challenges in microarray class discovery: a comprehensive examination of normalization, gene selection and clustering.
复制标题

DOI:
10.1186/1471-2105-11-503
复制
发表时间:
2010-10-11
期刊:
影响因子:
3
通讯作者:
Rydén P
Rydén P
中科院分区:
生物学4区
文献类型:
--
作者:
Freyhult E;Landfors M;Önskog J;Hvidsten TR;Rydén P

文献摘要

被引文献

相似文献

聚类分析,特别是层次聚类,被广泛用于从基因表达数据中提取信息。其目的是发现个体或基因的新类别或子类。执行聚类分析通常涉及如何决策;处理缺失值,标准化数据,选择基因。此外,预处理,包括各种类型的过滤和标准化程序,可以对发现生物学相关类的能力产生影响。在这里,我们从广义上考虑聚类分析,并对包括归一化在内的聚类分析的几个方面进行全面评估。我们评估了2780种聚类分析方法,这些方法采用了7种公开的双通道微阵列数据集,具有共同的参考设计。每种聚类分析方法在数据归一化(考虑了5种归一化)、缺失值归一化(2种)、数据标准化(2种)、基因选择(19种)或聚类方法(11种)方面存在差异。聚类分析使用已知的类别(如癌症类型)和调整后的Rand指数进行评估。不同分析的性能在不同的数据集之间有所不同,很难给出一般的建议。然而,归一化、基因选择和聚类方法都是影响性能的重要变量。特别是,基因选择很重要,为了获得良好的性能,通常需要包含相对大量的基因。选择高标准偏差的基因或采用主成分分析是首选的基因选择方法。采用Ward方法的分层聚类、k-means聚类和Mclust聚类是本文考虑的获得最高调整后Rand的聚类方法。归一化可以对聚类个体的能力产生显著的积极影响,并且有迹象表明背景校正是可取的,特别是如果基因选择成功的话。然而,这是一个需要进一步研究的领域,以便得出任何一般性结论。聚类分析的选择,特别是基因选择,对基于表达谱正确聚类个体的能力有很大影响。归一化具有积极的作用,但不同归一化的相对性能是一个需要更多研究的领域。总之,尽管聚类、基因选择和归一化被认为是生物信息学中的标准方法,但我们的综合分析表明,选择正确的方法和正确的方法组合远非微不足道,而且在被认为是最基本的基因组数据分析中还有很多尚未探索。
Cluster analysis, and in particular hierarchical clustering, is widely used to extract information from gene expression data. The aim is to discover new classes, or sub-classes, of either individuals or genes. Performing a cluster analysis commonly involve decisions on how to; handle missing values, standardize the data and select genes. In addition, pre-processing, involving various types of filtration and normalization procedures, can have an effect on the ability to discover biologically relevant classes. Here we consider cluster analysis in a broad sense and perform a comprehensive evaluation that covers several aspects of cluster analyses, including normalization. We evaluated 2780 cluster analysis methods on seven publicly available 2-channel microarray data sets with common reference designs. Each cluster analysis method differed in data normalization (5 normalizations were considered), missing value imputation (2), standardization of data (2), gene selection (19) or clustering method (11). The cluster analyses are evaluated using known classes, such as cancer types, and the adjusted Rand index. The performances of the different analyses vary between the data sets and it is difficult to give general recommendations. However, normalization, gene selection and clustering method are all variables that have a significant impact on the performance. In particular, gene selection is important and it is generally necessary to include a relatively large number of genes in order to get good performance. Selecting genes with high standard deviation or using principal component analysis are shown to be the preferred gene selection methods. Hierarchical clustering using Ward's method, k-means clustering and Mclust are the clustering methods considered in this paper that achieves the highest adjusted Rand. Normalization can have a significant positive impact on the ability to cluster individuals, and there are indications that background correction is preferable, in particular if the gene selection is successful. However, this is an area that needs to be studied further in order to draw any general conclusions. The choice of cluster analysis, and in particular gene selection, has a large impact on the ability to cluster individuals correctly based on expression profiles. Normalization has a positive effect, but the relative performance of different normalizations is an area that needs more research. In summary, although clustering, gene selection and normalization are considered standard methods in bioinformatics, our comprehensive analysis shows that selecting the right methods, and the right combinations of methods, is far from trivial and that much is still unexplored in what is considered to be the most basic analysis of genomic data.