Statistical and Biological Validation Methods in Cluster Analysis of Gene Expression

Statistical and Biological Validation Methods in Cluster Analysis of Gene Expression
复制标题

基因表达聚类分析中的统计和生物学验证方法

DOI:
--
复制
发表时间:
2007
期刊:
International Conference on Machine Learning and Applications
影响因子:
--
通讯作者:
Milton Pires Ramos
Milton Pires Ramos
中科院分区:
--
文献类型:
--
作者:
Daniele Yumi Sunaga;J. C. Nievola;Milton Pires Ramos

文献摘要

参考文献

被引文献

相似文献

数据聚类方法已成为基因表达数据分析的标准技术。它们被用于各种任务,从简单的数据预处理后验分析到重要信息的识别,如基因功能和/或一组基因在给定生物过程中的参与。从经济学的角度来看,数据聚类方法也为生物学家提供了优势,并且考虑到在没有智能计算方法的帮助下获得这类信息所必需的时间。这项工作旨在指导选择,以便从数据聚类中获得最佳可能的解决方案。为此,使用了不同方法的算法,即k-means和SOM算法属于一维方法,SAMBA算法属于二维方法。采用统计学和生物学验证的方法,选择最佳的数据聚类方案。结果表明,统计验证方法与生物学验证方法难以一致。此外,我们还观察到了SOM算法相对于k-means算法的一些优势。使用二维算法SAMBA揭示了单维算法无法识别的数据集结构。将有意义的生物信息聚合到功能未知的基因上是可能的。这项工作的所有内容,包括所有的数据聚类和详细分析都可以在URL http://www.ppgia.pucpr.br/~nievola/clusteranalysis上获得。
Data clustering methods have become standard techniques in the analysis of gene expression data. They are used in a variety of tasks ranging from simple data pre- treatment for posterior analysis to the identification of important information, such as gene function and/or the participation of a group of genes in a given biological process. Data clustering methods also offer advantages to the biologist from the economic point of view and given the time that would be necessary to obtain this type of information without the aid of intelligent computational methods. This work aims at guiding the choices in order to get the best possible solution from data clustering. To do so, algorithms from different approaches were used, i.e. k-means and SOM algorithms belong to the unidimentional approach and SAMBA algorithm, a bidimentional approach. Methods of statistical and biological validation were employed in order to choose the best data clustering solution. Results presented here demonstrated that the statistic validation methods were hardly in agreement with the biology validation method. Furthermore, some advantages of the SOM algorithm over the k-means algorithm were observed. Use of the bidimentional algorithm SAMBA revealed dataset structure not identified by the unidimentional algorithms. It was possible to aggregate meaningfull biological information to genes of unknown function. All the content of this work, including all the data clustering and detailed analysis are available at the URL http://www.ppgia.pucpr.br/~nievola/clusteranalysis.
DOI: 10.1091/mbc.9.12.3273
发表时间: 1998-12-01
影响因子: 3.3
作者:
Spellman, PT;Sherlock, G;Futcher, B
通讯作者: Futcher, B
DOI: 10.1073/pnas.95.25.14863
发表时间: 1998-12-08
影响因子: 11.1
作者:
Eisen, MB;Spellman, PT;Botstein, D
通讯作者: Botstein, D