GenClust: a genetic algorithm for clustering gene expression data.

GenClust: a genetic algorithm for clustering gene expression data.
复制标题

Genclust:用于聚类基因表达数据的遗传算法。

DOI:
10.1186/1471-2105-6-289
复制
发表时间:
2005-12-07
期刊:
影响因子:
3
通讯作者:
Scaturro, D
Scaturro, D
中科院分区:
生物学4区
文献类型:
--
作者:
Di Gesú, V;Giancarlo, R;Lo Bosco, G;Raimondi, A;Scaturro, D

文献摘要

参考文献

被引文献

相似文献

聚类是基因表达数据分析中的关键步骤,事实上,许多经典的聚类算法被使用,或者更多的创新算法已经被设计和验证。尽管人工智能技术在生物信息学和更普遍的数据分析中广泛使用,但基于遗传范式的聚类算法很少,但该范式在为困难的优化问题(如聚类)找到良好的启发式解决方案方面具有很大的潜力。GenClust是一种新的用于基因表达数据聚类的遗传算法。它有两个主要特点:(a)简单、紧凑且易于更新的搜索空间的新颖编码;(B)其可以自然地结合数据驱动的内部验证方法使用。我们已经试验了FOM方法,专门用于验证基因表达数据的聚类。GenClust的有效性已经在真实的数据集上进行了实验评估,既使用了验证措施,又与其他算法进行了比较,即,平均链接、投射、点击和K均值。实验表明,我们使用的算法在数据集和验证措施上都没有明显上级其他算法;即,在许多情况下,观察到的最差和最佳执行算法之间的差异在统计上可能是不显著的,并且它们可以被认为是等效的。然而,在某些情况下,一种算法可能比其他算法更好,因此是值得的。特别是,GenClust的实验表明,虽然简单的数据表示,它非常迅速地收敛到一个局部最优,它的能力,以确定有意义的集群是可比的,有时上级,更复杂的算法。此外,它非常适合与数据驱动的内部验证措施,特别是FOM方法结合使用。
Clustering is a key step in the analysis of gene expression data, and in fact, many classical clustering algorithms are used, or more innovative ones have been designed and validated for the task. Despite the widespread use of artificial intelligence techniques in bioinformatics and, more generally, data analysis, there are very few clustering algorithms based on the genetic paradigm, yet that paradigm has great potential in finding good heuristic solutions to a difficult optimization problem such as clustering. GenClust is a new genetic algorithm for clustering gene expression data. It has two key features: (a) a novel coding of the search space that is simple, compact and easy to update; (b) it can be used naturally in conjunction with data driven internal validation methods. We have experimented with the FOM methodology, specifically conceived for validating clusters of gene expression data. The validity of GenClust has been assessed experimentally on real data sets, both with the use of validation measures and in comparison with other algorithms, i.e., Average Link, Cast, Click and K-means. Experiments show that none of the algorithms we have used is markedly superior to the others across data sets and validation measures; i.e., in many cases the observed differences between the worst and best performing algorithm may be statistically insignificant and they could be considered equivalent. However, there are cases in which an algorithm may be better than others and therefore worthwhile. In particular, experiments for GenClust show that, although simple in its data representation, it converges very rapidly to a local optimum and that its ability to identify meaningful clusters is comparable, and sometimes superior, to that of more sophisticated algorithms. In addition, it is well suited for use in conjunction with data driven internal validation measures and, in particular, the FOM methodology.
DOI: 10.1016/s0031-3203(01)00108-x
发表时间: 2002-06-01
影响因子: 8
作者:
Bandyopadhyay, S;Maulik, U
通讯作者: Maulik, U
DOI: 10.1093/bioinformatics/btg330
发表时间: 2003-12-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Gat-Viks, I;Sharan, R;Shamir, R
通讯作者: Shamir, R
DOI: 10.1007/bf02294245
发表时间: 1985-01-01
期刊: PSYCHOMETRIKA
影响因子: 3
作者:
MILLIGAN, GW;COOPER, MC
通讯作者: COOPER, MC
DOI: 10.1016/s1097-2765(00)80114-8
发表时间: 1998-07-01
期刊: MOLECULAR CELL
影响因子: 16
作者:
Cho, RJ;Campbell, MJ;Davis, RW
通讯作者: Davis, RW
DOI: 10.1073/pnas.96.6.2907
发表时间: 1999-03-16
影响因子: 11.1
作者:
Tamayo, P;Slonim, D;Golub, TR
通讯作者: Golub, TR