A novel clustering approach and prediction of optimal number of clusters: global optimum search with enhanced positioning

A novel clustering approach and prediction of optimal number of clusters: global optimum search with enhanced positioning
复制标题

一种新颖的聚类方法和最佳聚类数量的预测:具有增强定位的全局最优搜索

DOI:
--
复制
发表时间:
2007
影响因子:
1.8
通讯作者:
C. Floudas
C. Floudas
中科院分区:
数学3区
文献类型:
--
作者:
M. P. Tan;J. Broach;C. Floudas

文献摘要

参考文献

被引文献

相似文献

来自DNA微阵列杂交研究的全基因组表达数据的聚类分析是用于鉴定生物学相关基因分组的有用工具(DeRisi et al. 1997; Weiler et al. 1997年)。因此,重要的是应用严格而直观的聚类算法来揭示这些基因组关系。在这项研究中,我们描述了一种新的聚类算法框架的基础上的一个变种的广义弯曲分解,表示为全局最优搜索(Floudas等。1989年; Floudas 1995年),其中包括一个程序,以确定最佳数量的集群使用。该方法涉及数据点的预聚类,以定义聚类的初始数量以及线性规划问题(原始问题)和混合线性规划问题(主问题)的迭代解,这些问题源自混合非线性规划问题的公式化。错误放置的数据点被删除以形成新的聚类,从而确保数据点之间的紧密分组,并增加聚类的数量,直到达到最佳数量。我们提出的聚类算法集中在Ras信号通路的酵母酿酒酵母的实验DNA微阵列数据,并比较与一些常用的聚类算法得到的结果。我们的算法相比,这些算法在类内相似性和类间不相似性方面,往往被认为是两个关键的聚类原则。此外,我们的算法可以预测的最佳数量的集群,并预测集群的生物一致性进行了分析,通过基因本体。
Cluster analysis of genome-wide expression data from DNA microarray hybridization studies is a useful tool for identifying biologically relevant gene groupings (DeRisi et al. 1997; Weiler et al. 1997). It is hence important to apply a rigorous yet intuitive clustering algorithm to uncover these genomic relationships. In this study, we describe a novel clustering algorithm framework based on a variant of the Generalized Benders Decomposition, denoted as the Global Optimum Search (Floudas et al. 1989; Floudas 1995), which includes a procedure to determine the optimal number of clusters to be used. The approach involves a pre-clustering of data points to define an initial number of clusters and the iterative solution of a Linear Programming problem (the primal problem) and a Mixed-Integer Linear Programming problem (the master problem), that are derived from a Mixed Integer Nonlinear Programming problem formulation. Badly placed data points are removed to form new clusters, thus ensuring tight groupings amongst the data points and incrementing the number of clusters until the optimum number is reached. We apply the proposed clustering algorithm to experimental DNA microarray data centered on the Ras signaling pathway in the yeast Saccharomyces cerevisiae and compare the results to that obtained with some commonly used clustering algorithms. Our algorithm compares favorably against these algorithms in the aspects of intra-cluster similarity and inter-cluster dissimilarity, often considered two key tenets of clustering. Furthermore, our algorithm can predict the optimal number of clusters, and the biological coherence of the predicted clusters is analyzed through gene ontology.
DOI: 10.1101/gr.9.11.1106
发表时间: 1999-11-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Heyer, LJ;Kruglyak, S;Yooseph, S
通讯作者: Yooseph, S
DOI: 10.1126/science.278.5338.680
发表时间: 1997-10-24
期刊: SCIENCE
影响因子: 56.9
作者:
DeRisi, JL;Iyer, VR;Brown, PO
通讯作者: Brown, PO
DOI: 10.1073/pnas.95.25.14863
发表时间: 1998-12-08
影响因子: 11.1
作者:
Eisen, MB;Spellman, PT;Botstein, D
通讯作者: Botstein, D